Tiejun Huang 0001

dblp:h/TiejunHuang · also Tie-Jun Huang 0001, Tie-jun Huang 0001, TieJun Huang 0001 · DBLP profile ↗
← Back
370ranked-venue papers
2as first author
184since 2021 · last 2026
0000-0002-4234-6099ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 240 · 103 since 2021Artificial intelligence and machine learning · 182 · 134 since 2021Databases, data management, data science and information retrieval · 22 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 6 since 2021Systems, architecture and hardware · 9 · 3 since 2021Computer networks · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Spike Stream Memory Transfer for Dynamic Scene Reconstruction
abstract
As a retina-inspired sensor with ultra-high temporal resolution, spike camera can continuously capture dynamic scenes with high-speed motion. It is a key task to restore clear images from spike streams. The quantization effects in spike readout bring degradation to the visual quality of restored images. To tackle the degradation without introducing motion blur, existing methods often employ a short-term temporal window to infer the light intensity at a certain time point. However, these methods only focus on the spike signals within the current window, which limits their performance. Motivated by the human-like memory mechanism for visual signals from the retina, we explore Spike Stream Memory Transfer (SSMT) to restore the dynamic scenes, considering spike signals beyond the window. Specifically, we design a framework that leverages temporal memory by transferring previously inferred light intensity and motion to enhance current reconstruction. The framework enables a long-term temporal perception of spike streams to handle the spike quantization effects. Besides, we utilize the estimated motion to suppress the potential blur from inter-stream clips, considering the underlying motion of spike streams. We also develop a spike interval-guided alignment module to tackle the blur from intra-stream clips. Experimental results on both synthetic and real-captured data demonstrate that our method can restore high-quality images from spike streams.
Yanchen Dong 0001, Ruiqin Xiong, Rui Zhao 0010, Xinfeng Zhang 0001, Tiejun Huang 0001
AAAI5
2026 Splats in Splats: Robust and Effective 3D Steganography Towards Gaussian Splatting
abstract
3D Gaussian splatting (3DGS) has demonstrated impressive 3D reconstruction performance with explicit scene representations. Given the widespread application of 3DGS in 3D reconstruction and generation tasks, there is an urgent need to protect the copyright of 3DGS assets. However, existing copyright protection techniques for 3DGS overlook the usability of 3D assets, posing challenges for practical deployment. Here we describe splats in splats, the first 3DGS steganography framework that embeds 3D content in 3DGS itself without modifying any attributes. To achieve this, we take a deep insight into spherical harmonics (SH) and devise an importance-graded SH coefficient encryption strategy to embed the hidden SH coefficients. Furthermore, we employ a convolutional autoencoder to establish a mapping between the original Gaussian primitives' opacity and the hidden Gaussian primitives' opacity. Extensive experiments indicate that our method significantly outperforms existing 3D steganography techniques, with 5.31% higher scene fidelity and 3x faster rendering speed, while ensuring security, robustness, and user experience.
Yijia Guo, Wenkai Huang 0003, Gaolei Li, Hang Zhang 0010, Liwen Hu 0002, Jianhua Li 0001, Tiejun Huang 0001, Lei Ma 0008
AAAI8
2026 Can Protective Watermarking Safeguard the Copyright of 3D Gaussian Splatting?
abstract
3D Gaussian Splatting (3DGS) has emerged as a powerful representation for 3D scenes, widely adopted due to its exceptional efficiency and high-fidelity visual quality. Given the significant value of 3DGS assets, recent works have introduced specialized watermarking schemes to ensure copyright protection and ownership verification. However, can existing 3D Gaussian watermarking approaches genuinely guarantee robust protection of the 3D assets? In this paper, for the first time, we systematically explore and validate possible vulnerabilities of 3DGS watermarking frameworks. We demonstrate that conventional watermark removal techniques designed for 2D images do not effectively generalize to the 3DGS scenario due to the specialized rendering pipeline and unique attributes of each gaussian primitives. Motivated by this insight, we propose GSPure, the first watermark purification framework specifically for 3DGS watermarking representations. By analyzing view-dependent rendering contributions and exploiting geometrically accurate feature clustering, GSPure precisely isolates and effectively removes watermark-related Gaussian primitives while preserving scene integrity. Extensive experiments demonstrate that our GSPure achieves the best watermark purification performance, reducing watermark PSNR by up to 16.34dB while minimizing degradation to original scene fidelity with less than 1dB PSNR loss. Moreover, it consistently outperforms existing methods in both effectiveness and generalization.
Wenkai Huang 0003, Yijia Guo, Gaolei Li, Lei Ma 0008, Hang Zhang 0010, Liwen Hu 0002, Jiazheng Wang 0001, Jianhua Li 0001, Tiejun Huang 0001
AAAI9
2026 Generalized Threshold Optimization with Harmony Multi-Threshold Neurons for Accurate ANN-to-SNN Conversion
abstract
Spiking Neural Networks(SNNs) are a promising paradigm designed to emulate the brain's energy efficient by incorporating the timing of spikes. Conversion is an efficient way to obtain high-performance SNNs from Artificial Neural Networks(ANNs). Existing conversion methods often face a trade-off between accuracy and time steps, which is largely caused by the incomplete release of residual membrane potentials. To minimize the conversion error, this paper proposed a harmonious mathematical property-based neuron, called Harmony Multi-Threshold Neurons (H-MT Neuron), which utilizes multiple spikes to minimize residual membrane potentials. The proposed neuron is further enhanced with an optional effective communication mechanism to achieve more accurate conversion. In addition, we propose a threshold optimization method applicable to a broader range cases of spiking neurons to to find the optimal neuron thresholds. Experiment results demonstrate that our method achieve superior accuracy on ImageNet benchmark datasets while significantly reducing the required time steps and energy consumption.
Zihan Huang, Tong Bu, Tiejun Huang 0001, Zhaofei Yu
AAAI4
2026 Spike Imaging Velocimetry: Dense Motion Estimation of Fluids Using Spike Streams
abstract
Particle Image Velocimetry (PIV) is a widely adopted non-invasive imaging technique that tracks the motion of tracer particles across image sequences to capture the velocity distribution of fluid flows. It is commonly employed to analyze complex flow structures and validate numerical simulations. This study explores the untapped potential of spike cameras—ultra-high-speed, high-dynamic-range vision sensors—in high-speed fluid velocimetry. We propose a deep learning framework, Spike Imaging Velocimetry (SIV), tailored for high-resolution fluid motion estimation. To enhance the network’s performance, we design three novel modules specifically adapted to the characteristics of fluid dynamics and spike streams: the Detail-Preserving Hierarchical Transform (DPHT), the Graph Encoder (GE), and the Multi-scale Velocity Refinement (MSVR). Furthermore, we introduce a spike-based PIV dataset, Particle Scenes with Spike and Displacement (PSSD), which contains labeled samples from three representative fluid-dynamics scenarios: steady turbulence, high-speed flow, and high-dynamic-range conditions. Our proposed method outperforms existing baselines across all these scenarios, demonstrating its effectiveness.
Yunzhong Zhang, Changqing Su, Zhen Cheng 0005, Zhaofei Yu, Tiejun Huang 0001, Xun Cao
AAAI7
2026 SpikeCV: open a continuous computer vision era
Yajing Zheng, Jiyuan Zhang 0005, Rui Zhao 0010, Jianhao Ding, Shiyan Chen, Weijian Wu, Ruiqin Xiong, Zhaofei Yu, Tiejun Huang 0001
Sci. China Inf. Sci.9
2026 S2E: Spatio-temporal filtering of spike streams for motion-selective event generation
Lingxiao Zheng, Yajing Zheng, Rui Zhao 0021, Tiejun Huang 0001
Neural Networks4
2026 Learn to Enhance Sparse Spike Streams
abstract
High-speed vision tasks have long been a challenge in computer vision. Recently, the spike camera has shown great potential in these tasks due to its high temporal resolution. Unlike traditional cameras, it emits asynchronous spike signals to capture visual information. However, under low-light conditions, spike signals become highly sparse, and the sparse spike stream severely hinders the effectiveness of existing spike-based methods in high-speed scenarios. To address this challenge, we introduce SS2DS, the first deep learning framework that enhances sparse spike streams into dense spike streams. SS2DS first estimates the spike firing frequency within sparse streams. Subsequently, the spike firing frequency is enhanced by a neural network. Finally, SS2DS decodes the enhanced spike stream from the enhanced spike firing frequency sequence. SS2DS can adjust the temporal distribution of sparse spike streams and improve the performance degradation of existing methods in low-light and high-speed scenarios. To evaluate sparse spike stream enhancement, we construct both synthetic and real sparse spike stream datasets. By comparing the reconstruction results, enhanced spike streams achieve an average improvement of +0.78 MA, -18.42 BRISQUE, and -1.42 NIQE over sparse spike streams. Moreover, the enhanced spike streams also benefit other spike-based vision tasks, such as 3D reconstruction (+1.325 dB PSNR, +0.005 SSIM, and -0.01 LPIPS) and superresolution (+0.63 MA, -13.67 BRISQUE, and -1.28 NIQE).
Liwen Hu 0002, Yijia Guo, Mianzhi Liu, Shengbo Chen, Lei Ma 0008, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack Against No-Reference Image Quality Assessment Models
abstract
No-Reference Image Quality Assessment (NR-IQA) models play an important role in various real-world applications. Recently, adversarial attacks against NR-IQA models have attracted increasing attention, as they provide valuable insights for revealing model vulnerabilities and guiding robust system design. Some effective attacks have been proposed against NR-IQA models in white-box settings, where the attacker has full access to the target model. However, these attacks often suffer from poor transferability to unknown target models in more realistic black-box scenarios, where the target model is inaccessible. This work makes the first attempt to address the challenge of low transferability in attacking NR-IQA models by proposing a transferable Signed Ensemble Gaussian black-box Attack (SEGA). The main idea is to approximate the gradient of the target model by applying Gaussian smoothing to source models and ensembling their smoothed gradients. To ensure the imperceptibility of adversarial perturbations, SEGA further removes inappropriate perturbations using a specially designed perturbation filter mask. Experimental results demonstrate the superior transferability of SEGA, validating its effectiveness in enabling successful transfer-based black-box attacks against NR-IQA models.
Yujia Liu 0005, Dingquan Li, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Spike Camera Optical Flow Estimation Based on Continuous Spike Streams
abstract
Spike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting binary spike streams. Optical flow estimation aims to estimate pixel-level correspondences between different moments, describing motion information along time, which is a key task of spike camera. High-quality optical flow is important since motion information is a foundation for analyzing spikes. However, extracting stable light-intensity information from spikes is difficult due to the randomness of binary spikes. Besides, the continuity of spikes can offer contextual information for optical flow. In this paper, we propose a network Spike2Flow++ to estimate optical flow for spike camera. In Spike2Flow++, we propose a differential of spike firing time (DSFT) to represent information in binary spikes. Moreover, we propose a dual DSFT representation and a dual correlation construction to extract stable light-intensity information for reliable correlations. To use the continuity of spikes as motion contextual information, we propose a joint correlation decoding (JCD) that jointly estimates a series of flow fields. To adaptively fuse different motions in JCD, we propose a global motion bank aggregation to construct an information bank for all motions and adaptively extract contexts from the bank for each iteration during recurrent decoding of each motion. To train and evaluate our network, we construct a real scene with spikes and flow++ (RSSF++) based on real-world scenes. Experiments demonstrate that our Spike2Flow++ achieves state-of-the-art performance on RSSF++, photo-realistic high-speed motion (PHM), and real-captured data.
Rui Zhao 0010, Ruiqin Xiong, Dongkai Wang, Shiyu Xuan, Jian Zhang 0018, Xiaopeng Fan 0001, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 SNNTracker: Online High-Speed Multi-Object Tracking With Spike Camera
abstract
Multi-object tracking (MOT) is crucial for applications such as autonomous driving and robotics, yet traditional image-based methods struggle in high-speed scenarios due to motion blur and temporal gaps caused by low frame rates. Spike cameras, with their ability to continuously record spatiotemporal signals, overcome these limitations. However, existing spike-based methods often rely on intermediate image reconstruction or discrete clustering, limiting real-time performance and temporal continuity. To address this, we propose SNNTracker, the first fully spiking neural network (SNN)-based MOT algorithm tailored for spike cameras. SNNTracker integrates a dynamic neural field (DNF)-based attention mechanism for target detection and a winner-take-all (WTA)-based tracking module with online spike-timing-dependent plasticity (STDP) for adaptive learning of object trajectories. By directly processing spike streams without reconstruction, SNNTracker reduces latency, computational overhead, and dependency on image quality, making it ideal for ultra-high-speed environments. It maintains robust, continuous tracking even under occlusions, severe lighting variations, or temporary object disappearance, by leveraging SNN-estimated motion predictions and long-term online clustering. We construct three types of spike-camera MOT datasets covering dense and sparse annotations across diverse real-world scenarios, including camera ego-motion, deformable and ultra-fast motion (up to 2600 RPM), occlusion, indoor/outdoor lighting changes, and low-visibility tracking. Extensive experiments demonstrate that SNNTracker consistently outperforms state-of-the-art MOT methods-both ANN- and SNN-based-achieving MOTA scores above 96% and up to 100% in many sequences. Our results highlight the advantages of spike-driven SNNs for low-latency, high-speed, and label-free multi-object tracking, advancing neuromorphic vision for real-time perception.
Yajing Zheng, Chengen Li, Jiyuan Zhang 0005, Zhaofei Yu, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Rethinking High-speed Image Reconstruction Framework with Spike Camera
abstract
Spike cameras, as innovative neuromorphic devices, generate continuous spike streams to capture high-speed scenes with lower bandwidth and higher dynamic range than traditional RGB cameras. However, reconstructing high-quality images from the spike input under low-light conditions remains challenging. Conventional learning-based methods often rely on the synthetic dataset as the supervision for training. Still, these approaches falter when dealing with noisy spikes fired under the low-light environment, leading to further performance degradation in the real-world dataset. This phenomenon is primarily due to inadequate noise modelling and the domain gap between synthetic and real datasets, resulting in recovered images with unclear textures, excessive noise, and diminished brightness. To address these challenges, we introduce a novel spike-to-image reconstruction framework SpikeCLIP that goes beyond traditional training paradigms. Leveraging the CLIP model's powerful capability to align text and images, we incorporate the textual description of the captured scene and unpaired high-quality datasets as the supervision. Textual descriptions provide additional context that guides the network's feature reconstruction, while high-quality datasets help produce sharp latent images. Our experiments on real-world low-light datasets U-CALTECH and U-CIFAR demonstrate that SpikeCLIP significantly enhances texture details and the luminance balance of recovered images. Furthermore, the reconstructed images are well-aligned with the broader visual features needed for downstream tasks, ensuring more robust and versatile performance in challenging environments.
Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu
AAAI3
2025 SpikeGS: Reconstruct 3D Scene Captured by a Fast-Moving Bio-Inspired Camera
abstract
3D Gaussian Splatting (3DGS) has been proven to exhibit exceptional performance in reconstructing 3D scenes. However, the effectiveness of 3DGS heavily relies on sharp images, and fulfilling this requirement presents challenges in real-world scenarios particularly when utilizing fast-moving cameras. This limitation severely constrains the practical application of 3DGS and may compromise the feasibility of real-time reconstruction. To mitigate these challenges, we proposed Spike Gaussian Splatting (SpikeGS), the first framework that integrates the Bayer-pattern spike streams into the 3DGS pipeline to reconstruct 3D scenes captured by a fast-moving high temporal color spike camera in one second. With accumulation rasterization, interval supervision, and a special designed pipeline, SpikeGS realizes continuous spatiotemporal perception while extracts detailed structure and texture from Bayer-pattern spike stream which is unstable and lacks details. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of SpikeGS compared with existing spike-based and deblur 3D scene reconstruction methods.
Yijia Guo, Liwen Hu 0002, Yuanxi Bai, Jiawei Yao, Lei Ma 0008, Tiejun Huang 0001
AAAI6
2025 Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and Analysis
abstract
Elucidating the functional mechanisms of the primary visual cortex (V1) remains a fundamental challenge in systems neuroscience. Current computational models face two critical limitations, namely the challenge of cross-modal integration between partial neural recordings and complex visual stimuli, and the inherent variability in neural characteristics across individuals, including differences in neuron populations and firing patterns. To address these challenges, we present a multi-modal identifiable variational autoencoder (miVAE) that employs a two-level disentanglement strategy to map neural activity and visual stimuli into a unified latent space. This framework enables robust identification of cross-modal correlations through refined latent space modeling. We complement this with a novel score-based attribution analysis that traces latent variables back to their origins in the source data space. Evaluation on a large-scale mouse V1 dataset demonstrates that our method achieves state-of-the-art performance in cross-individual latent representation and alignment, without requiring subject-specific fine-tuning, and exhibits improved performance with increasing data size. Significantly, our attribution algorithm successfully identifies distinct neuronal subpopulations characterized by unique temporal patterns and stimulus discrimination properties, while simultaneously revealing stimulus regions that show specific sensitivity to edge features and luminance variations. This scalable framework offers promising applications not only for advancing V1 research but also for broader investigations in neuroscience.
Yu Zhu 0008, Chunfeng Song, Wanli Ouyang, Tiejun Huang 0001
AAAI6
2025 MLVU: Benchmarking Multi-task Long Video Understanding
abstract
The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multitask Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, such as movies, surveillance, egocentric videos, and cartoons, reflects the models’ LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs’ key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today’s technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of LVU by providing a comprehensive and in-depth analysis of MLLMs. The code and dataset can be accessed from https://github.com/JUNJIE99/MLVU.
Junjie Zhou 0001, Bo Zhao 0015, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Yongping Xiong, Tiejun Huang 0001, Zheng Liu 0011
CVPR11
2025 Self-Supervised Learning for Color Spike Camera Reconstruction
abstract
Spike camera is a kind of neuromorphic camera with ultra-high temporal resolution, which can capture dynamic scenes by continuously firing spike signals. To capture color information, a color filter array (CFA) is employed on the sensor of the spike camera, resulting in Bayer-pattern spike streams. How to restore high-quality color images from the binary spike signals remains challenging. In this paper, we propose a motion-guided reconstruction method for spike cameras with CFA, utilizing color layout and estimated motion information. Specifically, we develop a joint motion estimation pipeline for the Bayer-pattern spike stream, exploiting the motion consistency of channels. We propose to estimate the missing pixels of each color channel according to temporally neighboring pixels of the corresponding color along the motion trajectory. As the spike signals are read out at discrete time points, there is quantization noise that impacts the image quality. Thus, we analyze the correlation of the noise in spatial and temporal domains and propose a self-supervised network utilizing a masked spike encoder to handle the noise. Experiments on real-world captured Bayer-pattern spike streams show that our method can restore color images with better visual quality, compared with state-of-the-art methods. The source codes are available at https://github.com/csycdong/SSL-CSC.
Yanchen Dong 0001, Ruiqin Xiong, Xiaopeng Fan 0001, Zhaofei Yu, Yonghong Tian 0001, Tiejun Huang 0001
CVPR6
2025 USP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian Splatting
abstract
Spike camera, as an innovative type of neuromorphic camera that captures scenes with 0-1 bit stream at 40 kHz, is increasingly being employed for the novel view synthesis task building on techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). Previous spike-based approaches typically follow a three-stage pipeline: I. Spike-to-image reconstruction based on established algorithms. II. Camera poses estimation. III. Novel view synthesis. However, the cascading framework suffers from substantial cumulative errors, i.e., the quality of the initially re-constructed images will impact pose estimation, ultimately limiting the fidelity of the 3D reconstruction. To address this limitation, we propose a synergistic optimization framework USP-Gaussian, which unifies spike-to-image reconstruction, pose correction, and gaussian splatting into an end-to-end pipeline. Leveraging the multi-view consistency afforded by 3DGS and the motion capture capability of the spike camera, our framework enables iterative optimization between the spike-to-image reconstruction network and 3DGS. Experiments on synthetic datasets demonstrate that our method surpasses previous approaches by effectively eliminating cascading errors. Moreover, in real-world scenarios, our method achieves robust 3D reconstruction benefiting from the integration of pose optimization. Our code, data, and trained models are available at https://github.com/chenkang455/USP-Gaussian.
Jiyuan Zhang 0005, Zecheng Hao, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu
CVPR5
2025 You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
abstract
Recent 3D generation models typically rely on limited-scale 3D ‘gold-labels’ or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrained 3D priors due to the lack of scalable learning paradigms. In this work, we present See3D, a visual-conditional multi-view diffusion model trained on large-scale Internet videos for open-world 3D creation. The model aims to Get 3D knowledge by solely Seeing the visual contents from the vast and rapidly growing video data — You See it, You Got it. To achieve this, we first scale up the training data using a proposed data curation pipeline that automatically filters out multi-view inconsistencies and insufficient observations from source videos. This results in a high-quality, richly diverse, large-scale dataset of multi-view images, termed WebVi3D, containing 320M frames from 16M video clips. Nevertheless, learning generic 3D priors from videos without explicit 3D geometry or camera pose annotations is nontrivial, and annotating poses for web-scale videos is prohibitively expensive. To eliminate the need for pose conditions, we introduce an innovative visual-condition - a purely 2D-inductive visual signal generated by adding time-dependent noise to the masked video data. Finally, we introduce a novel visual-conditional 3D generation framework by integrating See3D into a warping-based pipeline for high-fidelity 3D generation. Our numerical and visual comparisons on single and sparse reconstruction benchmarks show that See3D, trained on cost- effective and scalable video data, achieves notable zero-shot and open-world generation capabilities, markedly outperforming models trained on costly and constrained 3D datasets. Additionally, our model naturally supports other image-conditioned 3D creation tasks, such as 3D editing, without further fine-tuning. Please refer to our project page at: https://vision.baai.ac.cn/see3d.
Baorui Ma, Huachen Gao, Haoge Deng, Tiejun Huang 0001, Lulu Tang
CVPR5
2025 Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
abstract
Long video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos. Although several existing methods attempt to reduce visual tokens, their strategies encounter severe bottleneck, restricting MLLMs’ ability to perceive fine-grained visual details. In this work, we propose Video-XL, a novel approach that leverages MLLMs’ inherent key-value (KV) sparsification capacity to condense the visual input. Specifically, we introduce a new special token, the Visual Summarization Token (VST), for each interval of the video, which summarizes the visual information within the interval as its associated KV. The VST module is trained by instruction fine-tuning, where two optimizing strategies are offered. 1. Curriculum learning, where VST learns to make small (easy) and large compression (hard) progressively. 2. Composite data curation, which integrates single-image, multi-image, and synthetic data to overcome the scarcity of long-video instruction data. The compression quality is further improved by dynamic compression, which customizes compression granularity based on the information density of different video intervals. Video-XL’s effectiveness is verified from three aspects. First, it achieves a superior long-video understanding capability, outperforming state-of-the-art models of comparable sizes across multiple popular benchmarks. Second, it effectively preserves video information, with minimal compression loss even at 16 × compression ratio. Third, it realizes outstanding cost-effectiveness, enabling high-quality processing of thousands of frames on a single A100 GPU.
Zheng Liu 0011, Peitian Zhang, Minghao Qin, Junjie Zhou 0001, Zhengyang Liang, Tiejun Huang 0001, Bo Zhao 0015
CVPR7
2025 Spk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative Filtering
abstract
Spike camera is a kind of neuromorphic camera that records dynamic scenes by firing a stream of binary spikes with extremely high temporal resolution. It demonstrates great potential for vision tasks in high-speed scenarios. One limitation in its current implementation is the relatively low spatial resolution. This paper develops a network called Spk2SRImgNet to super-resolve high resolution images from low resolution spike stream. However, fluctuations in spike stream hinder the performance of spike camera super resolution. To address this issue, we propose a motion aligned collaborative filtering (MACF) module, which is motivated by key ideas in classic image restoration schemes to mitigate fluctuations in spike data. MACF leverages the temporal similarity of spike stream to acquire similar features from neighboring moments via motion alignment. To separate disturbances from features, MACF filters these similar features jointly in transform domain to exploit representation sparsity, and generates refinement features that will be used to update initial fluctuated features. Specifically, MACF designs an inverse motion alignment operation to map these refinement features back to their original positions. The initial features are aggregated with the repositioned refinement features to enhance reliability. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared with existing methods.
Yuanlin Wang, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Tiejun Huang 0001
CVPR7
2025 OmniGen: Unified Image Generation
abstract
The emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single framework remains largely unexplored. In this work, we introduce OmniGen, a new diffusion model for unified image generation. OmniGen is characterized by the following features: 1) Unification: OmniGen not only demonstrates text-to-image generation capabilities but also inherently supports various downstream tasks, such as image editing, subject-driven generation, and visual-conditional generation. 2) Simplicity: The architecture of OmniGen is highly simplified, eliminating the need for additional plugins. Moreover, compared to existing diffusion models, it is more user-friendly and can complete complex tasks end-to-end through instructions without the need for extra intermediate steps, greatly simplifying the image generation workflow. 3) Knowledge Transfer: Benefit from learning in a unified format, OmniGen effectively transfers knowledge across different tasks, manages unseen tasks and domains, and exhibits novel capabilities. We also explore the model’s reasoning capabilities and potential applications of the chain-of-thought mechanism. This work represents the first attempt at a general-purpose image generation model, and we will release our resources at https://github.com/VectorSpaceLab/OmniGen to foster future advancements.
Shitao Xiao, Yueze Wang, Junjie Zhou 0001, Huaying Yuan, Xingrun Xing, Ruiran Yan, Shuting Wang 0002, Tiejun Huang 0001, Zheng Liu 0011
CVPR9
2025 Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
abstract
Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent foreseeable ones proactively. To this end, we propose Code-as-Monitor (CaM), a novel paradigm leveraging the vision-language model (VLM) for both open-set reactive and proactive failure detection. The core of our method is to formulate both tasks as a unified set of spatio-temporal constraint satisfaction problems and use VLM-generated code to evaluate them for real-time monitoring. To enhance the accuracy and efficiency of monitoring, we further introduce constraint elements that abstract constraint-related entities or their parts into compact geometric elements. This approach offers greater generality, simplifies tracking, and facilitates constraint-aware visual programming by leveraging these elements as visual prompts. Experiments show that CaM achieves a 28.7% higher success rate and reduces execution time by 31.8% under severe disturbances compared to baselines across three simulators and a real-world setting. Moreover, CaM can be integrated with open-loop control policies to form closed-loop systems, enabling long-horizon tasks in cluttered scenes with dynamic environments. See the project page at https://zhoues.github.io/Code-as-Monitor.
Enshen Zhou, Cheng Chi 0001, Zhizheng Zhang 0011, Zhongyuan Wang 0006, Tiejun Huang 0001, Lu Sheng, He Wang 0010
CVPR6
2025 A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission
abstract
The large volume of data from the point cloud brings significant demands on network bandwidth. However, the current transmission framework only considers using lossy compression to control the size of data, while ignoring visually redundant information due to the setting of rendering devices. Based on the fact that point overlapping might occur for the case that a dense point cloud is rendered on a relatively low resolution 2D monitor, we propose a novel quality-aware sampling framework for point cloud transmission. When a target visual quality is determined, an optimal sampling module is designed to remove overlapped points with the help of a simple but effective quality model. By taking into account the impact of multiple factors (i.e., sampling, lossy compression, and client rendering resolution), this quality model can predict the final perceptual quality in the client. Based on a newly constructed dataset which consists of 420 samples, experiment results show that the proposed transmission framework can significantly reduce bandwidth cost (e.g., 6.10% to 84.43%) and processing time (e.g., 8.99% to 92.53%) without introducing noticeable distortion under certain rendering conditions, thus achieving higher bandwidth utilization and better real-time performance.
Puyue Hou, Qi Yang 0003, Yue Li 0015, Jianchao Yang, Yiling Xu, Tiejun Huang 0001
ICASSP7
2025 ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient Learning
Yuanlin Wang, Ruiqin Xiong, Rui Zhao 0010, Jin Wang 0023, Xiaopeng Fan 0001, Tiejun Huang 0001
ICCV6
2025 SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action Recognition
Rui Zhao 0010, Ruiqin Xiong, Xiaopeng Fan 0001, Tiejun Huang 0001
ICCV6
2025 SAN: Hypothesizing Long-Term Synaptic Development and Neural Engram Mechanism in Scalable Model's Parameter-Efficient Fine-Tuning
abstract
Advances in Parameter-efficient Fine-tuning (PEFT) bridged the performance gap with Full Fine-Tuning (FFT) through sophisticated analysis of pre-trained parameter spaces. Starting from drawing insights from Neural Engrams (NE) in Biological Neural Networks (BNNs), we establish a connection between the low-rank property observed during PEFT’s parameter space shifting and neurobiological mechanisms. This observation leads to our proposed method, Synapse and Neuron (SAN), which decomposes and propagates the scaling component from anterior feature adjustment vectors towards posterior weight matrices. Our approach is theoretically grounded in Long-Term Potentiation/Depression (LTP/D) phenomena, which govern synapse development through neurotransmitter release modulation. Extensive experiments demonstrate its effectiveness: on vision tasks across VTAB, FGVC, and GIC (25 datasets) using ViT, Swin-T and ConvNeXt architectures, SAN outperforms FFT up to 8.7% and LoRA by 3.2%; on language tasks using Commonsense Reasoning (8 datasets) with LLaMA models (all generations), surpassing ChatGPT up to 8.5% and LoRA by 4.7%; on vision-language tasks using Visual Instruction Tuning (7 datasets) with LLaVA models, it exceeds FFT up to 2.4% and LoRA by 1.9%. Our code and W&B log will be released
Gaole Dai, Chun-Kai Fan, Zhi Zhang 0009, Yuan Zhang 0020, Yulu Gan, Qizhe Zhang, Cheng-Ching Tseng, Shanghang Zhang, Tiejun Huang 0001
ICML10
2025 Faster and Stronger: When ANN-SNN Conversion Meets Parallel Spiking Calculation
abstract
Spiking Neural Network (SNN), as a brain-inspired and energy-efficient network, is currently facing the pivotal challenge of exploring a suitable and efficient learning framework. The predominant training methodologies, namely Spatial-Temporal Back-propagation (STBP) and ANN-SNN Conversion, are encumbered by substantial training overhead or pronounced inference latency, which impedes the advancement of SNNs in scaling to larger networks and navigating intricate application domains. In this work, we propose a novel parallel conversion learning framework, which establishes a mathematical mapping relationship between each time-step of the parallel spiking neurons and the cumulative spike firing rate. We theoretically validate the lossless and sorting properties of the conversion process, as well as pointing out the optimal shifting distance for each step. Furthermore, by integrating the above framework with the distribution-aware error calibration technique, we can achieve efficient conversion towards more general activation functions or training-free circumstance. Extensive experiments have confirmed the significant performance advantages of our method for various conversion cases under ultra-low time latency. To our best knowledge, this is the first work which jointly utilizes parallel spiking calculation and ANN-SNN Conversion, providing a highly promising approach for SNN supervised training. Code is available at https://github.com/hzc1208/Parallel_Conversion.
Zecheng Hao, Zhaofei Yu, Tiejun Huang 0001
ICML6
2025 Differential Coding for Training-Free ANN-to-SNN Conversion
abstract
Spiking Neural Networks (SNNs) exhibit significant potential due to their low energy consumption. Converting Artificial Neural Networks (ANNs) to SNNs is an efficient way to achieve high-performance SNNs. However, many conversion methods are based on rate coding, which requires numerous spikes and longer time-steps compared to directly trained SNNs, leading to increased energy consumption and latency. This article introduces differential coding for ANN-to-SNN conversion, a novel coding scheme that reduces spike counts and energy consumption by transmitting changes in rate information rather than rates directly, and explores its application across various layers. Additionally, the threshold iteration method is proposed to optimize thresholds based on activation distribution when converting Rectified Linear Units (ReLUs) to spiking neurons. Experimental results on various Convolutional Neural Networks (CNNs) and Transformers demonstrate that the proposed differential coding significantly improves accuracy while reducing energy consumption, particularly when combined with the threshold iteration method, achieving state-of-the-art performance. The source codes of the proposed method are available at https://github.com/h-z-h-cell/ANN-to-SNN-DCGS.
Zihan Huang, Wei Fang 0006, Tong Bu, Zecheng Hao, Wenxuan Liu 0008, Yuanhong Tang, Zhaofei Yu, Tiejun Huang 0001
ICML9
2025 Neural Representational Consistency Emerges from Probabilistic Neural-Behavioral Representation Alignment
abstract
Individual brains exhibit striking structural and physiological heterogeneity, yet neural circuits can generate remarkably consistent functional properties across individuals, an apparent paradox in neuroscience. While recent studies have observed preserved neural representations in motor cortex through manual alignment across subjects, the zero-shot validation of such preservation and its generalization to more cortices remain unexplored. Here we present PNBA (Probabilistic Neural-Behavioral Representation Alignment), a new framework that leverages probabilistic modeling to address hierarchical variability across trials, sessions, and subjects, with generative constraints preventing representation degeneration. By establishing reliable cross-modal representational alignment, PNBA reveals robust preserved neural representations in monkey primary motor cortex (M1) and dorsal premotor cortex (PMd) through zero-shot validation. We further establish similar representational preservation in mouse primary visual cortex (V1), reflecting a general neural basis. These findings resolve the paradox of neural heterogeneity by establishing zero-shot preserved neural representations across cortices and species, enriching neural coding insights and enabling zero-shot behavior decoding.
Yu Zhu 0008, Chunfeng Song, Wanli Ouyang, Tiejun Huang 0001
ICML5
2025 SOTA: Spike-Navigated Optimal TrAnsport Saliency Region Detection in Composite-bias Videos
abstract
Existing saliency detection methods struggle in real-world scenarios due to motion blur and occlusions. In contrast, spike cameras, with their high temporal resolution, significantly enhance visual saliency maps. However, the composite noise inherent to spike camera imaging introduces discontinuities in saliency detection. Low-quality samples further distort model predictions, leading to saliency bias. To address these challenges, we propose Spike-navigated Optimal TrAnsport Saliency Region Detection (SOTA), a framework that leverages the strengths of spike cameras while mitigating biases in both spatial and temporal dimensions. Our method introduces Spike-based Micro-debias (SM) to capture subtle frame-to-frame variations and preserve critical details, even under minimal scene or lighting changes. Additionally, Spike-based Global-debias (SG) refines predictions by reducing inconsistencies across diverse conditions. Extensive experiments on real and synthetic datasets demonstrate that SOTA outperforms existing methods by eliminating composite noise bias. Our code and dataset will be released at https://github.com/lwxfight/sota.
Wenxuan Liu 0008, Xian Zhong, Zhaofei Yu, Tiejun Huang 0001
IJCAI6
2025 Orochi: Versatile Biomedical Image Processor
abstract
Deep learning has emerged as a pivotal tool for accelerating research in the life sciences, with the low-level processing of biomedical images (e.g., registration, fusion, restoration, super-resolution) being one of its most critical applications. Platforms such as ImageJ (Fiji) and napari have enabled the development of customized plugins for various models. However, these plugins are typically based on models that are limited to specific tasks and datasets, making them less practical for biologists. To address this challenge, we introduce **Orochi**, the first application-oriented, efficient, and versatile image processor designed to overcome these limitations. Orochi is pre-trained on patches/volumes extracted from the raw data of over 100 publicly available studies using our Random Multi-scale Sampling strategy. We further propose Task-related Joint-embedding Pre-Training (TJP), which employs biomedical task-related degradation for self-supervision rather than relying on Masked Image Modelling (MIM), which performs poorly in downstream tasks such as registration. To ensure computational efficiency, we leverage Mamba's linear computational complexity and construct Multi-head Hierarchy Mamba. Additionally, we provide a three-tier fine-tuning framework (Full, Normal, and Light) and demonstrate that Orochi achieves comparable or superior performance to current state-of-the-art specialist models, even with lightweight parameter-efficient options. We hope that our study contributes to the development of an all-in-one workflow, thereby relieving biologists from the overwhelming task of selecting among numerous models. Our pre-trained weights and code will be released.
Gaole Dai, Chenghao Zhou, Rongyu Zhang, Yuan Zhang 0020, Chengkai Hou, Tiejun Huang 0001, Jianxu Chen 0001, Shanghang Zhang
NeurIPS7
2025 RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
abstract
Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained VLMs, recent approaches are still not qualified to accurately understand the complex 3D scenes and dynamically reason about the instruction-indicated locations for interaction. To this end, we propose RoboRefer, a 3D-aware vision language model (VLM) that can first achieve precise spatial understanding by integrating a disentangled but dedicated depth encoder via supervised fine-tuning (SFT). Moreover, RoboRefer advances generalized multi-step spatial reasoning via reinforcement fine-tuning (RFT), with metric-sensitive process reward functions tailored for spatial referring tasks. To support SFT and RFT training, we introduce RefSpatial, a large-scale dataset of 20M QA pairs (2x prior), covering 31 spatial relations (vs. 15 prior) and supporting complex reasoning processes (up to 5 steps). In addition, we introduce RefSpatial-Bench, a challenging benchmark filling the gap in evaluating spatial referring with multi-step reasoning. Experiments show that SFT-trained RoboRefer achieves state-of-the-art spatial understanding, with an average success rate of 89.6%. RFT-trained RoboRefer further outperforms all other baselines by a large margin, even surpassing Gemini-2.5-Pro by 12.4% in average accuracy on RefSpatial-Bench. Notably, RoboRefer can be integrated with various control policies to execute long-horizon, dynamic tasks across diverse robots (e,g., UR5, G1 humanoid) in cluttered real-world scenes.
Enshen Zhou, Jingkun An, Cheng Chi 0001, Shanyu Rong, Pengwei Wang 0004, Zhongyuan Wang 0006, Tiejun Huang 0001, Lu Sheng, Shanghang Zhang
NeurIPS9
2025 High Dynamic Range Imaging with Time-Encoding Spike Camera
abstract
As a bio-inspired vision sensor, spike camera records light intensity by accumulating photons and firing a spike once a preset threshold is reached. For high-light regions, the accumulated photons may reach the threshold multiple times within a readout interval, while only one spike can be stored and read out, resulting in incorrect intensity representation and a limited dynamic range. Multi-level (ML) spike camera enhances the dynamic range by introducing a spike-firing counter (SFC) to count spikes within each readout interval for each pixel, and uses different spike symbols to represent the arrival of different amounts of photons. However, when the light intensity becomes even higher, each pixel requires an SFC with a higher bit depth, causing great cost to the manufacturing process. To address these issues, we propose time-encoding (TE) spike camera, which transforms the counting of spikes to recording of the time at which a specific number of spikes (i.e., an overflow) is reached. To encode time information with as few bits as possible, instead of directly utilising a timer, we leverage a periodic timing signal with a higher frequency than the readout signal. Then the recording of overflow moment can be transformed into recording the number of accumulated timing signal cycles until the overflow occurs. Additionally, we propose an image reconstruction scheme for TE spike camera, which leverages the multi-scale gradient features of spike data. This scheme includes a similarity-based pyramid alignment module to align spike streams across the temporal domain and a light intensity-based refinement module, which utilises the guidance of light intensity to fuse spatial features of the spike data. Experimental results demonstrate that TE spike camera effectively improves the dynamic range of spike camera.
Zhenkun Zhu 0001, Ruiqin Xiong, Jiyu Xie, Yuanlin Wang, Xinfeng Zhang 0001, Tiejun Huang 0001
NeurIPS6
2025 Spk2ImgMamba: Spiking Camera Image Reconstruction with Multi-Scale State Space Models
abstract
As a bio-inspired vision sensor, the spiking camera has showcased remarkable capability in high-speed imaging with a sampling rate of 40,000 Hz. Reconstructing clear images from continuous spike streams, which is obtained by each photosensor continuously detecting photons and firing them asynchronously, has garnered significant attention. Despite promising results, existing spike-to-image reconstruction methods face challenges in balancing global receptive fields and efficient computation due to the inherent limitations of their backbones. Recently, due to powerful long-range modeling and linear complexity, the state space model (SSM) has emerged as a competitive alternative to CNNs and Transformers. In this paper, we propose a lightweight spike-to-image reconstruction network that harnesses Mamba as the backbone. Our approach sequentially executes three core modules: temporal information integration, spatial feature enhancement, and progressive image reconstruction. The former accumulates cues across diverse temporal windows to explore both long-term and short-term contexts. Subsequently, to model global dependencies while heightening local detail perception, we develop a multi-scale SSM block characterized by multi-scale multi-direction scanning, which effectively boosts spatial feature representations. Finally, intensity images are decoded progressively from the enhanced light-intensity features. Extensive experiments on both synthetic and real-captured data demonstrate that our approach achieves state-of-the-art performance, with only 10% of the network parameters and nearly two orders of magnitude less computational effort. The code will be available at https://github.com/interstellarH/Spk2ImgMamba.
Jiaoyang Yin, Bin Fan 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
WACV4
2025 MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
abstract
Processing long contexts presents a significant challenge for large language models (LLMs). While recent advancements allow LLMs to handle much longer contexts than before (e.g., 32K or 128K tokens), it is computationally expensive and can still be insufficient for many applications. Retrieval-Augmented Generation (RAG) is considered a promising strategy to address this problem. However, conventional RAG methods face inherent limitations because of two underlying requirements: 1) explicitly stated queries, and 2) well-structured knowledge. These conditions, however, do not hold in general long-context processing tasks.
Hongjin Qian, Zheng Liu 0011, Peitian Zhang, Kelong Mao, Defu Lian, Zhicheng Dou, Tiejun Huang 0001
WWW7
2025 Masked Channel Modeling for Bootstrapping Visual Pre-training
Yang Liu 0357, Muzhi Zhu, Yue Cao 0001, Tiejun Huang 0001, Chunhua Shen
Int. J. Comput. Vis.5
2025 A Norm Regularization Training Strategy for Robust Image Quality Assessment Models
Yujia Liu 0005, Chenxi Yang 0004, Dingquan Li, Tingting Jiang 0001, Tiejun Huang 0001
Int. J. Comput. Vis.5
2025 Implementing feature binding through dendritic networks of a single neuron
abstract
A single neuron receives an extensive array of synaptic inputs through its dendrites, raising the fundamental question of how these inputs undergo integration and summation, culminating in the initiation of spikes in the soma. Experimental and computational investigations have revealed various modes of integration operations that include linear, superlinear, and sublinear summation. Interestingly, different types of neurons exhibit diverse patterns of dendritic integration depending on the spatial distribution of dendrites. The functional implications of these specific integration modalities remain largely unexplored. In this study, we employ the Purkinje cell (PC) as a model system to investigate these complex questions. Our findings reveal that PCs generally exhibit sublinear summation across their expansive dendrites. Both spatial and temporal input dynamically modulates the degree of sublinearity. Strong sublinearity necessitates the synaptic distribution in PCs to be globally scattered sensitive, whereas weak sublinearity facilitates the generation of complex firing patterns in PCs. Using dendritic branches characterized by strong sublinearity as computational units, we demonstrate that a neuron can successfully address the feature binding problem. Taken together, these results offer a systematic perspective on the functional role of dendritic sublinearity, inspiring a broader understanding of dendritic integration in various neuronal types.
Yuanhong Tang, Shanshan Jia 0001, Tiejun Huang 0001, Zhaofei Yu, Jian K. Liu
Neural Networks3
2025 Corrigendum to "Implementing feature binding through dendritic networks of a single neuron" [Neural Networks(2025) 107555]
Yuanhong Tang, Shanshan Jia 0001, Tiejun Huang 0001, Zhaofei Yu, Jian K. Liu
Neural Networks3
2025 Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic Segmentation
abstract
The divergence between labeled training data and unlabeled testing data is a significant challenge for recent deep learning models. Unsupervised domain adaptation (UDA) attempts to solve such problem. Recent works show that self-training is a powerful approach to UDA. However, existing methods have difficulty in balancing the scalability and performance. In this paper, we propose a hard-aware instance adaptive self-training framework for UDA on the task of semantic segmentation. To effectively improve the quality and diversity of pseudo-labels, we develop a novel pseudo-label generation strategy with an instance adaptive selector. We further enrich the hard class pseudo-labels with inter-image information through a skillfully designed hard-aware pseudo-label augmentation. Besides, we propose the region-adaptive regularization to smooth the pseudo-label region and sharpen the non-pseudo-label region. For the non-pseudo-label region, consistency constraint is also constructed to introduce stronger supervision signals during model optimization. Our method is so concise and efficient that it is easy to be generalized to other UDA methods. Experiments on GTA5 $\rightarrow$→ Cityscapes, SYNTHIA $\rightarrow$→ Cityscapes, and Cityscapes $\rightarrow$→ Oxford RobotCar demonstrate the superior performance of our approach compared with the state-of-the-art methods.
Chuang Zhu, Kebin Liu 0002, Wenqi Tang, Ke Mei, Jiaqi Zou, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Enhancing NR-IQA Model Robustness Through Simple Image Compression Techniques
abstract
No-Reference Image Quality Assessment (NR-IQA) plays a crucial role in various real-world applications by predicting image quality scores without the need for reference images. Despite the impressive performance of deep learning-based NR-IQA models, they remain vulnerable to adversarial attacks, which introduce imperceptible perturbations to input images, causing significant changes in predicted scores. In this study, we explore the use of simple JPEG compression techniques, as well as their combination with norm regularization training, to defend against these adversarial attacks. Our results demonstrate that image compression is an effective method to enhance model robustness, and it can further improve the robustness of NR-IQA models when combined with appropriate training strategies. Since excessive compression may reduce performance on clean images, it is essential to strike a balance. This work provides valuable insights into designing effective image compression methods for NR-IQA models.
Yujia Liu 0005, Chenxi Yang 0004, Zhaofei Yu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 High Dynamic Range Imaging for Dynamic Scenes Based on Multi-Level Spike Camera
abstract
Spike camera is a retina-inspired neuromorphic camera which can capture dynamic scenes of high-speed motion by firing a continuous stream of spikes at an extremely high temporal resolution. The limitation in the current design is that each spike only represents the arrival of a fixed amount of photons. It can not deal with strong light areas in which the amount of accumulated photons reaches the pre-specified threshold multiple times within a single readout interval. In this paper, we propose a new spike camera model of high-speed imaging for high dynamic range scenarios. In this scheme, each pixel accumulates the incoming photons persistently and generates a new type of spike stream in which each spike symbol may be associated with different levels, indicating the arrival of different amounts of photons since the last readout. This enables the camera to support dynamic scenes with wider dynamic range. To achieve this, we propose a two-level buffer mechanism, one for photon accumulation and one for spike-firing encoding. We use a register to hold the number of spike-firings which has not been read out yet. At each readout time, the major part in the counter is read out via a carefully designed exponential encoding and the counter is updated. Such encoding and readout strategy enables a very efficient expansion of the dynamic range using a small number of encoding bits. Furthermore, we propose an image reconstruction scheme for the proposed camera, utilizing both spike intervals and spike levels to recover the light intensity. We incorporate Mamba and propose a temporal-spatial selective scan mechanism to extract temporal-spatial correlation within spike streams. We employ a pyramid adaptive filtering and alignment module to achieve coarse-to-fine feature alignment. Experimental results show that the proposed scheme can achieve better imaging quality and outperform the existing spike camera in high dynamic range scenarios.
Zhenkun Zhu 0001, Ruiqin Xiong, Jing Zhao 0011, Rui Zhao 0010, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Color Spike Camera Reconstruction via Long Short-Term Temporal Aggregation of Spike Signals
abstract
With the prevalence of emerging computer vision applications, the demand for capturing dynamic scenes with high-speed motion has increased. A kind of neuromorphic sensor called spike camera shows great potential in this aspect since it generates a stream of binary spikes to describe the dynamic light intensity with a very high temporal resolution. Color spike camera (CSC) was recently invented to capture the color information of dynamic scenes via a color filter array (CFA) on the sensor. This paper proposes a long short-term temporal aggregation strategy of spike signals. First, we utilize short-term temporal correlation to adaptively extract temporal features of each time point. Then we align the features and aggregate them to exploit long-term temporal correlation, suppressing undesired motion blur. To implement the strategy, we design a CSC reconstruction network. Based on adaptive short-term temporal aggregation, we propose a spike representation module to extract temporal features of each color channel, leveraging multiple temporal scales. Considering the long-term temporal correlation, we develop an alignment module to align the temporal features. In particular, we perform motion alignment of red and blue channels with the guidance of the higher-sampling-rate green channel, leveraging motion consistency among color channels. Besides, we propose a module to aggregate the aligned temporal features for the restored color image, which exploits color channel correlation. We have also developed a CSC simulator for data generation. Experimental results demonstrate that our method can restore color images with fine texture details, achieving state-of-the-art CSC reconstruction performance.
Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Xiaopeng Fan 0001, Xinfeng Zhang 0001, Tiejun Huang 0001
IEEE Trans. Image Process.6
2025 Super-Resolving Dynamic Scenes With Spike Camera via Multi-Frame Sequential Alignment With Motion Propagation
abstract
Spike camera is a neuromorphic sensor that can capture high-speed dynamic scenes by firing a continuous stream of binary spikes with extremely high temporal resolution, essentially forming a dense sampling in the temporal dimension. Due to the relative motion between camera and scene, each pixel is actually sampling at a large number of different spatial positions on the object in a short period. Converting this dense sampling from temporal dimension to spatial domain, high resolution images can be reconstructed from the spike stream. However, spike fluctuations and large motion in high-speed scenes pose great challenges for this task, especially for intensity information extraction and temporal alignment. In this paper, we propose a spike camera super resolution network to address these issues. Considering the local temporal correlation of spike stream and correlation consistency within a local region, we introduce a representation module that performs region-adaptive temporal filtering on spikes to mitigate fluctuations and extract stable intensity information from binary data. Additionally, we develop a module for multi-frame feature alignment, leveraging the long-term temporal information of spike stream. To handle large motions, we propagate the motion information from neighboring moment to current feature alignment module, which provides a prior that helps to narrow the search range for current motion offset, improving the accuracy of temporal alignment. Experimental results demonstrate that the proposed network achieves state-of-the-art performance on synthetic and real-captured spike data.
Yuanlin Wang, Ruiqin Xiong, Jian Zhang 0018, Xinfeng Zhang 0001, Tiejun Huang 0001
IEEE Trans. Image Process.6
2025 Source-Free Semantic Regularization Learning for Semi-Supervised Domain Adaptation
abstract
Semi-supervised domain adaptation (SSDA) has been extensively researched due to its ability to improve classification performance and generalization ability of models by using a small amount of labeled data on the target domain. However, existing methods cannot effectively adapt to the target domain due to difficulty in fully learning rich and complex target semantic information and relationships. In this paper, we propose a novel SSDA learning framework called semantic regularization learning (SERL), which captures the target semantic information from multiple perspectives of regularization learning to achieve adaptive fine-tuning of the source pre-trained model on the target domain. SERL includes three robust semantic regularization techniques. Firstly, semantic probability contrastive regularization (SPCR) helps the model learn more discriminative feature representations from a probabilistic perspective, using semantic information on the target domain to understand the similarities and differences between samples. Additionally, adaptive weights in SPCR can help the model learn the semantic distribution correctly through the probabilities of different samples. To further comprehensively understand the target semantic distribution, we introduce hard-sample mixup regularization (HMR), which uses easy samples as guidance to mine the latent target knowledge contained in hard samples, thereby learning more complete and complex target semantic knowledge. Finally, target prediction regularization (TPR) regularizes the target predictions of the model by maximizing the correlation between the current prediction and the past learned objective, thereby mitigating the misleading of semantic information caused by erroneous pseudo-labels. Extensive experiments on three benchmark datasets demonstrate that our SERL method achieves state-of-the-art performance.
Chuang Zhu, Ruiying Ren, Tiejun Huang 0001
IEEE Trans. Multim.5
2025 Fully Spiking Actor Network With Intralayer Connections for Reinforcement Learning
abstract
With the help of special neuromorphic hardware, spiking neural networks (SNNs) are expected to realize artificial intelligence (AI) with less energy consumption. It provides a promising energy-efficient way for realistic control tasks by combining SNNs with deep reinforcement learning (DRL). In this article, we focus on the task where the agent needs to learn multidimensional deterministic policies to control, which is very common in real scenarios. Recently, the surrogate gradient method has been utilized for training multilayer SNNs, which allows SNNs to achieve comparable performance with the corresponding deep networks in this task. Most existing spike-based reinforcement learning (RL) methods take the firing rate as the output of SNNs, and convert it to represent continuous action space (i.e., the deterministic policy) through a fully connected (FC) layer. However, the decimal characteristic of the firing rate brings the floating-point matrix operations to the FC layer, making the whole SNN unable to deploy on the neuromorphic hardware directly. To develop a fully spiking actor network (SAN) without any floating-point matrix operations, we draw inspiration from the nonspiking interneurons found in insects and employ the membrane voltage of the nonspiking neurons to represent the action. Before the nonspiking neurons, multiple population neurons are introduced to decode different dimensions of actions. Since each population is used to decode a dimension of action, we argue that the neurons in each population should be connected in time domain and space domain. Hence, the intralayer connections are used in output populations to enhance the representation capacity. This mechanism exists extensively in animals and has been demonstrated effectively. Finally, we propose a fully SAN with intralayer connections (ILC-SAN). Extensive experimental results demonstrate that the proposed method outperforms the state-of-the-art performance on continuous control tasks from OpenAI gym. Moreover, we estimate the theoretical energy consumption when deploying ILC-SAN on neuromorphic chips to illustrate its high energy efficiency.
Peixi Peng, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Assisting Training of Deep Spiking Neural Networks With Parameter Initialization
abstract
Spiking neural networks (SNNs) exhibit significant advantages in terms of information encoding, computational capabilities, and power usage. We regard initializing weight distribution as a key problem for effective SNN training. When backpropagation (BP) through time is used in the initial training phase, it has a significant impact on gradient generation. We first derive an asymptotic formula for the response curve of spiking neurons, which approximates the real neuron response distribution. To avoid gradient vanishing, we then provide an initialization technique based on the slant asymptote. Finally, validations on classification tasks on the MNIST and CIFAR10 datasets demonstrate that our strategy can significantly speed up training and improve the model accuracy compared with other initialization methods. Further testing on various neuron configurations and training hyperparameters demonstrates comparable versatility and superiority to other methods. Based on the analyses, some recommendations for SNN training are made.
Jianhao Ding, Jiyuan Zhang 0005, Tiejun Huang 0001, Jian K. Liu, Zhaofei Yu
IEEE Trans. Neural Networks Learn. Syst.3
2025 Robust Decoding of Rich Dynamical Visual Scenes With Retinal Spikes
abstract
Sensory information transmitted to the brain activates neurons to create a series of coping behaviors. Understanding the mechanisms of neural computation and reverse engineering the brain to build intelligent machines requires establishing a robust relationship between stimuli and neural responses. Neural decoding aims to reconstruct the original stimuli that trigger neural responses. With the recent upsurge of artificial intelligence, neural decoding provides an insightful perspective for designing novel algorithms of brain-machine interface. For humans, vision is the dominant contributor to the interaction between the external environment and the brain. In this study, utilizing the retinal neural spike data collected over multi trials with visual stimuli of two movies with different levels of scene complexity, we used a neural network decoder to quantify the decoded visual stimuli with six different metrics for image quality assessment establishing comprehensive inspection of decoding. With the detailed and systematical study of the effect and single and multiple trials of data, different noise in spikes, and blurred images, our results provide an in-depth investigation of decoding dynamical visual scenes using retinal spikes. These results provide insights into the neural coding of visual scenes and services as a guideline for designing next-generation decoding algorithms of neuroprosthesis and other devices of brain-machine interface.
Zhaofei Yu, Tong Bu, Yijun Zhang 0003, Shanshan Jia 0001, Tiejun Huang 0001, Jian K. Liu
IEEE Trans. Neural Networks Learn. Syst.5
2024 Joint Demosaicing and Denoising for Spike Camera
abstract
As a neuromorphic camera with high temporal resolution, spike camera can capture dynamic scenes with high-speed motion. Recently, spike camera with a color filter array (CFA) has been developed for color imaging. There are some methods for spike camera demosaicing to reconstruct color images from Bayer-pattern spike streams. However, the demosaicing results are bothered by severe noise in spike streams, to which previous works pay less attention. In this paper, we propose an iterative joint demosaicing and denoising network (SJDD-Net) for spike cameras based on the observation model. Firstly, we design a color spike representation (CSR) to learn latent representation from Bayer-pattern spike streams. In CSR, we propose an offset-sharing deformable convolution module to align temporal features of color channels. Then we develop a spike noise estimator (SNE) to obtain features of the noise distribution. Finally, a color correlation prior (CCP) module is proposed to utilize the color correlation for better details. For training and evaluation, we designed a spike camera simulator to generate Bayer-pattern spike streams with synthesized noise. Besides, we captured some Bayer-pattern spike streams, building the first real-world captured dataset to our knowledge. Experimental results show that our method can restore clean images from Bayer-pattern spike streams. The source codes and dataset are available at https://github.com/csycdong/SJDD-Net.
Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
AAAI7
2024 Optical Flow for Spike Camera with Hierarchical Spatial-Temporal Spike Fusion
abstract
As an emerging neuromorphic camera with an asynchronous working mechanism, spike camera shows good potential for high-speed vision tasks. Each pixel in spike camera accumulates photons persistently and fires a spike whenever the accumulation exceeds a threshold. Such high-frequency fine-granularity photon recording facilitates the analysis and recovery of dynamic scenes with high-speed motion. This paper considers the optical flow estimation problem for spike cameras. Due to the Poisson nature of incoming photons, the occurrence of spikes is random and fluctuating, making conventional image matching inefficient. We propose a Hierarchical Spatial-Temporal (HiST) fusion module for spike representation to pursue reliable feature matching and develop a robust optical flow network, dubbed as HiST-SFlow. The HiST extracts features at multiple moments and hierarchically fuses the spatial-temporal information. We also propose an intra-moment filtering module to further extract the feature and suppress the influence of randomness in spikes. A scene loss is proposed to ensure that this hierarchical representation recovers the essential visual information in the scene. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared with the existing methods. The source codes are available at https://github.com/ruizhao26/HiST-SFlow.
Rui Zhao 0010, Ruiqin Xiong, Jian Zhang 0018, Xinfeng Zhang 0001, Zhaofei Yu, Tiejun Huang 0001
AAAI6
2024 Enhancing the Robustness of Spiking Neural Networks with Stochastic Gating Mechanisms
abstract
Spiking neural networks (SNNs) exploit neural spikes to provide solutions for low-power intelligent applications on neuromorphic hardware. Although SNNs have high computational efficiency due to spiking communication, they still lack resistance to adversarial attacks and noise perturbations. In the brain, neuronal responses generally possess stochasticity induced by ion channels and synapses, while the role of stochasticity in computing tasks is poorly understood. Inspired by this, we elaborate a stochastic gating spiking neural model for layer-by-layer spike communication, introducing stochasticity to SNNs. Through theoretical analysis, our gating model can be viewed as a regularizer that prevents error amplification under attacks. Meanwhile, our work can explain the robustness of Poisson coding. Experimental results prove that our method can be used alone or with existing robust enhancement algorithms to improve SNN robustness and reduce SNN energy consumption. We hope our work will shed new light on the role of stochasticity in the computation of SNNs. Our code is available at https://github.com/DingJianhao/StoG-meets-SNN/.
Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001, Jian K. Liu
AAAI3
2024 Evidential Uncertainty-Guided Mitochondria Segmentation for 3D EM Images
abstract
Recent advances in deep learning have greatly improved the segmentation of mitochondria from Electron Microscopy (EM) images. However, suffering from variations in mitochondrial morphology, imaging conditions, and image noise, existing methods still exhibit high uncertainty in their predictions. Moreover, in view of our findings, predictions with high levels of uncertainty are often accompanied by inaccuracies such as ambiguous boundaries and amount of false positive segments. To deal with the above problems, we propose a novel approach for mitochondria segmentation in 3D EM images that leverages evidential uncertainty estimation, which for the first time integrates evidential uncertainty to enhance the performance of segmentation. To be more specific, our proposed method not only provides accurate segmentation results, but also estimates associated uncertainty. Then, the estimated uncertainty is used to help improve the segmentation performance by an uncertainty rectification module, which leverages uncertainty maps and multi-scale information to refine the segmentation. Extensive experiments conducted on four challenging benchmarks demonstrate the superiority of our proposed method over existing approaches.
Ruohua Shi, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001
AAAI3
2024 Transient Glimpses: Unveiling Occluded Backgrounds through the Spike Camera
abstract
The de-occlusion problem, involving extracting clear background images by removing foreground occlusions, holds significant practical importance but poses considerable challenges. Most current research predominantly focuses on generating discrete images from calibrated camera arrays, but this approach often struggles with dense occlusions and fast motions due to limited perspectives and motion blur. To overcome these limitations, an effective solution requires the integration of multi-view visual information. The spike camera, as an innovative neuromorphic sensor, shows promise with its ultra-high temporal resolution and dynamic range. In this study, we propose a novel approach that utilizes a single spike camera for continuous multi-view imaging to address occlusion removal. By rapidly moving the spike camera, we capture a dense stream of spikes from occluded scenes. Our model, SpkOccNet, processes these spikes by integrating multi-view spatial-temporal information via long-short-window feature extractor (LSW) and employs a novel cross-view mutual attention-based module (CVA) for effective fusion and refinement. Additionally, to facilitate research in occlusion removal, we introduce the S-OCC dataset, which consists of real-world spike-based data. Experimental results demonstrate the efficiency and generalization capabilities of our model in effectively removing dense occlusions across diverse scenes. Public project page: https://github.com/Leozhangjiyuan/SpikeDeOcclusion.
Jiyuan Zhang 0005, Shiyan Chen, Yajing Zheng, Zhaofei Yu, Tiejun Huang 0001
AAAI5
2024 Recognizing Ultra-High-Speed Moving Objects with Bio-Inspired Spike Camera
abstract
Bio-inspired spike camera mimics the sampling principle of primate fovea. It presents high temporal resolution and dynamic range, showing great promise in fast-moving object recognition. However, the physical limit of CMOS technology in spike cameras still hinders their capability of recognizing ultra-high-speed moving objects, e.g., extremely fast motions cause blur during the imaging process of spike cameras. This paper presents the first theoretical analysis for the causes of spiking motion blur and proposes a robust representation that addresses this issue through temporal-spatial context learning. The proposed method leverages multi-span feature aggregation to capture temporal cues and employs residual deformable convolution to model spatial correlation among neighbouring pixels. Additionally, this paper contributes an original real-captured spiking recognition dataset consisting of 12,000 ultra-high-speed (equivalent speed > 500 km/h) moving objects. Experimental results show that the proposed method achieves 73.2% accuracy in recognizing 10 classes of ultra-high-speed moving objects, outperforming all existing spike-based recognition methods. Resources will be available at https://github.com/Evin-X/UHSR.
Junwei Zhao 0003, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001
AAAI4
2024 E/I Balanced Adaptive Sequential Neural Posterior Estimation for Inferring the Connection Weights in Mouse V1 Model
abstract
Effectively utilizing biological firing rate data to estimate the numerous connection weights in the mouse primary visual cortex (V1) model from the Allen Institute is a challenging task. The existing iterative grid-search algorithm cannot enable the mouse V1 model to better fit the biological firing rate data. To tackle this issue, we propose an excitation-inhibition balanced adaptive sequential neural posterior estimation (E/I balanced ASNPE) approach to accurately infer the connection weights of the mouse V1 model, allowing the neurons’ firing rates to converge to the given biological data. This method fully leverages the structural information of the mouse V1 model, reducing the dimensionality of the weight parameters to be optimized. Initially, sampling is performed in the prior distribution based on the proposed non-dominated sorting adaptive genetic algorithm (NSAGA). This algorithm optimizes the sorting, crossover and mutation processes based on the fitness scores of the current samples and updates the proposal distribution based on these samples, increasing the likelihood of identifying high posterior probability regions in the prior distribution. To avoid bad simulations, we also explore the E/I balance in each layer of the mouse V1 model, adding biological constraints during weight inference with Automatic Posterior Transformation (APT). Experimental results confirm that the proposed E/I balanced ASNPE method significantly outperforms the baseline in all five firing rate fitness scores in the mouse V1 model. This study is pioneering in applying non-dominated sorting genetic algorithms combined with sequential neural posterior estimation to optimize connection weights in large-scale complex biological models.
Luntian Mou, Peize Li, Lei Ma 0008, Tiejun Huang 0001
BIBM5
2024 Super-Resolution Reconstruction from Bayer-Pattern Spike Streams
abstract
Spike camera is a neuromorphic vision sensor that can capture highly dynamic scenes by generating a continuous stream of binary spikes to represent the arrival of photons at very high temporal resolution. Equipped with Bayer color filter array (CFA), color spike camera (CSC) has been invented to capture color information. Although spike camera has already demonstrated great potential for high-speed imaging, its spatial resolution is limited compared with conventional digital cameras. This paper proposes a Color Spike Camera Super-Resolution (CSCSR) network to super-resolve higher-resolution color images from spike camera streams with Bayer CFA. To be specific, we first propose a representation for Bayer-pattern spike streams, exploring local temporal information with global perception to represent the binary data. Then we exploit the CFA layout and sub-pixel level motion to collect temporal pixels for the spatial super-resolution of each color channel. In particular, a residual-based module for feature refinement is developed to reduce the impact of motion estimation errors. Considering color correlation, we jointly utilize the multi-stage temporal-pixel features of color channels to reconstruct the high-resolution color image. Experimental results demonstrate that the proposed scheme can reconstruct satisfactory color images with both high temporal and spatial resolution from low-resolution Bayerpattern spike streams. The source codes are available at https://github.com/csycdong/CSCSR.
Yanchen Dong 0001, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
CVPR7
2024 Boosting Spike Camera Image Reconstruction from a Perspective of Dealing with Spike Fluctuations
abstract
As a bio-inspired vision sensor with ultra-high speed, spike cameras exhibit great potential in recording dynamic scenes with high-speed motion or drastic light changes. Different from traditional cameras, each pixel in spike cam-eras records the arrival of photons continuously by firing binary spikes at an ultra-fine temporal granularity. In this process, multiple factors impact the imaging, including the photons' Poisson arrival, thermal noises from circuits, and quantization effects in spike readout. These factors intro-duce fluctuations to spikes, making the recorded spike in-tervals unstable and unable to reflect accurate light intensi-ties. In this paper, we present an approach to deal with spike fluctuations and boost spike camera image reconstruction. We first analyze the quantization effects and reveal the unbi-ased estimation attribute of the reciprocal of differential of spike firing time (DSFT). Based on this, we propose a spike representation module to use DSFT with multiple orders for fluctuation suppression, where DSFT with higher or-ders indicates spike integration duration between multiple spikes. We also propose a module for inter-moment feature alignment at multiple granularities. The coarser alignment is based on patch-level cross-attention with a local search strategy, and the finer alignment is based on deformable convolution at the pixel level. Experimental results demon-strate the effectiveness of our method on both synthetic and real-captured data. The source code and dataset are avail-able at https://github.com/ruizhao26/BSF.
Rui Zhao 0010, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Zhaofei Yu, Tiejun Huang 0001
CVPR7
2024 Towards HDR and HFR Video from Rolling-Mixed-Bit Spikings
abstract
The spiking cameras offer the benefits of high dynamic range (HDR), high temporal resolution, and low data redundancy. However, reconstructing HDR videos in high-speed conditions using single-bit spikings presents challenges due to the limited bit depth. Increasing the bit depth of the spikings is advantageous for boosting HDR performance, but the readout efficiency will be decreased, which is unfavorable for achieving a high frame rate (HFR) video. To address these challenges, we propose a readout mechanism to obtain rolling-mixed-bit (RMB) spikings, which involves inter-leaving multi-bit spikings within the single-bit spikings in a rolling manner, thereby combining the characteristics of high bit depth and efficient readout. Furthermore, we introduce RMB-Net for reconstructing HDR and HFR videos. RMB-Net comprises a cross-bit attention block for fusing mixed-bit spikings and a cross-time attention block for achieving temporal fusion. Extensive experiments conducted on synthetic and real-synthetic data demonstrate the superiority of our method. For instance, pure 3 -bit spikings result in 3 times of data volume, whereas our method achieves comparable performance with less than 2% increase in data volume.
Yakun Chang, Yeliduosi Xiaokaiti, Yujia Liu 0005, Bin Fan 0002, Zhaojun Huang, Tiejun Huang 0001, Boxin Shi
CVPR6
2024 Exploring Efficient Asymmetric Blind-Spots for Self-Supervised Denoising in Real-World Scenarios
abstract
Self-supervised denoising has attracted widespread at-tention due to its ability to train without clean images. How-ever, noise in real-world scenarios is often spatially cor-related, which causes many self-supervised algorithms that assume pixel-wise independent noise to perform poorly. Re-cent works have attempted to break noise correlation with downsampling or neighborhood masking. However, denoising on downsampled subgraphs can lead to aliasing effects and loss of details due to a lower sampling rate. Further-more, the neighborhood masking methods either come with high computational complexity or do not consider local spatial preservation during inference. Through the analy-sis of existing methods, we point out that the key to obtaining high-quality and texture-rich results in real-world self-supervised denoising tasks is to train at the original input resolution structure and use asymmetric operations during training and inference. Based on this, we propose Asymmet-ric Tunable Blind-Spot Network (AT-BSN), where the blind-spot size can be freely adjusted, thus better balancing noise correlation suppression and image local spatial destruction during training and inference. In addition, we regard the pre-trained AT-BSN as a meta-teacher network capable of generating various teacher networks by sampling different blind-spots. We propose a blind-spot based multi-teacher distillation strategy to distill a lightweight network, signif-icantly improving performance. Experimental results on multiple datasets prove that our method achieves state-of-the-art, and is superior to other self-supervised algorithms in terms of computational overhead and visual effects.
Shiyan Chen, Jiyuan Zhang 0005, Zhaofei Yu, Tiejun Huang 0001
CVPR4
2024 Intensity-Robust Autofocus for Spike Camera
abstract
Spike cameras, a novel neuromorphic visual sensor, can capture full-time spatial information through spike stream, offering ultra-high temporal resolution and an extensive dy-namic range. Autofocus control (AC) plays a pivotal role in a camera to efficiently capture information in challenging real-world scenarios. Nevertheless, due to disparities in data modality and information characteristics compared to frame stream and event stream, the current lack of effi-cient AC methods has made it challenging for spike cam-eras to adapt to intricate real-world conditions. To ad-dress this challenge, we introduce a spike-based autofo-cus framework that includes a spike-specific focus measure called spike dispersion (SD), which effectively mitigates the influence of variations in scene light intensity during the focusing process by leveraging the spike camera's ability to record full-time spatial light intensity. Additionally, the framework integrates a fast search strategy called spike-based goldenfast search (SGFS), allowing rapidfocal positioning without the need for a complete focus range traver-sal. To validate the performance of our method, we have collected a spike-based autofocus dataset (SAD) containing synthetic data and real-world data under varying scene brightness and motion scenarios. Experimental results on these datasets demonstrate that our method offers state-of-the-art accuracy and efficiency. Furthermore, experiments with data captured under varying scene brightness levels illustrate the robustness of our method to changes in light intensity during the focusing process.
Changqing Su, Zhiyuan Ye, Yongsheng Xiao, Zhen Cheng 0005, Zhaofei Yu, Tiejun Huang 0001
CVPR8
2024 Generative Multimodal Models are In-Context Learners
abstract
The human ability to easily solve multimodal tasks in context (i.e., with only a few demonstrations or simple instructions), is what current multimodal systems have largely struggled to imitate. In this work, we demonstrate that the task-agnostic in-context learning capabilities of large multimodal models can be significantly enhanced by effective scaling-up. We introduce Emu2, a generative multimodal model with 37 billion parameters, trained on large-scale multimodal sequences with a unified autoregressive objective. Emu2 exhibits strong multimodal in-context learning abilities, even emerging to solve tasks that require on-the-fly reasoning, such as visual prompting and object-grounded generation. The model sets a new record on multiple multimodal understanding tasks in few-shot settings. When instruction-tuned to follow specific instructions, Emu2 further achieves new state-of-the-art on challenging tasks such as question answering benchmarks for large multimodal models and open-ended subject-driven generation. These achievements demonstrate that Emu2 can serve as a base model and general-purpose interface for a wide range of multimodal tasks. Code and models are publicly available to facilitate future research.
Yufeng Cui, Qiying Yu, Yueze Wang, Yongming Rao, Tiejun Huang 0001
CVPR9
2024 Spike-guided Motion Deblurring with Unknown Modal Spatiotemporal Alignment
abstract
The traditional frame-based cameras that rely on exposure windows for imaging experience motion blur in high-speed scenarios. Frame-based deblurring methods lack reliable motion cues to restore sharp images under extreme blur conditions. The spike camera is a novel neuromorphic visual sensor that outputs spike streams with ultra-high temporal resolution. It can supplement the temporal information lost in traditional cameras and guide motion deblurring. However, in real-world scenarios, aligning discrete RGB images and continuous spike streams along both temporal and spatial axes is challenging due to the complexity of calibrating their coordinates, device displacements in vibrations, and time deviations. Misalignment of pixels leads to severe degradation of deblurring. We introduce the first framework for spike-guided motion deblurring without knowing the spatiotemporal alignment between spikes and images. To address the problem, we first propose a novel three-stage network containing a basic deblurring net, a carefully designed bi-directional deformable aligning module, and a flow-based multi-scale fusion net. Experimental results demonstrate that our approach can effectively guide the image deblurring with unknown alignment, surpassing the performance of other methods. Public project page: https://github.com/Leozhangjiyuan/UaSDN.
Jiyuan Zhang 0005, Shiyan Chen, Yajing Zheng, Zhaofei Yu, Tiejun Huang 0001
CVPR5
2024 Learning to Robustly Reconstruct Dynamic Scenes from Low-Light Spike Streams
Liwen Hu 0002, Ziluo Ding, Mianzhi Liu, Lei Ma 0008, Tiejun Huang 0001
ECCV (17)5
2024 Region-Native Visual Tokenization
Meng Wang 0001, Yuyao Huang 0002, Henghui Ding, Tiejun Huang 0001, Yao Zhao 0001, Yunchao Wei, Shuicheng Yan
ECCV (74)5
2024 Reconstruct Dynamic Scene for Spike Camera Based on 3D Space Time Similarity
abstract
Spike camera is a neuromorphic camera that recurrently accumulates photons and fires spikes to record the incident light intensity at very high temporal resolution, making it particularly suitable for recording high dynamic scenes. This paper addresses the problem of image reconstruction for spike camera. Due to the Poisson effect of photon arrival and the quantization effect of spike readout, the spike interval calculated from a single spike cycle cannot reflect the light intensity accurately. Firstly, this paper analyzes the error of spike interval estimation under static light intensity. Then, this paper focuses on the temporal correlation of continuous spikes under dynamic light intensity. Specifically, it considers 3D space time similarity to weighted average multiple continuous spike intervals for the intensity estimation at a certain pixel. Experimental results demonstrate that the proposed method achieves better performance in both objective and subjective aspects compared with previous reconstruction methods.
Yuanlin Wang, Ruiqin Xiong, Jing Zhao 0011, Tiejun Huang 0001
ICIP4
2024 Threaten Spiking Neural Networks through Combining Rate and Temporal Information
abstract
Spiking Neural Networks (SNNs) have received widespread attention in academic communities due to their superior spatio-temporal processing capabilities and energy-efficient characteristics. With further in-depth application in various fields, the vulnerability of SNNs under adversarial attack has become a focus of concern. In this paper, we draw inspiration from two mainstream learning algorithms of SNNs and observe that SNN models reserve both rate and temporal information. To better understand the capabilities of these two types of information, we conduct a quantitative analysis separately for each. In addition, we note that the retention degree of temporal information is related to the parameters and input settings of spiking neurons. Building on these insights, we propose a hybrid adversarial attack based on rate and temporal information (HART), which allows for dynamic adjustment of the rate and temporal attributes. Experimental results demonstrate that compared to previous works, HART attack can achieve significant superiority under different attack scenarios, data types, network architecture, time-steps, and model hyper-parameters. These findings call for further exploration into how both types of information can be effectively utilized to enhance the reliability of SNNs. Code is available at [https://github.com/hzc1208/HART_Attack](https://github.com/hzc1208/HART_Attack).
Zecheng Hao, Tong Bu, Xinyu Shi 0004, Zihan Huang, Zhaofei Yu, Tiejun Huang 0001
ICLR6
2024 A Progressive Training Framework for Spiking Neural Networks with Learnable Multi-hierarchical Model
abstract
Spiking Neural Networks (SNNs) have garnered considerable attention due to their energy efficiency and unique biological characteristics. However, the widely adopted Leaky Integrate-and-Fire (LIF) model, as the mainstream neuron model in current SNN research, has been revealed to exhibit significant deficiencies in deep-layer gradient calculation and capturing global information on the time dimension. In this paper, we propose the Learnable Multi-hierarchical (LM-H) model to address these issues by dynamically regulating its membrane-related factors. We point out that the LM-H model fully encompasses the information representation range of the LIF model while offering the flexibility to adjust the extraction ratio between historical and current information. Additionally, we theoretically demonstrate the effectiveness of the LM-H model and the functionality of its internal parameters, and propose a progressive training algorithm tailored specifically for the LM-H model. Furthermore, we devise an efficient training framework for our novel advanced model, encompassing hybrid training and time-slicing online training. Through extensive experiments on various datasets, we validate the remarkable superiority of our model and training algorithm compared to previous state-of-the-art approaches. Code is available at [https://github.com/hzc1208/STBP_LMH](https://github.com/hzc1208/STBP_LMH).
Zecheng Hao, Xinyu Shi 0004, Zihan Huang, Tong Bu, Zhaofei Yu, Tiejun Huang 0001
ICLR6
2024 Emu: Generative Pretraining in Multimodality
abstract
We present Emu, a multimodal foundation model that seamlessly generates images and text in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscriminately (e.g., interleaved image, text and video) through a one-model-for-all autoregressive training process. First, visual signals are encoded into embeddings, and together with text tokens form an interleaved input sequence. Emu is end-to-end trained with a unified objective of classifying the next text token or regressing the next visual embedding in the multimodal sequence. This versatile multimodality empowers the leverage of diverse pretraining data sources at scale, such as videos with interleaved frames and text, webpages with interleaved images and text, as well as web-scale image-text pairs and video-text pairs. Emu can serve as a generalist multimodal interface for both image-to-text and text-to-image tasks, supporting in-context image and text generation. Across a broad range of zero-shot/few-shot tasks including image captioning, visual question answering, video question answering and text-to-image generation, Emu demonstrates superb performance compared to state-of-the-art large multimodal models. Extended capabilities such as multimodal assistants via instruction tuning are also demonstrated with impressive performance.
Qiying Yu, Yufeng Cui, Yueze Wang, Hongcheng Gao, Tiejun Huang 0001
ICLR9
2024 Uni3D: Exploring Unified 3D Representation at Scale
abstract
Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively unexplored. In this work, we present Uni3D, a 3D foundation model to explore the unified 3D representation at scale. Uni3D uses a 2D initialized ViT end-to-end pretrained to align the 3D point cloud features with the image-text aligned features. Via the simple architecture and pretext task, Uni3D can leverage abundant 2D pretrained models as initialization and image-text aligned models as the target, unlocking the great potential of 2D model zoos and scaling-up strategies to the 3D world. We efficiently scale up Uni3D to one billion parameters, and set new records on a broad range of 3D tasks, such as zero-shot classification, few-shot classification, open-world understanding and zero-shot part segmentation. We show that the strong Uni3D representation also enables applications such as 3D painting and retrieval in the wild. We believe that Uni3D provides a new direction for exploring both scaling up and efficiency of the representation in 3D domain.
Junsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu, Tiejun Huang 0001
ICLR5
2024 Online Stabilization of Spiking Neural Networks
abstract
Spiking neural networks (SNNs), attributed to the binary, event-driven nature of spikes, possess heightened biological plausibility and enhanced energy efficiency on neuromorphic hardware compared to analog neural networks (ANNs). Mainstream SNN training schemes apply backpropagation-through-time (BPTT) with surrogate gradients to replace the non-differentiable spike emitting process during backpropagation. While achieving competitive performance, the requirement for storing intermediate information at all time-steps incurs higher memory consumption and fails to fulfill the online property crucial to biological brains. Our work focuses on online training techniques, aiming for memory efficiency while preserving biological plausibility. The limitation of not having access to future information in early time steps in online training has constrained previous efforts to incorporate advantageous modules such as batch normalization. To address this problem, we propose Online Spiking Renormalization (OSR) to ensure consistent parameters between testing and training, and Online Threshold Stabilizer (OTS) to stabilize neuron firing rates across time steps. Furthermore, we design a novel online approach to compute the sample mean and variance over time for OSR. Experiments conducted on various datasets demonstrate the proposed method's superior performance among SNN online training algorithms. Our code is available at https://github.com/zhuyaoyu/SNN-online-normalization.
Yaoyu Zhu, Jianhao Ding, Tiejun Huang 0001, Zhaofei Yu
ICLR3
2024 Spike-NeRF: Neural Radiance Field Based On Spike Camera
abstract
As a neuromorphic sensor with high temporal resolution, spike cameras offer notable advantages over traditional cameras in high-speed vision applications such as high-speed optical estimation, depth estimation, and object tracking. Inspired by the success of the spike camera, we proposed Spike-NeRF, the first Neural Radiance Field derived from spike data, to achieve 3D reconstruction and novel viewpoint synthesis for high-speed scenes. Instead of the multi-view images at the same as time of NeRF, the inputs of Spike-NeRF are continuous spike streams captured by a moving spike camera in a very short time. To reconstruct a correct and stable 3D scene from high-frequency but unstable spike data, we devised spike masks along with a distinctive loss function. We evaluate our method qualitatively and quantitatively on several challenging synthetic scenes generated using Blender with the spike camera simulator. Our results demonstrate that Spike-NeRF produces more visually appealing results than the existing methods and the baseline we proposed in high-speed scenes. Our code is available at https://github.com/yijiaguo02/SpikeNerf
Yijia Guo, Yuanxi Bai, Liwen Hu 0002, Mianzhi Liu, Lei Ma 0008, Tiejun Huang 0001
ICME7
2024 SCSim: A Realistic Spike Cameras Simulator
abstract
Spike cameras, with their exceptional temporal resolution, are revolutionizing high-speed visual applications. Large-scale synthetic datasets have significantly accelerated the development of these cameras, particularly in reconstruction and optical flow. However, current synthetic datasets for spike cameras lack sophistication. Addressing this gap, we introduce SCSim, a novel and more realistic spike camera simulator with a comprehensive noise model. SCSim is adept at autonomously generating driving scenarios and synthesizing corresponding spike streams. To enhance the fidelity of these streams, we’ve developed a comprehensive noise model tailored to the unique circuitry of spike cameras. Our evaluations demonstrate that SCSim outperforms existing simulation methods in generating authentic spike streams. Crucially, SCSim simplifies the creation of datasets, thereby greatly advancing spike-based visual tasks like reconstruction. Our project refers to https://github.com/Acnext/SCSim.
Liwen Hu 0002, Lei Ma 0008, Yijia Guo, Tiejun Huang 0001
ICME4
2024 Robust Stable Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are gaining popularity in deep learning due to their low energy budget on neuromorphic hardware. However, they still face challenges in lacking sufficient robustness to guard safety-critical applications such as autonomous driving. Many studies have been conducted to defend SNNs from the threat of adversarial attacks. This paper aims to uncover the robustness of SNN through the lens of the stability of nonlinear systems. We are inspired by the fact that searching for parameters altering the leaky integrate-and-fire dynamics can enhance their robustness. Thus, we dive into the dynamics of membrane potential perturbation and simplify the formulation of the dynamics. We present that membrane potential perturbation dynamics can reliably convey the intensity of perturbation. Our theoretical analyses imply that the simplified perturbation dynamics satisfy input-output stability. Thus, we propose a training framework with modified SNN neurons and to reduce the mean square of membrane potential perturbation aiming at enhancing the robustness of SNN. Finally, we experimentally verify the effectiveness of the framework in the setting of Gaussian noise training and adversarial training on the image classification task. Please refer to https://github.com/DingJianhao/stable-snn for our code implementation.
Jianhao Ding, Yujia Liu 0005, Zhaofei Yu, Tiejun Huang 0001
ICML5
2024 Enhancing Adversarial Robustness in SNNs with Sparse Gradients
abstract
Spiking Neural Networks (SNNs) have attracted great attention for their energy-efficient operations and biologically inspired structures, offering potential advantages over Artificial Neural Networks (ANNs) in terms of energy efficiency and interpretability. Nonetheless, similar to ANNs, the robustness of SNNs remains a challenge, especially when facing adversarial attacks. Existing techniques, whether adapted from ANNs or specifically designed for SNNs, exhibit limitations in training SNNs or defending against strong attacks. In this paper, we propose a novel approach to enhance the robustness of SNNs through gradient sparsity regularization. We observe that SNNs exhibit greater resilience to random perturbations compared to adversarial perturbations, even at larger scales. Motivated by this, we aim to narrow the gap between SNNs under adversarial and random perturbations, thereby improving their overall robustness. To achieve this, we theoretically prove that this performance gap is upper bounded by the gradient sparsity of the probability associated with the true label concerning the input image, laying the groundwork for a practical strategy to train robust SNNs by regularizing the gradient sparsity. We validate the effectiveness of our approach through extensive experiments on both image-based and event-based datasets. The results demonstrate notable improvements in the robustness of SNNs. Our work highlights the importance of gradient sparsity in SNNs and its role in enhancing robustness.
Yujia Liu 0005, Tong Bu, Jianhao Ding, Zecheng Hao, Tiejun Huang 0001, Zhaofei Yu
ICML5
2024 Unsupervised Spike Depth Estimation via Cross-modality Cross-domain Knowledge Transfer
abstract
Neuromorphic spike data, an upcoming modality with high temporal resolution, has shown promising potential in autonomous driving by mitigating the challenges posed by high-velocity motion blur. However, training the spike depth estimation network holds significant challenges in two aspects: sparse spatial information for pixel-wise tasks and difficulties in achieving paired depth labels for temporally intensive spike streams. Therefore, we introduce open-source RGB data to support spike depth estimation, leveraging its annotations and spatial information. The inherent differences in modalities and data distribution make it challenging to directly apply transfer learning from open-source RGB to target spike data. To this end, we propose a cross-modality cross-domain (BiCross) framework to realize unsupervised spike depth estimation by introducing simulated mediate source spike data. Specifically, we design a Coarse-to-Fine Knowledge Distillation (CFKD) approach to facilitate comprehensive cross-modality knowledge transfer while preserving the unique strengths of both modalities, utilizing a spike-oriented uncertainty scheme. Then, we propose a Self-Correcting Teacher-Student (SCTS) mechanism to screen out reliable pixel-wise pseudo labels and ease the domain shift of the student model, which avoids error accumulation in target spike data. To verify the effectiveness of BiCross, we conduct extensive experiments on four scenarios, including Synthetic to Real, Extreme Weather, Scene Changing, and Real Spike. Our method achieves state-of-the-art (SOTA) performances, compared with RGB-oriented unsupervised depth estimation methods. Code and dataset: https://github.com/Theia-4869/BiCross.
Jiaming Liu 0003, Qizhe Zhang, Xiaoqi Li 0020, Jianing Li 0001, Guanqun Wang, Ming Lu 0002, Tiejun Huang 0001, Shanghang Zhang
ICRA7
2024 ShapeMamba-EM: Fine-Tuning Foundation Model with Local Shape Descriptors and Mamba Blocks for 3D EM Image Segmentation
Ruohua Shi, Qiufan Pang, Lei Ma 0008, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001
MICCAI (12)5
2024 PRTGS: Precomputed Radiance Transfer of Gaussian Splats for Real-Time High-Quality Relighting
abstract
We proposed Precomputed Radiance Transfer of Gaussian Splats (PRTGS), a real-time high-quality relighting method for Gaussian splats in low-frequency lighting environments that captures soft shadows and interreflections by precomputing 3D Gaussian splats' radiance transfer. Existing studies have demonstrated that 3D Gaussian splatting (3DGS) outperforms neural fields in efficiency for dynamic lighting scenarios. However, the current relighting method based on 3DGS is still struggling to compute high-quality shadow and indirect illumination in real time for dynamic light, leading to unrealistic rendering results. We solve this problem by precomputing the expensive transport simulations required for complex transfer functions like shadowing, the resulting transfer functions are represented as dense sets of vectors or matrices for every Gaussian splat. We introduce distinct precomputing methods tailored for training and rendering stages, along with unique ray tracing and indirect lighting precomputation techniques for 3D Gaussian splats to accelerate training speed and compute accurate indirect lighting related to environment light. Experimental analyses demonstrate that our approach achieves state-of-the-art visual quality while maintaining competitive training times and importantly allows high-quality real-time (30+ fps) relighting for dynamic light and relatively complex scenes at 1080p resolution.
Yijia Guo, Yuanxi Bai, Liwen Hu 0002, Mianzhi Liu, Yu Cai 0008, Tiejun Huang 0001, Lei Ma 0008
ACM Multimedia7
2024 Towards High-performance Spiking Transformers from ANN to SNN Conversion
abstract
Spiking neural networks (SNNs) show great potential due to their energy efficiency, fast processing capabilities, and robustness. There are two main approaches to constructing SNNs. Direct training methods require much memory, while conversion methods offer a simpler and more efficient option. However, current conversion methods mainly focus on converting convolutional neural networks (CNNs) to SNNs. Converting Transformers to SNN is challenging because of the presence of non-linear modules. In this paper, we propose an Expectation Compensation Module to preserve the accuracy of the conversion. The core idea is to use information from the previous T time-steps to calculate the expected output at time-step T. We also propose a Multi-Threshold Neuron and the corresponding Parallel Parameter normalization to address the challenge of large time steps needed for high accuracy, aiming to reduce network latency and power consumption. Our experimental results demonstrate that our approach achieves state-of-the-art performance. For example, we achieve a top-1 accuracy of 88.60% with only a 1% loss in accuracy using 4 time steps while consuming only 35% of the original power of the Transformer. To our knowledge, this is the first successful Artificial Neural Network (ANN) to SNN conversion for Spiking Transformers that achieves high accuracy, low latency, and low power consumption on complex datasets. The source codes of the proposed method are available at https://github.com/h-z-h-cell/Transformer-to-SNN-ECMT.
Zihan Huang, Xinyu Shi 0004, Zecheng Hao, Tong Bu, Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001
ACM Multimedia7
2024 Real-time Parameter Evaluation of High-speed Microfluidic Droplets using Continuous Spike Streams
abstract
Droplet-based microfluidic devices, with their high throughput and low power consumption, have found wide-ranging applications in the life sciences, such as drug discovery and cancer detection. However, the lack of real-time methods for accurately estimating droplet generation parameters has resulted in droplet microfluidic systems remaining largely offline-controlled, making it challenging to achieve efficient feedback in droplet generation. To meet the real-time requirements, it's imperative to minimize the data throughput of the collection system while employing parameter estimation algorithms that are both resource-efficient and highly effective. Spike camera, as an innovative form of neuromorphic camera, facilitates high temporal resolution scene capture with comparatively low data throughput. In this paper, we propose a real-time evaluation method for high-speed droplet parameters based on spike-based microfluidic flow-focusing, named RTDE, that integrates spike camera into the droplet collection system to efficiently capture information using spike stream. To process the spike stream effectively, we develop a spike-based estimation algorithm for real-time droplet generation parameters. To validate the performance of our method, we collected spike-based droplet datasets (SDD), comprising synthetic and real data with varying flow velocities, frequencies, and droplet sizes. Experiments result on these datasets consistently demonstrate that our method achieves parameter estimations that closely match the ground truth values, showcasing high precision. Furthermore, comparative experiments with image-based parameter estimation methods highlight the superior time efficiency of our method, enabling real-time calculation of parameter estimations. Cdoe and datasets are avaliable at: https://github.com/Onetism/RTDE
Changqing Su, Yanqin Chen, Zhen Cheng 0005, Zhaofei Yu, Tiejun Huang 0001
ACM Multimedia8
2024 SpikeGS: 3D Gaussian Splatting from Spike Streams with High-Speed Camera Motion
Jiyuan Zhang 0005, Shiyan Chen, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu
ACM Multimedia5
2024 Towards Low-latency Event-based Visual Recognition with Hybrid Step-wise Distillation Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have garnered significant attention for their low power consumption and high biological interpretability. Their rich spatio-temporal information processing capability and event-driven nature make them ideally well-suited for neuromorphic datasets. However, current SNNs struggle to balance accuracy and latency in classifying these datasets. In this paper, we propose Hybrid Step-wise Distillation (HSD) method, tailored for neuromorphic datasets, to mitigate the notable decline in performance at lower time steps. Our work disentangles the dependency between the number of event frames and the time steps of SNNs, utilizing more event frames during the training stage to improve performance, while using fewer event frames during the inference stage to reduce latency. Nevertheless, the average output of SNNs across all time steps is susceptible to individual time step with abnormal outputs, particularly at extremely low time steps. To tackle this issue, we implement Step-wise Knowledge Distillation (SKD) module that considers variations in the output distribution of SNNs at each time step. Empirical evidence demonstrates that our method yields competitive performance in classification tasks on neuromorphic datasets, especially at lower time steps. Our code will be available at: https://github.com/hsw0929/HSD.
Xian Zhong, Shengwang Hu, Wenxuan Liu 0008, Wenxin Huang, Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001
ACM Multimedia7
2024 Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking Cameras
abstract
The spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network architectures that overlook the intrinsic collaboration of spatio-temporal complementary information. In this paper, we propose an efficient spatio-temporal interactive reconstruction network to jointly perform inter-frame feature alignment and intra-frame feature filtering in a coarse-to-fine manner. Specifically, it starts by extracting hierarchical features from a concise hybrid spike representation, then refines the motion fields and target frames scale-by-scale, ultimately obtaining a full-resolution output. Meanwhile, we introduce a symmetric interactive attention block and a multi-motion field estimation block to further enhance the interaction capability of the overall network. Experiments on synthetic and real-captured data show that our approach exhibits excellent performance while maintaining low model complexity.
Bin Fan 0002, Jiaoyang Yin, Yuchao Dai, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
NeurIPS5
2024 Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
abstract
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou
NeurIPS35
2024 SpikeReveal: Unlocking Temporal Sequences from Real Blurry Inputs with Spike Streams
abstract
Reconstructing a sequence of sharp images from the blurry input is crucial for enhancing our insights into the captured scene and poses a significant challenge due to the limited temporal features embedded in the image. Spike cameras, sampling at rates up to 40,000 Hz, have proven effective in capturing motion features and beneficial for solving this ill-posed problem. Nonetheless, existing methods fall into the supervised learning paradigm, which suffers from notable performance degradation when applied to real-world scenarios that diverge from the synthetic training data domain. To address these challenges, we propose the first self-supervised framework for the task of spike-guided motion deblurring. Our approach begins with the formulation of a spike-guided deblurring model that explores the theoretical relationships among spike streams, blurry images, and their corresponding sharp sequences. We subsequently develop a self-supervised cascaded framework to alleviate the issues of spike noise and spatial-resolution mismatching encountered in the deblurring model. With knowledge distillation and re-blurring loss, we further design a lightweight deblur network to generate high-quality sequences with brightness and texture consistency with the original input. Quantitative and qualitative experiments conducted on our real-world and synthetic datasets with spikes validate the superior generalization of the proposed framework. Our code, data and trained models are available at \url{https://github.com/chenkang455/S-SDM}.
Shiyan Chen, Jiyuan Zhang 0005, Baoyue Zhang, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu
NeurIPS6
2024 Learning from Pattern Completion: Self-supervised Controllable Generation
abstract
The human brain exhibits a strong ability to spontaneously associate different visual attributes of the same or similar visual scene, such as associating sketches and graffiti with real-world visual objects, usually without supervising information. In contrast, in the field of artificial intelligence, controllable generation methods like ControlNet heavily rely on annotated training datasets such as depth maps, semantic segmentation maps, and poses, which limits the method’s scalability. Inspired by the neural mechanisms that may contribute to the brain’s associative power, specifically the cortical modularization and hippocampal pattern completion, here we propose a self-supervised controllable generation (SCG) framework. Firstly, we introduce an equivariance constraint to promote inter-module independence and intra-module correlation in a modular autoencoder network, thereby achieving functional specialization. Subsequently, based on these specialized modules, we employ a self-supervised pattern completion approach for controllable generation training. Experimental results demonstrate that the proposed modular autoencoder effectively achieves functional specialization, including the modular processing of color, brightness, and edge detection, and exhibits brain-like features including orientation selectivity, color antagonism, and center-surround receptive fields. Through self-supervised training, associative generation capabilities spontaneously emerge in SCG, demonstrating excellent zero-shot generalization ability to various tasks such as superresolution, dehaze and associative or conditional generation on painting, sketches, and ancient graffiti. Compared to the previous representative method ControlNet, our proposed approach not only demonstrates superior robustness in more challenging high-noise scenarios but also possesses more promising scalability potential due to its self-supervised manner. Codes are released on Github and Gitee.
Guofan Fan, Jinying Gao, Lei Ma 0008, Tiejun Huang 0001
NeurIPS6
2024 SegVol: Universal and Interactive Volumetric Medical Image Segmentation
abstract
Precise image segmentation provides clinical study with instructive information. Despite the remarkable progress achieved in medical image segmentation, there is still an absence of a 3D foundation segmentation model that can segment a wide range of anatomical categories with easy user interaction. In this paper, we propose a 3D foundation segmentation model, named SegVol, supporting universal and interactive volumetric medical image segmentation. By scaling up training data to 90K unlabeled Computed Tomography (CT) volumes and 6K labeled CT volumes, this foundation model supports the segmentation of over 200 anatomical categories using semantic and spatial prompts. To facilitate efficient and precise inference on volumetric images, we design a zoom-out-zoom-in mechanism. Extensive experiments on 22 anatomical segmentation tasks verify that SegVol outperforms the competitors in 19 tasks, with improvements up to 37.24\% compared to the runner-up methods. We demonstrate the effectiveness and importance of specific designs by ablation study. We expect this foundation model can promote the development of volumetric medical image analysis. The model and code are publicly available at https://github.com/BAAI-DCAI/SegVol.
Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015
NeurIPS3
2024 LM-HT SNN: Enhancing the Performance of SNN to ANN Counterpart through Learnable Multi-hierarchical Threshold Model
abstract
Compared to traditional Artificial Neural Network (ANN), Spiking Neural Network (SNN) has garnered widespread academic interest for its intrinsic ability to transmit information in a more energy-efficient manner. However, despite previous efforts to optimize the learning algorithm of SNNs through various methods, SNNs still lag behind ANNs in terms of performance. The recently proposed multi-threshold model provides more possibilities for further enhancing the learning capability of SNNs. In this paper, we rigorously analyze the relationship among the multi-threshold model, vanilla spiking model and quantized ANNs from a mathematical perspective, then propose a novel LM-HT model, which is an equidistant multi-threshold model that can dynamically regulate the global input current and membrane potential leakage on the time dimension. The LM-HT model can also be transformed into a vanilla single threshold model through reparameterization, thereby achieving more flexible hardware deployment. In addition, we note that the LM-HT model can seamlessly integrate with ANN-SNN Conversion framework under special initialization. This novel hybrid learning framework can effectively improve the relatively poor performance of converted SNNs under low time latency. Extensive experimental results have demonstrated that our model can outperform previous state-of-the-art works on various types of datasets, which promote SNNs to achieve a brand-new level of performance comparable to quantized ANNs. Code is available at https://github.com/hzc1208/LMHT_SNN.
Zecheng Hao, Xinyu Shi 0004, Yujia Liu 0005, Zhaofei Yu, Tiejun Huang 0001
NeurIPS5
2024 Retrospective for the Dynamic Sensorium Competition for predicting large-scale mouse primary visual cortex activity from videos
abstract
Understanding how biological visual systems process information is challenging because of the nonlinear relationship between visual input and neuronal responses. Artificial neural networks allow computational neuroscientists to create predictive models that connect biological and machine vision.Machine learning has benefited tremendously from benchmarks that compare different models on the same task under standardized conditions. However, there was no standardized benchmark to identify state-of-the-art dynamic models of the mouse visual system.To address this gap, we established the SENSORIUM 2023 Benchmark Competition with dynamic input, featuring a new large-scale dataset from the primary visual cortex of ten mice. This dataset includes responses from 78,853 neurons to 2 hours of dynamic stimuli per neuron, together with behavioral measurements such as running speed, pupil dilation, and eye movements.The competition ranked models in two tracks based on predictive performance for neuronal responses on a held-out test set: one focusing on predicting in-domain natural stimuli and another on out-of-distribution (OOD) stimuli to assess model generalization.As part of the NeurIPS 2023 Competition Track, we received more than 160 model submissions from 22 teams. Several new architectures for predictive models were proposed, and the winning teams improved the previous state-of-the-art model by 50\%. Access to the dataset as well as the benchmarking infrastructure will remain online at www.sensorium-competition.net.
Polina Turishcheva, Paul G. Fahey, Michaela Vystrcilová, Laura Hansel, Rachel Froebe, Kayla Ponder, Yongrong Qiu, Konstantin Willeke, Mohammad Bashiri, Ruslan Baikulov, Yu Zhu 0008, Lei Ma 0008, Tiejun Huang 0001, Bryan Li, Wolf De Wulf, Nina Kudryashova, Matthias H. Hennig, Nathalie Rochefort, Arno Onken, Eric Y. Wang, Zhiwei Ding, Andreas S. Tolias, Fabian H. Sinz, Alexander S. Ecker
NeurIPS14
2024 Continuous Spatiotemporal Events Decoupling through Spike-based Bayesian Computation
abstract
Numerous studies have demonstrated that the cognitive processes of the human brain can be modeled using the Bayesian theorem for probabilistic inference of the external world. Spiking neural networks (SNNs), capable of performing Bayesian computation with greater physiological interpretability, offer a novel approach to distributed information processing in the cortex. However, applying these models to real-world scenarios to harness the advantages of brain-like computation remains a challenge. Recently, bio-inspired sensors with high dynamic range and ultra-high temporal resolution have been widely used in extreme vision scenarios. Event streams, generated by various types of motion, represent spatiotemporal data. Inferring motion targets from these streams without prior knowledge remains a difficult task. The Bayesian inference-based Expectation-Maximization (EM) framework has proven effective for motion segmentation in event streams, allowing for decoupling without prior information about the motion or its source. This work demonstrates that Bayesian computation based on spiking neural networks can decouple event streams of different motions. The Winner-Take-All (WTA) circuits in the constructed network implement an equivalent E-step, while STDP achieves an equivalent optimization in M-step. Through theoretical analysis and experiments, we show that STDP-based learning can maximize the contrast of warped events under mixed motion models. Experimental results show that the constructed spiking network can effectively segment the motion contained in event streams.
Yajing Zheng, Jiyuan Zhang 0005, Zhaofei Yu, Tiejun Huang 0001
NeurIPS4
2024 Iterative Fine-Grained Genetic Algorithm for Inferring Connection Weights in Large-Scale Biophysical Mouse V1 Model
Peize Li, Tiejun Huang 0001
PRICAI (4)4
2024 Light Flickering Guided Reflection Removal
Yuchen Hong, Yakun Chang, Jinxiu Liang, Lei Ma 0008, Tiejun Huang 0001, Boxin Shi
Int. J. Comput. Vis.5
2024 EVA-02: A visual representation for neon genesis
Xinggang Wang, Tiejun Huang 0001, Yue Cao 0001
Image Vis. Comput.4
2024 Hybrid All-in-Focus Imaging From Neuromorphic Focal Stack
abstract
Creating an image focal stack requires multiple shots, which captures images at different depths within the same scene. Such methods are not suitable for scenes undergoing continuous changes. Achieving an all-in-focus image from a single shot poses significant challenges, due to the highly ill-posed nature of rectifying defocus and deblurring from a single image. In this paper, to restore an all-in-focus image, we introduce the neuromorphic focal stack, which is defined as neuromorphic signal streams captured by an event/ a spike camera during a continuous focal sweep, aiming to restore an all-in-focus image. Given an RGB image focused at any distance, we harness the high temporal resolution of neuromorphic signal streams. From neuromorphic signal streams, we automatically select refocusing timestamps and reconstruct corresponding refocused images to form a focal stack. Guided by the neuromorphic signal around the selected timestamps, we can merge the focal stack using proper weights and restore a sharp all-in-focus image. We test our method on two distinct neuromorphic cameras. Experimental results from both synthetic and real datasets demonstrate a marked improvement over existing State-of-the-Art methods.
Minggui Teng, Hanyue Lou, Yixin Yang 0008, Tiejun Huang 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 A Bayesian Approach Toward Robust Multidimensional Ellipsoid-Specific Fitting
abstract
This work presents a novel and effective method for fitting multidimensional ellipsoids (i.e., ellipsoids embedded in [Formula: see text]) to scattered data in the contamination of noise and outliers. Unlike conventional algebraic or geometric fitting paradigms that assume each measurement point is a noisy version of its nearest point on the ellipsoid, we approach the problem as a Bayesian parameter estimate process and maximize the posterior probability of a certain ellipsoidal solution given the data. We establish a more robust correlation between these points based on the predictive distribution within the Bayesian framework, i.e., considering each model point as a potential source for generating each measurement. Concretely, we incorporate a uniform prior distribution to constrain the search for primitive parameters within an ellipsoidal domain, ensuring ellipsoid-specific results regardless of inputs. We then establish the connection between measurement point and model data via Bayes' rule to enhance the method's robustness against noise. Due to independent of spatial dimensions, the proposed method not only delivers high-quality fittings to challenging elongated ellipsoids but also generalizes well to multidimensional spaces. To address outlier disturbances, often overlooked by previous approaches, we further introduce a uniform distribution on top of the predictive distribution to significantly enhance the algorithm's robustness against outliers. Thanks to the uniform prior, our maximum a posterior probability coincides with a more tractable maximum likelihood estimation problem, which is subsequently solved by a numerically stable Expectation Maximization (EM) framework. Moreover, we introduce an ε-accelerated technique to expedite the convergence of EM considerably. We also investigate the relationship between our algorithm and conventional least-squares-based ones, during which we theoretically prove our method's superior robustness. To the best of our knowledge, this is the first comprehensive method capable of performing multidimensional ellipsoid-specific fitting within the Bayesian optimization paradigm under diverse disturbances. We evaluate it across lower and higher dimensional spaces in the presence of heavy noise, outliers, and substantial variations in axis ratios. Also, we apply it to a wide range of practical applications such as microscopy cell counting, 3D reconstruction, geometric shape approximation, and magnetometer calibration tasks. In all these test contexts, our method consistently delivers flexible, robust, ellipsoid-specific performance, and achieves the state-of-the-art results.
Mingyang Zhao 0001, Xiaohong Jia 0001, Lei Ma 0008, Yuke Shi, Jingen Jiang 0001, Qizhai Li, Dong-Ming Yan 0001, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 Spike Camera Image Reconstruction Using Deep Spiking Neural Networks
abstract
Spike camera is a bio-inspired sensor with ultra-high temporal resolution and low energy consumption. It captures visual signals using an “integrate-and-fire" mechanism and outputs a continuous stream of binary spikes. Reconstructing image sequence from spikes streams is critical for spike camera. Several reconstruction methods have been proposed in recent years. However, the computational cost of these methods is relatively high. Inspired by the fact that spiking neural networks (SNNs) are energy efficient and support time-series signal processing inherently, we propose a lightweight SNN for spike camera image reconstruction (abbreviated to SSIR). Experimental results show that SSIR achieves comparable performance with the state-of-the-art (SOTA) methods at much lower computation and energy cost.
Rui Zhao 0010, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Shuyuan Zhu, Lei Ma 0008, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 SpiReco: Fast and Efficient Recognition of High-Speed Moving Objects With Spike Camera
abstract
Benefited from the high temporal resolution and high dynamic range, spike cameras have shown great potential in recognizing high-speed moving objects. However, the computer vision community has not explored this task due to the lack of spike data and annotations of high-speed moving objects. This paper contributes a novel dataset, namedSpiReco(Spiking datasets forRecognition), by recording high-speed moving objects using a spike camera. To annotate the dataset, image labels from established datasets such as MNIST, CIFAR10, and CALTECH101 are utilized. Based on this new dataset, this paper proposes the first spike-based object recognition framework. The proposed framework includes a denoise module, which is designed to suppress spike noise by learning spatio-temporal correlation from neighbouring pixels. Additionally, a motion enhancement module is introduced to address high-speed and random motions. Afterward, binarized neural networks are adopted to save computation costs. These efforts result in a fast and efficient processing framework for spiking data. Experimental results demonstrate the effectiveness of the proposed methods. For example, the proposed spike-based recognition framework achieves 80.2% accuracy in recognizing 101 classes of high-speed moving objects using only 2.2ms of spike streams. The SpiReco is available at https://github.com/Evin-X/SpiReco.
Junwei Zhao 0003, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Learning a Deep Demosaicing Network for Spike Camera With Color Filter Array
abstract
For capturing dynamic scenes with ultra-fast motion, neuromorphic cameras with extremely high temporal resolution have demonstrated their great capability and potential. Different from the event cameras that only record relative changes in light intensity, spike camera fires a stream of spikes according to a full-time accumulation of photons so that it can recover the texture details for both static areas and dynamic areas. Recently, color spike camera has been invented to record color information of dynamic scenes using a color filter array (CFA). However, demosaicing for color spike cameras is an open and challenging problem. In this paper, we develop a demosaicing network, called CSpkNet, to reconstruct dynamic color visual signals from the spike stream captured by the color spike camera. Firstly, we develop a light inference module to convert binary spike streams to intensity estimates. In particular, a feature-based channel attention module is proposed to reduce the noises caused by quantization errors. Secondly, considering both the Bayer configuration and object motion, we propose a motion-guided filtering module to estimate the missing pixels of each color channel, without undesired motion blur. Finally, we design a refinement module to improve the intensity and details, utilizing the color correlation. Experimental results demonstrate that CSpkNet can reconstruct color images from the Bayer-pattern spike stream with promising visual quality.
Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
IEEE Trans. Image Process.7
2024 GIN: Generative INvariant Shape Prior for Amodal Instance Segmentation
abstract
Amodal instance segmentation (AIS) predicts the complete shape of the occluded object, including both visible and occluded regions. Because visual clues are lacking, the occluded region is difficult to segment accurately. In human amodal perception, shape-prior knowledge is helpful for AIS. The previous method uses a 2D shape prior byrote memorizing, establishing a shape dictionary and retrieving the closest mask to the segmentation result. However, this approach cannot obtain the shape prior, which is not prestored in the shape dictionary. In this article, to improve generalization ability, we propose a generative invariant shape-prior network (GIN), simulating the human perception process that learns the basic shape, which is invariant to transformations, including translation, rotation, and scaling. We designa novel framework that decouples the learning of shape priors from transformation. GIN is end-to-end trainable and needs no dictionary establishment, making the whole pipeline efficient. GIN outperforms state-of-the-art methods on three public datasets (D2SA, COCOA-cls, and KINS) with large margins.
Weining Ye, Tingting Jiang 0001, Tiejun Huang 0001
IEEE Trans. Multim.4
2023 Self-Supervised Joint Dynamic Scene Reconstruction and Optical Flow Estimation for Spiking Camera
abstract
Spiking camera, a novel retina-inspired vision sensor, has shown its great potential for capturing high-speed dynamic scenes with a sampling rate of 40,000 Hz. The spiking camera abandons the concept of exposure window, with each of its photosensitive units continuously capturing photons and firing spikes asynchronously. However, the special sampling mechanism prevents the frame-based algorithm from being used to spiking camera. It remains to be a challenge to reconstruct dynamic scenes and perform common computer vision tasks for spiking camera. In this paper, we propose a self-supervised joint learning framework for optical flow estimation and reconstruction of spiking camera. The framework reconstructs clean frame-based spiking representations in a self-supervised manner, and then uses them to train the optical flow networks. We also propose an optical flow based inverse rendering process to achieve self-supervision by minimizing the difference with respect to the original spiking temporal aggregation image. The experimental results demonstrate that our method bridges the gap between synthetic and real-world scenes and achieves desired results in real-world scenarios. To the best of our knowledge, this is the first attempt to jointly reconstruct dynamic scenes and estimate optical flow for spiking camera from a self-supervised learning perspective.
Shiyan Chen, Zhaofei Yu, Tiejun Huang 0001
AAAI3
2023 Reducing ANN-SNN Conversion Error through Residual Membrane Potential
abstract
Spiking Neural Networks (SNNs) have received extensive academic attention due to the unique properties of low power consumption and high-speed computing on neuromorphic chips. Among various training methods of SNNs, ANN-SNN conversion has shown the equivalent level of performance as ANNs on large-scale datasets. However, unevenness error, which refers to the deviation caused by different temporal sequences of spike arrival on activation layers, has not been effectively resolved and seriously suffers the performance of SNNs under the condition of short time-steps. In this paper, we make a detailed analysis of unevenness error and divide it into four categories. We point out that the case of the ANN output being zero while the SNN output being larger than zero accounts for the largest percentage. Based on this, we theoretically prove the sufficient and necessary conditions of this case and propose an optimization strategy based on residual membrane potential to reduce unevenness error. The experimental results show that the proposed method achieves state-of-the-art performance on CIFAR-10, CIFAR-100, and ImageNet datasets. For example, we reach top-1 accuracy of 64.32% on ImageNet with 10-steps. To the best of our knowledge, this is the first time ANN-SNN conversion can simultaneously achieve high accuracy and ultra-low-latency on the complex dataset. Code is available at https://github.com/hzc1208/ANN2SNN_SRP.
Zecheng Hao, Tong Bu, Jianhao Ding, Tiejun Huang 0001, Zhaofei Yu
AAAI4
2023 SVFI: Spiking-Based Video Frame Interpolation for High-Speed Motion
abstract
Occlusion and motion blur make it challenging to interpolate video frame, since estimating complex motions between two frames is hard and unreliable, especially in highly dynamic scenes. This paper aims to address these issues by exploiting spike stream as auxiliary visual information between frames to synthesize target frames. Instead of estimating motions by optical flow from RGB frames, we present a new dual-modal pipeline adopting both RGB frames and the corresponding spike stream as inputs (SVFI). It extracts the scene structure and objects' outline feature maps of the target frames from spike stream. Those feature maps are fused with the color and texture feature maps extracted from RGB frames to synthesize target frames. Benefited by the spike stream that contains consecutive information between two frames, SVFI can directly extract the information in occlusion and motion blur areas of target frames from spike stream, thus it is more robust than previous optical flow-based methods. Experiments show SVFI outperforms the SOTA methods on wide variety of datasets. For instance, in 7 and 15 frame skip evaluations, it shows up to 5.58 dB and 6.56 dB improvements in terms of PSNR over the corresponding second best methods BMBC and DAIN. SVFI also shows visually impressive performance in real-world scenes.
Lujie Xia, Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001
AAAI4
2023 Learning Temporal-Ordered Representation for Spike Streams Based on Discrete Wavelet Transforms
abstract
Spike camera, a new type of neuromorphic visual sensor that imitates the sampling mechanism of the primate fovea, can capture photons and output 40000 Hz binary spike streams. Benefiting from the asynchronous sampling mechanism, the spike camera can record fast-moving objects and clear images can be recovered from the spike stream at any specified timestamps without motion blurring. Despite these, due to the dense time sequence information of the discrete spike stream, it is not easy to directly apply the existing algorithms of traditional cameras to the spike camera. Therefore, it is necessary and interesting to explore a universally effective representation of dense spike streams to better fit various network architectures. In this paper, we propose to mine temporal-robust features of spikes in time-frequency space with wavelet transforms. We present a novel Wavelet-Guided Spike Enhancing (WGSE) paradigm consisting of three consecutive steps: multi-level wavelet transform, CNN-based learnable module, and inverse wavelet transform. With the assistance of WGSE, the new streaming representation of spikes can be learned. We demonstrate the effectiveness of WGSE on two downstream tasks, achieving state-of-the-art performance on the image reconstruction task and getting considerable performance on semantic segmentation. Furthermore, We build a new spike-based synthesized dataset for semantic segmentation. Code and Datasets are available at https://github.com/Leozhangjiyuan/WGSE-SpikeCamera.
Jiyuan Zhang 0005, Shanshan Jia 0001, Zhaofei Yu, Tiejun Huang 0001
AAAI4
2023 Learning to Super-resolve Dynamic Scenes for Neuromorphic Spike Camera
abstract
Spike camera is a kind of neuromorphic sensor that uses a novel ``integrate-and-fire'' mechanism to generate a continuous spike stream to record the dynamic light intensity at extremely high temporal resolution. However, as a trade-off for high temporal resolution, its spatial resolution is limited, resulting in inferior reconstruction details. To address this issue, this paper develops a network (SpikeSR-Net) to super-resolve a high-resolution image sequence from the low-resolution binary spike streams. SpikeSR-Net is designed based on the observation model of spike camera and exploits both the merits of model-based and learning-based methods. To deal with the limited representation capacity of binary data, a pixel-adaptive spike encoder is proposed to convert spikes to latent representation to infer clues on intensity and motion. Then, a motion-aligned super resolver is employed to exploit long-term correlation, so that the dense sampling in temporal domain can be exploited to enhance the spatial resolution without introducing motion blur. Experimental results show that SpikeSR-Net is promising in super-resolving higher-quality images for spike camera.
Jing Zhao 0011, Ruiqin Xiong, Jian Zhang 0018, Rui Zhao 0010, Hangfan Liu, Tiejun Huang 0001
AAAI6
2023 1000 FPS HDR Video with a Spike-RGB Hybrid Camera
abstract
Capturing high frame rate and high dynamic range (HFR&HDR) color videos in high-speed scenes with conventional frame-based cameras is very challenging. The increasing frame rate is usually guaranteed by using shorter exposure time so that the captured video is severely interfered by noise. Alternating exposures can alleviate the noise issue but sacrifice frame rate due to involving long-exposure frames. The neuromorphic spiking camera records high-speed scenes of high dynamic range without colors using a completely different sensing mechanism and visual representation. We introduce a hybrid camera system composed of a spiking and an alternating-exposure RGB camera to capture HFR&HDR scenes with high fidelity. Our insight is to bring each camera's superiority into full play. The spike frames, with accurate fast motion information encoded, are firstly reconstructed for motion representation, from which the spike-based optical flows guide the recovery of missing temporal information for long-exposure RGB images while retaining their reliable color appearances. With the strong temporal constraint estimated from spike trains, both missing and distorted colors cross RGB frames are recovered to generate time-consistent and HFR color frames. We collect a new Spike-RGB dataset that contains 300 sequences of synthetic data and 20 groups of real-world data to demonstrate 1000 FPS HDR videos outperforming HDR video reconstruction methods and commercial high-speed cameras.
Yakun Chang, Chu Zhou, Yuchen Hong, Liwen Hu 0002, Chao Xu 0002, Tiejun Huang 0001, Boxin Shi
CVPR6
2023 EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
abstract
We launch EVA, a vision-centric foundation model to Explore the limits of Visual representation at scAle using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches. Via this pretext task, we can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks, such as image recognition, video action recognition, object detection, instance segmentation and semantic segmentation without heavy supervised training. Moreover, we observe quantitative changes in scaling EVA result in qualitative changes in transfer learning performance that are not present in other models. For instance, EVA takes a great leap in the challenging large vocabulary instance segmentation task: our model achieves almost the same state-of-the-art performance on LVIS dataset with over a thousand categories and COCO dataset with only eighty categories. Beyond a pure vision encoder, EVA can also serve as a vision-centric, multi-modal pivot to connect images and text. We find initializing the vision tower of a giant CLIP from EVA can greatly stabilize the training and outperform the training from scratch counterpart with much fewer samples and less compute, providing a new direction for scaling up and accelerating the costly training of multi-modal foundation models.
Wen Wang 0015, Binhui Xie, Ledell Wu, Xinggang Wang, Tiejun Huang 0001, Yue Cao 0001
CVPR7
2023 Images Speak in Images: A Generalist Painter for In-Context Visual Learning
abstract
In-context learning, as a new paradigm in NLP, allows the model to rapidly adapt to various tasks with only a handful of prompts and examples. But in computer vision, the difficulties for in-context learning lie in that tasks vary significantly in the output representations, thus it is unclear how to define the general-purpose task prompts that the vision model can understand and transfer to out-of-domain tasks. In this work, we present Painter, a generalist model which addresses these obstacles with an “image”-centric solution, that is, to redefine the output of core vision tasks as images, and specify task prompts as also images. With this idea, our training process is extremely simple, which performs standard masked image modeling on the stitch of input and output image pairs. This makes the model capable of performing tasks conditioned on visible image patches. Thus, during inference, we can adopt a pair of input and output images from the same task as the input condition, to indicate which task to perform. Without bells and whistles, our generalist Painter can achieve competitive performance compared to well-established task-specific models, on seven representative vision tasks ranging from high-level visual understanding to low-level image processing. In addition, Painter significantly outperforms recent generalist models on several challenging tasks.
Wen Wang 0015, Yue Cao 0001, Chunhua Shen, Tiejun Huang 0001
CVPR5
2023 OAFormer: Learning Occlusion Distinguishable Feature for Amodal Instance Segmentation
abstract
The Amodal Instance Segmentation (AIS) task aims to infer the complete mask of occluded instance. Under many circumstances, existing methods treat occluded objects as unoccluded ones, and vice versa, leading to inaccurate predictions. This is because existing AIS methods do not explicitly utilize the occlusion rates of each object as supervision. However, occlusion information is critical for the methods to recognize whether the target objects are occluded. Hence we believe it is vital for the method to be distinguishable about the degree of occlusion for each instance. In this paper, a simple yet effective Occlusion-aware transformer-based model, OAFormer, is proposed for accurate amodal instance segmentation. The goal of OAFormer is to learn the occlusion discriminative features. Novel components are proposed to enable OAFormer to be occlusion distinguishable. We conduct extensive experiments on two challenging AIS datasets to evaluate the effectiveness of our method. OAFormer outperforms state-of-the-art methods by large margins.
Ruohua Shi, Tiejun Huang 0001, Tingting Jiang 0001
ICASSP3
2023 MUVA: A New Large-Scale Benchmark for Multi-view Amodal Instance Segmentation in the Shopping Scenario
abstract
Amodal Instance Segmentation (AIS) endeavors to accurately deduce complete object shapes that are partially or fully occluded. However, the inherent ill-posed nature of single-view datasets poses challenges in determining occluded shapes. A multi-view framework may help alleviate this problem, as humans often adjust their perspective when encountering occluded objects. At present, this approach has not yet been explored by existing methods and datasets. To bridge this gap, we propose a new task called Multi-view Amodal Instance Segmentation (MAIS) and introduce the MUVA dataset, the first MUlti-View AIS dataset that takes the shopping scenario as instantiation. MUVA provides comprehensive annotations, including multi-view amodal/visible segmentation masks, 3D models, and depth maps, making it the largest image-level AIS dataset in terms of both the number of images and instances. Additionally, we propose a new method for aggregating representative features across different instances and views, which demonstrates promising results in accurately predicting occluded objects from one viewpoint by leveraging information from other viewpoints. Besides, we also demonstrate that MUVA can benefit the AIS task in real-world scenarios.1
Weining Ye, Juan R. Terven, Zachary Bennett, Tingting Jiang 0001, Tiejun Huang 0001
ICCV7
2023 SegGPT: Towards Segmenting Everything In Context
abstract
We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images. The training of SegGPT is formulated as an in-context coloring problem with random color mapping for each data sample. The objective is to accomplish diverse tasks according to the context, rather than relying on specific colors. After training, SegGPT can perform arbitrary segmentation tasks in images or videos via in-context inference, such as object instance, stuff, part, contour, and text. SegGPT is evaluated on a broad range of tasks, including few-shot semantic segmentation, video object segmentation, semantic segmentation, and panoptic segmentation. Our results show strong capabilities in segmenting in-domain and out-of-domain targets, either qualitatively or quantitatively.
Yue Cao 0001, Wen Wang 0015, Chunhua Shen, Tiejun Huang 0001
ICCV6
2023 Bridging the Gap between ANNs and SNNs by Calibrating Offset Spikes
Zecheng Hao, Jianhao Ding, Tong Bu, Tiejun Huang 0001, Zhaofei Yu
ICLR4
2023 Entity Divider with Language Grounding in Multi-Agent Reinforcement Learning
abstract
We investigate the use of natural language to drive the generalization of policies in multi-agent settings. Unlike single-agent settings, the generalization of policies should also consider the influence of other agents. Besides, with the increasing number of entities in multi-agent settings, more agent-entity interactions are needed for language grounding, and the enormous search space could impede the learning process. Moreover, given a simple general instruction, e.g., beating all enemies, agents are required to decompose it into multiple subgoals and figure out the right one to focus on. Inspired by previous work, we try to address these issues at the entity level and propose a novel framework for language grounding in multi-agent reinforcement learning, entity divider (EnDi). EnDi enables agents to independently learn subgoal division at the entity level and act in the environment based on the associated entities. The subgoal division is regularized by agent modeling to avoid subgoal conflicts and promote coordinated strategies. Empirically, EnDi demonstrates the strong generalization ability to unseen games with new dynamics and expresses the superiority over existing methods. The code is available at https://github.com/PKU-RL/EnDi.
Ziluo Ding, Wanpeng Zhang 0002, Junpeng Yue, Tiejun Huang 0001, Zongqing Lu 0002
ICML5
2023 Optimization-Inspired Deep Network for Image Restoration from Partial Random Samples
abstract
Image Restoration from Partial Random Samples (RRS) has been studied in many image restoration works. There are also some attempts to use convolutional neural networks (CNNs) to handle it. However, most existing neural network-based methods perform poorly in generalization and we need to train a specific model for each degradation situation. Besides, the sampling mask which represents the positions of the sampled pixels is not used effectively in these methods. To address the problems, we propose an optimization-inspired network called RRSNet based on our derivation of the iterative optimization formulas for RRS. In our method, we design a CNN with two encoders and one decoder for training, setting up a flexible and effective prior. To make the most of the sampling information, we concatenate the degraded image with the mask and input them into one encoder for better generalization. Then we split the pixels into two groups according to the mask and extract their features as the input of another encoder. Experiments demonstrate that our RRSNet with the mask input can handle various sampling ratios using only one trained model and achieve the best restoration performance among all comparison methods.
Yanchen Dong 0001, Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Xiaopeng Fan 0001, Tiejun Huang 0001
ISCAS6
2023 HumVis: Human-Centric Visual Analysis System
abstract
Human-centric visual analysis is a fundamental task for many multimedia and computer vision applications, such as self-driving, multimedia retrieval, and augmented reality, etc. Based on our recent research efforts on fine-grained human visual analysis, we develop a robust and efficient human-centric visual analysis system named as HumVis. HumVis is built on a simple yet efficient contextual instance decoupling (CID) module, which can effectively separate different persons in an input image and output corresponding person structure information for visual analysis. Based on CID, HumVis achieves accurate multi-person pose estimation, multi-person foreground segmentation, multi-person part segmentation and 3D human mesh recovery for user-uploaded images/videos and support live stream presentation.
Dongkai Wang, Shiliang Zhang, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ACM Multimedia5
2023 Recognizing High-Speed Moving Objects with Spike Camera
abstract
Spike camera is a novel bio-inspired vision sensor that mimics the sampling mechanism of the primate fovea. It presents high temporal resolution and dynamic range, showing great potentials in the high-speed moving object recognition task, which has not been fully explored in the Multimedia community due to the lack of data and annotations. This paper contributes the first large-scale High-Speed Spiking Recognition (HSSR) dataset, by recording high-speed moving objects using a spike camera. The HSSR dataset contains 135,000 indoor objects annotated using ImageNet labels and 3,100 outdoor objects collected from real-world scenarios. Furthermore, we propose an original spiking recognition framework, which employs long-term spike stream features to supervise the feature learning from short-term spike streams. This framework improves the recognition accuracy, meanwhile substantially decreasing the recognition latency, making our method can accurately recognize moving objects at an equivalent speed of 514 km/h, using only 1 ms of spike stream. Experimental results show that, the proposed method achieves 76.5% accuracy for recognizing 100 fine-grained indoor objects and 84.3% accuracy for recognizing 8 outdoor objects using 1 ms of spike streams. Resources will be available at https://github.com/Evin-X/HSSR.
Junwei Zhao 0003, Jianming Ye, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001
ACM Multimedia5
2023 Enhancing Motion Deblurring in High-Speed Scenes with Spike Streams
abstract
Traditional cameras produce desirable vision results but struggle with motion blur in high-speed scenes due to long exposure windows. Existing frame-based deblurring algorithms face challenges in extracting useful motion cues from severely blurred images. Recently, an emerging bio-inspired vision sensor known as the spike camera has achieved an extremely high frame rate while preserving rich spatial details, owing to its novel sampling mechanism. However, typical binary spike streams are relatively low-resolution, degraded image signals devoid of color information, making them unfriendly to human vision. In this paper, we propose a novel approach that integrates the two modalities from two branches, leveraging spike streams as auxiliary visual cues for guiding deblurring in high-speed motion scenes. We propose the first spike-based motion deblurring model with bidirectional information complementarity. We introduce a content-aware motion magnitude attention module that utilizes learnable mask to extract relevant information from blurry images effectively, and we incorporate a transposed cross-attention fusion module to efficiently combine features from both spike data and blurry RGB images. Furthermore, we build two extensive synthesized datasets for training and validation purposes, encompassing high-temporal-resolution spikes, blurry images, and corresponding sharp images. The experimental results demonstrate that our method effectively recovers clear RGB images from highly blurry scenes and outperforms state-of-the-art deblurring algorithms in multiple settings.
Shiyan Chen, Jiyuan Zhang 0005, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu
NeurIPS4
2023 Slow and Weak Attractor Computation Embedded in Fast and Strong E-I Balanced Neural Dynamics
abstract
Attractor networks require neuronal connections to be highly structured in order to maintain attractor states that represent information, while excitation and inhibition balanced networks (E-INNs) require neuronal connections to be random and sparse to generate irregular neuronal firings. Despite being regarded as canonical models of neural circuits, both types of networks are usually studied in isolation, and it remains unclear how they coexist in the brain, given their very different structural demands. In this study, we investigate the compatibility of continuous attractor neural networks (CANNs) and E-INNs. In line with recent experimental data, we find that a neural circuit can exhibit both the traits of CANNs and E-INNs if the neuronal synapses consist of two sets: one set is strong and fast for irregular firing, and the other set is weak and slow for attractor dynamics. Our results from simulations and theoretical analysis reveal that the network also exhibits enhanced performance compared to the case of using only one set of synapses, with accelerated convergence of attractor states and retained E-I balanced condition for localized input. We also apply the network model to solve a real-world tracking problem and demonstrate that it can track fast-moving objects well. We hope that this study provides insight into how structured neural computations are realized by irregular firings of neurons.
Xiaohan Lin, Liyuan Li, Boxin Shi, Tiejun Huang 0001, Yuanyuan Mi, Si Wu 0001
NeurIPS4
2023 Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike Camera
abstract
Efficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timing representation for spike streams. Based on multi-layers architecture, it applies dilated convolutions on temporal dimension to extract features on multi-temporal scales with few parameters. And we design layer attention to dynamically fuse these features. Moreover, we propose an unsupervised learning method for optical flow estimation in a spike-based manner to break the dependence on labeled data. In addition, to verify the robustness, we also build a spike-based synthetic validation dataset for extreme scenarios in autonomous driving, denoted as SSES dataset. It consists of various corner cases. Experiments show that our method can predict optical flow from spike streams in different high-speed scenes, including real scenes. For instance, our method achieves $15\%$ and $19\%$ error reduction on PHM dataset compared to the best spike-based work, SCFlow, in $\Delta t=10$ and $\Delta t=20$ respectively, using the same settings as in previous works. The source code and dataset are available at \href{https://github.com/Bosserhead/USFlow}{https://github.com/Bosserhead/USFlow}.
Lujie Xia, Ziluo Ding, Rui Zhao 0010, Jiyuan Zhang 0005, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001, Ruiqin Xiong
NeurIPS7
2023 Exploring Loss Functions for Time-based Training Strategy in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are considered promising brain-inspired energy-efficient models due to their event-driven computing paradigm. The spatiotemporal spike patterns used to convey information in SNNs consist of both rate coding and temporal coding, where the temporal coding is crucial to biological-plausible learning rules such as spike-timing-dependent-plasticity. The time-based training strategy is proposed to better utilize the temporal information in SNNs and learn in an asynchronous fashion. However, some recent works train SNNs by the time-based scheme with rate-coding-dominated loss functions. In this paper, we first map rate-based loss functions to time-based counterparts and explain why they are also applicable to the time-based training scheme. After that, we infer that loss functions providing adequate positive overall gradients help training by theoretical analysis. Based on this, we propose the enhanced counting loss to replace the commonly used mean square counting loss. In addition, we transfer the training of scale factor in weight standardization into thresholds. Experiments show that our approach outperforms previous time-based training methods in most datasets. Our work provides insights for training SNNs with time-based schemes and offers a fresh perspective on the correlation between rate coding and temporal coding. Our code is available at https://github.com/zhuyaoyu/SNN-temporal-training-losses.
Yaoyu Zhu, Wei Fang 0006, Tiejun Huang 0001, Zhaofei Yu
NeurIPS4
2023 PS-Net: human perception-guided segmentation network for EM cell membrane
abstract
MOTIVATION: Cell membrane segmentation in electron microscopy (EM) images is a crucial step in EM image processing. However, while popular approaches have achieved performance comparable to that of humans on low-resolution EM datasets, they have shown limited success when applied to high-resolution EM datasets. The human visual system, on the other hand, displays consistently excellent performance on both low and high resolutions. To better understand this limitation, we conducted eye movement and perceptual consistency experiments. Our data showed that human observers are more sensitive to the structure of the membrane while tolerating misalignment, contrary to commonly used evaluation criteria. Additionally, our results indicated that the human visual system processes images in both global-local and coarse-to-fine manners. RESULTS: Based on these observations, we propose a computational framework for membrane segmentation that incorporates these characteristics of human perception. This framework includes a novel evaluation metric, the perceptual Hausdorff distance (PHD), and an end-to-end network called the PHD-guided segmentation network (PS-Net) that is trained using adaptively tuned PHD loss functions and a multiscale architecture. Our subjective experiments showed that the PHD metric is more consistent with human perception than other criteria, and our proposed PS-Net outperformed state-of-the-art methods on both low- and high-resolution EM image datasets as well as other natural image datasets. AVAILABILITY AND IMPLEMENTATION: The code and dataset can be found at https://github.com/EmmaSRH/PS-Net.
Ruohua Shi, Keyan Bi, Lei Ma 0008, Fang Fang 0003, Ling-Yu Duan, Tingting Jiang 0001, Tiejun Huang 0001
Bioinform.8
2023 Heuristic Tree-Partition-Based Parallel Method for Biophysically Detailed Neuron Simulation
abstract
Biophysically detailed neuron simulation is a powerful tool to explore the mechanisms behind biological experiments and bridge the gap between various scales in neuroscience research. However, the extremely high computational complexity of detailed neuron simulation restricts the modeling and exploration of detailed network models. The bottleneck is solving the system of linear equations. To accelerate detailed simulation, we propose a heuristic tree-partition-based parallel method (HTP) to parallelize the computation of the Hines algorithm, the kernel for solving linear equations, and leverage the strong parallel capability of the graphic processing unit (GPU) to achieve further speedup. We formulate the problem of how to get a fine parallel process as a tree-partition problem. Next, we present a heuristic partition algorithm to obtain an effective partition to efficiently parallelize the equation-solving process in detailed simulation. With further optimization on GPU, our HTP method achieves 2.2 to 8.5 folds speedup compared to the state-of-the-art GPU method and 36 to 660 folds speedup compared to the typical Hines algorithm.
Yichen Zhang 0002, Tiejun Huang 0001
Neural Comput.3
2023 Visual information processing through the interplay between fine and coarse signal pathways
abstract
Object recognition is often viewed as a feedforward, bottom-up process in machine learning, but in real neural systems, object recognition is a complicated process which involves the interplay between two signal pathways. One is the parvocellular pathway (P-pathway), which is slow and extracts fine features of objects; the other is the magnocellular pathway (M-pathway), which is fast and extracts coarse features of objects. It has been suggested that the interplay between the two pathways endows the neural system with the capacity of processing visual information rapidly, adaptively, and robustly. However, the underlying computational mechanism remains largely unknown. In this study, we build a two-pathway model to elucidate the computational properties associated with the interactions between two visual pathways. Specifically, we model two visual pathways using two convolution neural networks: one mimics the P-pathway, referred to as FineNet, which is deep, has small-size kernels, and receives detailed visual inputs; the other mimics the M-pathway, referred to as CoarseNet, which is shallow, has large-size kernels, and receives blurred visual inputs. We show that CoarseNet can learn from FineNet through imitation to improve its performance, FineNet can benefit from the feedback of CoarseNet to improve its robustness to noise; and the two pathways interact with each other to achieve rough-to-fine information processing. Using visual backward masking as an example, we further demonstrate that our model can explain visual cognitive behaviors that involve the interplay between two pathways. We hope that this study gives us insight into understanding the interaction principles between two visual pathways.
Xiaolong Zou, Zilong Ji, Tianqiu Zhang, Tiejun Huang 0001, Si Wu 0001
Neural Networks4
2023 Ultra-High Temporal Resolution Visual Reconstruction From a Fovea-Like Spike Camera via Spiking Neuron Model
abstract
Neuromorphic vision sensor is a new bio-inspired imaging paradigm emerged in recent years. It uses the asynchronous spike signals instead of the traditional frame-based manner to achieve ultra-high speed sampling. Unlike the dynamic vision sensor (DVS) that perceives movement by imitating the retinal periphery, the spike camera was developed recently to perceive fine textures by simulating a small retinal region called the fovea. For this new type of neuromorphic camera, how to reconstruct ultra-high speed visual images from spike data becomes an important yet challenging issue in visual scene perception, analysis, and recognition applications. In this paper, a bio-inspired visual reconstruction framework for the spike camera is proposed for the first time. Its core idea is to use the biologically inspired adaptive adjustment mechanisms, combined with the spatiotemporal spike information extracted by the proposed model, to reconstruct the full texture of natural scenes in an ultra-high temporal resolution. Specifically, the proposed model consists of a motion local excitation layer, a spike refining layer and a visual reconstruction layer motivated by the bio-realistic leaky integrate-and-fire (LIF) neurons and synapse connection with spike-timing dependent plasticity (STDP) rule. To evaluate the performance, a spike dataset was constructed for normal and high-speed scenes in real-world recorded by the spike camera. The experimental results show that the proposed approach can reconstruct the visual images with 40,000 frames per second in both normal and high-speed scenes, while achieving high dynamic range and high image quality.
Lin Zhu 0012, Siwei Dong, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 NeuroZoom: Denoising and Super Resolving Neuromorphic Events and Spikes
abstract
Neuromorphic cameras are emerging imaging technology that has advantages over conventional imaging sensors in several aspects including dynamic range, sensing latency, and power consumption. However, the signal-to-noise level and the spatial resolution still fall behind the state of conventional imaging sensors. In this article, we address the denoising and super-resolution problem for modern neuromorphic cameras. We employ 3D U-Net as the backbone neural architecture for such a task. The networks are trained and tested on two types of neuromorphic cameras: a dynamic vision sensor and a spike camera. Their pixels generate signals asynchronously, the former is based on perceived light changes and the latter is based on accumulated light intensity. To collect the datasets for training such networks, we design a display-camera system to record high frame-rate videos at multiple resolutions, providing supervision for denoising and super-resolution. The networks are trained in a noise-to-noise fashion, where the two ends of the network are unfiltered noisy data. The output of the networks has been tested for downstream applications including event-based visual object tracking and image reconstruction. Experimental results demonstrate the effectiveness of improving the quality of neuromorphic events and spikes, and the corresponding improvement to downstream applications with state-of-the-art performance.
Peiqi Duan 0002, Yi Ma 0001, Xinyu Shi 0004, Zihao W. Wang, Tiejun Huang 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Hybrid High Dynamic Range Imaging fusing Neuromorphic and Conventional Images
abstract
Reconstruction of high dynamic range image from a single low dynamic range image captured by a conventional RGB camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras like event camera and spike camera can record high dynamic range scenes in the form of intensity maps, but with much lower spatial resolution and no color information. In this article, we propose a hybrid imaging system (denoted as NeurImg) that captures and fuses the visual information from a neuromorphic camera and ordinary images from an RGB camera to reconstruct high-quality high dynamic range images and videos. The proposed NeurImg-HDR+ network consists of specially designed modules, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images to reconstruct high-resolution, high dynamic range images and videos. We capture a test dataset of hybrid signals on various HDR scenes using the hybrid camera, and analyze the advantages of the proposed fusing strategy by comparing it to state-of-the-art inverse tone mapping methods and merging two low dynamic range images approaches. Quantitative and qualitative experiments on both synthetic data and real-world scenarios demonstrate the effectiveness of the proposed hybrid high dynamic range imaging system. Code and dataset can be found at: https://github.com/hjynwa/NeurImg-HDR.
Jin Han 0001, Yixin Yang 0008, Peiqi Duan 0002, Chu Zhou, Lei Ma 0008, Chao Xu 0006, Tiejun Huang 0001, Imari Sato, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Capture the Moment: High-Speed Imaging With Spiking Cameras Through Short-Term Plasticity
abstract
High-speed imaging can help us understand some phenomena that are too fast to be captured by our eyes. Although ultra-fast frame-based cameras (e.g., Phantom) can record millions of fps at reduced resolution, they are too expensive to be widely used. Recently, a retina-inspired vision sensor, spiking camera, has been developed to record external information at 40, 000 Hz. The spiking camera uses the asynchronous binary spike streams to represent visual information. Despite this, how to reconstruct dynamic scenes from asynchronous spikes remains challenging. In this paper, we introduce novel high-speed image reconstruction models based on the short-term plasticity (STP) mechanism of the brain, termed TFSTP and TFMDSTP. We first derive the relationship between states of STP and spike patterns. Then, in TFSTP, by setting up the STP model at each pixel, the scene radiance can be inferred by the states of the models. In TFMDSTP, we use the STP to distinguish the moving and stationary regions, and then use two sets of STP models to reconstruct them respectively. In addition, we present a strategy for correcting error spikes. Experimental results show that the STP-based reconstruction methods can effectively reduce noise with less computing time, and achieve the best performances on both real-world and simulated datasets.
Yajing Zheng, Lingxiao Zheng, Zhaofei Yu, Tiejun Huang 0001, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Driver Emotion Recognition With a Hybrid Attentional Multimodal Fusion Framework
abstract
Negative emotions may induce dangerous driving behaviors leading to extremely serious traffic accidents. Therefore, it is necessary to establish a system that can automatically recognize driver emotions so that some actions can be taken to avoid traffic accidents. Existing studies on driver emotion recognition have mainly used facial data and physiological data. However, there are fewer studies on multimodal data with contextual characteristics of driving. In addition, fully fusing multimodal data in the feature fusion layer to improve the performance of emotion recognition is still a challenge. To this end, we propose to recognize driver emotion using a novel multimodal fusion framework based on convolutional long-short term memory network (ConvLSTM), and hybrid attention mechanism to fuse non-invasive multimodal data of eye, vehicle, and environment. In order to verify the effectiveness of the proposed method, extensive experiments have been carried out on a dataset collected using an advanced driving simulator. The experimental results demonstrate the effectiveness of the proposed method. Finally, a preliminary exploration on the correlation between driver emotion and stress is performed.
Luntian Mou, Yiyuan Zhao, Bahareh Nakisa, Mohammad Naim Rastgoo, Lei Ma 0008, Tiejun Huang 0001, Ramesh Jain 0001, Wen Gao 0001
IEEE Trans. Affect. Comput.7
2023 Learning Super-Resolution Reconstruction for High Temporal Resolution Spike Stream
abstract
Spike camera is a new type of bio-inspired vision sensor, each pixel of which perceives the brightness of the scene independently, and finally outputs 3-dimensional spatiotemporal spike streams. To bridge the spike camera and traditional frame-based vision, there is some works to reconstruct spike streams into regular images. However, the low spatial resolution ($400\times 250$) of the spike camera limits the quality of the reconstructed images. Thus, it is meaningful to explore a super-resolution reconstruction for spike streams. In this paper, we propose an end-to-end network to reconstruct high-resolution images from low-resolution spike streams. To utilize more spatiotemporal features of spike streams, our network adopts a multi-level features learning mechanism, including intra-stream feature extraction by spike encoder, inter-stream dependencies extraction based on optical flow module, and joint features learning via spike-based iterative projection. Experimental results demonstrate that our network is superior to the combination of state-of-the-art intensity image reconstruction methods and super-resolution networks on simulated and real datasets.
Xijie Xiang, Lin Zhu 0012, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Spike-Based Motion Estimation for Object Tracking Through Bio-Inspired Unsupervised Learning
abstract
Neuromorphic vision sensors, whose pixels output events/spikes asynchronously with a high temporal resolution according to the scene radiance change, are naturally appropriate for capturing high-speed motion in the scenes. However, how to utilize the events/spikes to smoothly track high-speed moving objects is still a challenging problem. Existing approaches either employ time-consuming iterative optimization, or require large amounts of labeled data to train the object detector. To this end, we propose a bio-inspired unsupervised learning framework, which takes advantage of the spatiotemporal information of events/spikes generated by neuromorphic vision sensors to capture the intrinsic motion patterns. Without off-line training, our models can filter the redundant signals with dynamic adaption module based on short-term plasticity, and extract the motion patterns with motion estimation module based on the spike-timing-dependent plasticity. Combined with the spatiotemporal and motion information of the filtered spike stream, the traditional DBSCAN clustering algorithm and Kalman filter can effectively track multiple targets in extreme scenes. We evaluate the proposed unsupervised framework for object detection and tracking tasks on synthetic data, publicly available event-based datasets, and spiking camera datasets. The experiment results show that the proposed model can robustly detect and smoothly track the moving targets on various challenging scenarios and outperforms state-of-the-art approaches.
Yajing Zheng, Zhaofei Yu, Song Wang 0002, Tiejun Huang 0001
IEEE Trans. Image Process.4
2023 A Residual Learning Approach to Deblur and Generate High Frame Rate Video With an Event Camera
abstract
Event cameras are bio-inspired cameras that can measure the intensity change asynchronously with high temporal resolution. One of the advantages of event cameras is that they suffer less from motion blur than traditional frame cameras when recording daily scenes with fast-moving objects. In this paper, we formulate the deblurring task on traditional cameras directed by events to be a residual learning one, and propose corresponding network architectures for effective learning of deblurring and high frame rate video generation tasks. We first train a modified U-Net network to restore a sharp image from a blurry image using the corresponding events. Then we train another similar network by replacing the downsampling blocks with blocks of the convolutional long short-term memory (Conv-LSTM) to recurrently generate high frame rate video using the restored sharp image and part of the events. Benefitting from the blur-free events and the proposed learning strategy, the experimental results show that the proposed method outperforms state-of-the-art methods for generating sharp images and high frame rate videos.
Minggui Teng, Boxin Shi, Yizhou Wang 0001, Tiejun Huang 0001
IEEE Trans. Multim.5
2023 Asynchronous Spatiotemporal Spike Metric for Event Cameras
abstract
Event cameras as bioinspired vision sensors have shown great advantages in high dynamic range and high temporal resolution in vision tasks. Asynchronous spikes from event cameras can be depicted using the marked spatiotemporal point processes (MSTPPs). However, how to measure the distance between asynchronous spikes in the MSTPPs still remains an open issue. To address this problem, we propose a general asynchronous spatiotemporal spike metric considering both spatiotemporal structural properties and polarity attributes for event cameras. Technically, the conditional probability density function is first introduced to describe the spatiotemporal distribution and polarity prior in the MSTPPs. Besides, a spatiotemporal Gaussian kernel is defined to capture the spatiotemporal structure, which transforms discrete spikes into the continuous function in a reproducing kernel Hilbert space (RKHS). Finally, the distance between asynchronous spikes can be quantified by the inner product in the RKHS. The experimental results demonstrate that the proposed approach outperforms the state-of-the-art methods and achieves significant improvement in computational efficiency. Especially, it is able to better depict the changes involving spatiotemporal structural properties and polarity attributes.
Jianing Li 0001, Yihua Fu, Siwei Dong, Zhaofei Yu, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Improving Multispike Learning With Plastic Synaptic Delays
abstract
Emulating the spike-based processing in the brain, spiking neural networks (SNNs) are developed and act as a promising candidate for the new generation of artificial neural networks that aim to produce efficient cognitions as the brain. Due to the complex dynamics and nonlinearity of SNNs, designing efficient learning algorithms has remained a major difficulty, which attracts great research attention. Most existing ones focus on the adjustment of synaptic weights. However, other components, such as synaptic delays, are found to be adaptive and important in modulating neural behavior. How could plasticity on different components cooperate to improve the learning of SNNs remains as an interesting question. Advancing our previous multispike learning, we propose a new joint weight-delay plasticity rule, named TDP-DL, in this article. Plastic delays are integrated into the learning framework, and as a result, the performance of multispike learning is significantly improved. Simulation results highlight the effectiveness and efficiency of our TDP-DL rule compared to baseline ones. Moreover, we reveal the underlying principle of how synaptic weights and delays cooperate with each other through a synthetic task of interval selectivity and show that plastic delays can enhance the selectivity and flexibility of neurons by shifting information across time. Due to this capability, useful information distributed away in the time domain can be effectively integrated for a better accuracy performance, as highlighted in our generalization tasks of the image, speech, and event-based object recognitions. Our work is thus valuable and significant to improve the performance of spike-based neuromorphic computing.
Qiang Yu 0005, Jialu Gao, Jianguo Wei, Kay Chen Tan, Tiejun Huang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 AMSA: Adaptive Multimodal Learning for Sentiment Analysis
abstract
Efficient recognition of emotions has attracted extensive research interest, which makes new applications in many fields possible, such as human-computer interaction, disease diagnosis, service robots, and so forth. Although existing work on sentiment analysis relying on sensors or unimodal methods performs well for simple contexts like business recommendation and facial expression recognition, it does far below expectations for complex scenes, such as sarcasm, disdain, and metaphors. In this article, we propose a novel two-stage multimodal learning framework, called AMSA, to adaptively learn correlation and complementarity between modalities for dynamic fusion, achieving more stable and precise sentiment analysis results. Specifically, a multiscale attention model with a slice positioning scheme is proposed to get stable quintuplets of sentiment in images, texts, and speeches in the first stage. Then a Transformer-based self-adaptive network is proposed to assign weights flexibly for multimodal fusion in the second stage and update the parameters of the loss function through compensation iteration. To quickly locate key areas for efficient affective computing, a patch-based selection scheme is proposed to iteratively remove redundant information through a novel loss function before fusion. Extensive experiments have been conducted on both machine weakly labeled and manually annotated datasets of self-made Video-SA, CMU-MOSEI, and CMU-MOSI. The results demonstrate the superiority of our approach through comparison with baselines.
Luntian Mou, Lei Ma 0008, Tiejun Huang 0001, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Optimized Potential Initialization for Low-Latency Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have been attached great importance due to the distinctive properties of low power consumption, biological plausibility, and adversarial robustness. The most effective way to train deep SNNs is through ANN-to-SNN conversion, which have yielded the best performance in deep network structure and large-scale datasets. However, there is a trade-off between accuracy and latency. In order to achieve high precision as original ANNs, a long simulation time is needed to match the firing rate of a spiking neuron with the activation value of an analog neuron, which impedes the practical application of SNN. In this paper, we aim to achieve high-performance converted SNNs with extremely low latency (fewer than 32 time-steps). We start by theoretically analyzing ANN-to-SNN conversion and show that scaling the thresholds does play a similar role as weight normalization. Instead of introducing constraints that facilitate ANN-to-SNN conversion at the cost of model capacity, we applied a more direct way by optimizing the initial membrane potential to reduce the conversion loss in each layer. Besides, we demonstrate that optimal initialization of membrane potentials can implement expected error-free ANN-to-SNN conversion. We evaluate our algorithm on the CIFAR-10 dataset and CIFAR-100 dataset and achieve state-of-the-art accuracy, using fewer time-steps. For example, we reach top-1 accuracy of 93.38% on CIFAR-10 with 16 time-steps. Moreover, our method can be applied to other ANN-SNN conversion methodologies and remarkably promote performance when the time-steps is small.
Tong Bu, Jianhao Ding, Zhaofei Yu, Tiejun Huang 0001
AAAI4
2022 Spatio-Temporal Recurrent Networks for Event-Based Optical Flow Estimation
abstract
Event camera has offered promising alternative for visual perception, especially in high speed and high dynamic range scenes. Recently, many deep learning methods have shown great success in providing model-free solutions to many event-based problems, such as optical flow estimation. However, existing deep learning methods did not address the importance of temporal information well from the perspective of architecture design and cannot effectively extract spatio-temporal features. Another line of research that utilizes Spiking Neural Network suffers from training issues for deeper architecture. To address these points, a novel input representation is proposed that captures the events temporal distribution for signal enhancement. Moreover, we introduce a spatio-temporal recurrent encoding-decoding neural network architecture for event-based optical flow estimation, which utilizes Convolutional Gated Recurrent Units to extract feature maps from a series of event images. Besides, our architecture allows some traditional frame-based core modules, such as correlation layer and iterative residual refine scheme, to be incorporated. The network is end-to-end trained with self-supervised learning on the Multi-Vehicle Stereo Event Camera dataset. We have shown that it outperforms all the existing state-of-the-art methods by a large margin.
Ziluo Ding, Rui Zhao 0010, Jiyuan Zhang 0005, Tianxiao Gao, Ruiqin Xiong, Zhaofei Yu, Tiejun Huang 0001
AAAI7
2022 Retinomorphic Object Detection in Asynchronous Visual Streams
abstract
Due to high-speed motion blur and challenging illumination, conventional frame-based cameras have encountered an important challenge in object detection tasks. Neuromorphic cameras that output asynchronous visual streams instead of intensity frames, by taking the advantage of high temporal resolution and high dynamic range, have brought a new perspective to address the challenge. In this paper, we propose a novel problem setting, retinomorphic object detection, which is the first trial that integrates foveal-like and peripheral-like visual streams. Technically, we first build a large-scale multimodal neuromorphic object detection dataset (i.e., PKU-Vidar-DVS) over 215.5k spatio-temporal synchronized labels. Then, we design temporal aggregation representations to preserve the spatio-temporal information from asynchronous visual streams. Finally, we present a novel bio-inspired unifying framework to fuse two sensing modalities via a dynamic interaction mechanism. Our experimental evaluation shows that our approach has significant improvements over the state-of-the-art methods with the single-modality, especially in high-speed motion and low-light scenarios. We hope that our work will attract further research into this newly identified, yet crucial research direction. Our dataset can be available at https://www.pkuml.org/resources/pku-vidar-dvs.html.
Jianing Li 0001, Xiao Wang 0014, Lin Zhu 0012, Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001
AAAI5
2022 Event-based Video Reconstruction via Potential-assisted Spiking Neural Network
abstract
Neuromorphic vision sensor is a new bio-inspired imaging paradigm that reports asynchronous, continuously perpixel brightness changes called ‘events’ with high temporal resolution and high dynamic range. So far, the event-based image reconstruction methods are based on artificial neural networks (ANN) or hand-crafted spatiotemporal smoothing techniques. In this paper, we first implement the image reconstruction work via deep spiking neural network (SNN) architecture. As the bio-inspired neural networks, SNNs operating with asynchronous binary spikes distributed over time, can potentially lead to greater computational efficiency on event-driven hardware. We propose a novel Event-based Video reconstruction framework based on a fully Spiking Neural Network (EVSNN), which utilizes Leaky-Integrate-and-Fire (LIF) neuron and Membrane Potential (MP) neuron. We find that the spiking neurons have the potential to store useful temporal information (memory) to complete such time-dependent tasks. Further-more, to better utilize the temporal information, we propose a hybrid potential-assisted framework (PAEVSNN) using the membrane potential of spiking neuron. The proposed neuron is referred as Adaptive Membrane Potential (AMP) neuron, which adaptively updates the membrane potential according to the input spikes. The experimental results demonstrate that our models achieve comparable performance to ANN-based models on IJRR, MVSEC, and HQF datasets. The energy consumptions of EVSNN and PAEVSNN are$19.36\times$and$7.75\times$more computationally ef-ficient than their ANN architectures, respectively. The code and pretrained model are available at https://sites.google.com/view/evsnn.
Lin Zhu 0012, Xiao Wang 0014, Yi Chang 0002, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001
CVPR5
2022 Optical Flow Estimation for Spiking Camera
abstract
As a bio-inspired sensor with high temporal resolution, the spiking camera has an enormous potential in real applications, especially for motion estimation in high-speed scenes. However, frame-based and event-based methods are not well suited to spike streams from the spiking camera due to the different data modalities. To this end, we present, SCFlow, a tailored deep learning pipeline to estimate optical flow in high-speed scenes from spike streams. Importantly, a novel input representation is introduced which can adaptively remove the motion blur in spike streams according to the prior motion. Further, for training SCFlow, we synthesize two sets of optical flow data for the spiking camera, SPIkingly Flying Things and Photo-realistic Highspeed Motion, denoted as SPIFT and PHM respectively, corresponding to random high-speed and well-designed scenes. Experimental results show that the SCFlow can predict optical flow from spike streams in different high-speed scenes. Moreover, SCFlow shows promising generalization on real spike streams. Codes and datasets refer to https://github.com/Acnext/Optical-Flow-For-Spiking-Camera.
Liwen Hu 0002, Rui Zhao 0010, Ziluo Ding, Lei Ma 0008, Boxin Shi, Ruiqin Xiong, Tiejun Huang 0001
CVPR7
2022 Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
abstract
We present Point-BERT, a new paradigm for learning Transformers to generalize the concept of BERT [8] to 3D point cloud. Inspired by BERT, we devise a Masked Point Modeling (MPM) task to pre-train point cloud Transformers. Specifically, we first divide a point cloud into several local point patches, and a point cloud Tokenizer with a discrete Variational AutoEncoder (dVAE) is designed to generate discrete point tokens containing meaningful local information. Then, we randomly mask out some patches of input point clouds and feed them into the backbone Transformers. The pre-training objective is to recover the original point tokens at the masked locations under the supervision of point tokens obtained by the Tokenizer. Extensive experiments demonstrate that the proposed BERT-style pre-training strategy significantly improves the performance of standard point cloud Transformers. Equipped with our pre-training strategy, we show that a pure Transformer architecture attains 93.8% accuracy on ModelNet40 and 83.1% accuracy on the hardest setting of ScanObjectNN, surpassing carefully designed point cloud models with much fewer hand-made designs. We also demonstrate that the representations learned by Point-BERT transfer well to new tasks and domains, where our models largely advance the state-of-the-art of few-shot point cloud classification task. The code and pre-trained models are available at https://github.com/lulutang0608/Point-BERT.
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 0001, Jie Zhou 0001, Jiwen Lu
CVPR4
2022 2D Amodal Instance Segmentation Guided by 3D Shape Prior
Weining Ye, Tingting Jiang 0001, Tiejun Huang 0001
ECCV (29)4
2022 Spike Transformer: Monocular Depth Estimation for Spiking Camera
Jiyuan Zhang 0005, Lulu Tang, Zhaofei Yu, Jiwen Lu, Tiejun Huang 0001
ECCV (7)5
2022 Modeling The Detection Capability Of High-Speed Spiking Cameras
abstract
The novel working principle enables spiking cameras to capture high-speed moving objects. However, the applications of spiking cameras can be affected by many factors, such as brightness intensity, detectable distance, and the maximum speed of moving targets. Improper settings such as weak ambient brightness and too short object-camera distance, will lead to failure in the application of such cameras. To address the issue, this paper proposes a modeling algorithm that studies the detection capability of spiking cameras. The algorithm deduces the maximum detectable speed of spiking cameras corresponding to different scenario settings (e.g., brightness intensity, camera lens, and object-camera distance) based on the basic technical parameters of cameras (e.g., pixel size, spatial and temporal resolution). Thereby, the proper camera settings for various applications can be determined. Extensive experiments verify the effectiveness of the modeling algorithm. To our best knowledge, it is the first work to investigate the detection capability of spiking cameras.
Junwei Zhao 0003, Zhaofei Yu, Lei Ma 0008, Ziluo Ding, Shiliang Zhang, Yonghong Tian 0001, Tiejun Huang 0001
ICASSP7
2022 Transformer-Based Domain Adaptation for Event Data Classification
abstract
Event cameras encode the change of brightness into events, differing from conventional frame cameras. The novel working principle makes them to have stronger potential in high-speed applications. However, the lack of labeled event annotations limits the applications of such cameras in deep learning frameworks, making it appealing to study more efficient deep learning algorithms and architectures. This paper devises the Convolutional Transformer Network (CTN) for processing event data. The CTN enjoys the advantages of convolution networks and transformers, presenting stronger capability in event-based classification tasks compared with existing models. To address the insufficiency issue of annotated event data, we propose to train the CTN via the source-free Unsupervised Domain Adaptation (UDA) algorithm leveraging large-scale labeled image data. Extensive experiments verify the effectiveness of the UDA algorithm. And our CTN outperforms recent state-of-the-art methods on event-based classification tasks, suggesting that it is an effective model for this task. To our best acknowledge, it is an early attempt of employing vision transformers with the source-free UDA algorithm to process event data.
Junwei Zhao 0003, Shiliang Zhang, Tiejun Huang 0001
ICASSP3
2022 3D Residual Interpolation for Spike Camera Demosaicing
abstract
The recently invented spike camera can capture high-speed motion in dynamic scenes by accumulating incoming photons continuously and firing spikes at very high temporal resolution. This paper addresses the demosaicing problem in spike camera color imaging. Specifically, we propose the 3D residual interpolation (3DRI) method to convert raw spike frames to color image frames. Due to the Poisson effect of photon arrivals and the quantization effect of spike readout, the instantaneous intensity recovered from the spike stream may suffer from undesired noise. To handle the noise, we estimate the missing color pixels along motion trajectories to exploit the temporal correlation among neighboring frames. In addition, by utilizing the color channels correlation, we design a residual-based demosaicing pipeline that uses the green pixels to guide the estimation of the red or blue missing pixels. Experimental results demonstrate our proposed 3DRI can produce color images from spike streams, achieving a good objective and perceptual quality for high-motion scenes.
Yanchen Dong 0001, Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001
ICIP4
2022 Optimal ANN-SNN Conversion for High-accuracy and Ultra-low-latency Spiking Neural Networks
Tong Bu, Wei Fang 0006, Jianhao Ding, Penglin Dai, Zhaofei Yu, Tiejun Huang 0001
ICLR6
2022 Learning Stereo Depth Estimation with Bio-Inspired Spike Cameras
abstract
Bio-inspired spike cameras, offering high temporal resolution spike streams, have brought a new perspective to address common challenges (e.g.,high-speed motion blur) in depth estimation tasks. In this paper, we propose a novel problem setting, spike-based stereo depth estimation, which is the first trail that explores an end-to-end network to learn stereo depth estimation with transformers for spike cameras, named Spike-based Stereo Depth Estimation Transformer (SSDEFormer). We first build a hybrid camera platform and provide a new stereo depth estimation dataset (i.e.,PKU-Spike-Stereo) with spatiotemporal synchronized labels. Then, we propose a novel spike representation to effectively exploit spatiotemporal information from spike streams. Finally, a transformer-based network is designed to generate dense depth maps without a fixed-disparity cost volume. Empirically, it shows that our approach is extremely effective on both synthetic and real-world datasets. The results verify that spike cameras can perform robust depth estimation even in cases where conventional cameras and event cameras fail in fast motion scenarios.
Jianing Li 0001, Lin Zhu 0012, Xijie Xiang, Tiejun Huang 0001, Yonghong Tian 0001
ICME5
2022 Temporal Up-Sampling for Asynchronous Events
abstract
The event camera is a novel bio-inspired vision sensor. When the brightness change exceeds the preset threshold, the sensor generates events asynchronously. The number of valid events directly affects the performance of event-based tasks, such as reconstruction, detection, and recognition. However, when in low-brightness or slow-moving scenes, events are often sparse and accompanied by noise, which poses challenges for event-based tasks. To solve these challenges, we propose an event temporal up-sampling algorithm11Code: https://github.com/XIJIE-XIANG/Event-Temporal-Up-sampling to generate more effective and reliable events. The main idea of our algorithm is to generate up-sampling events on the event motion trajectory. First, we estimate the event motion trajectory by contrast maximization algorithm and then up-sampling the events by temporal point processes. Experimental results show that up-sampling events can provide more effective information and improve the performance of downstream tasks, such as improving the quality of reconstructed images and increasing the accuracy of object detection.
Xijie Xiang, Lin Zhu 0012, Jianing Li 0001, Yonghong Tian 0001, Tiejun Huang 0001
ICME5
2022 State Transition of Dendritic Spines Improves Learning of Sparse Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are considered a promising alternative to Artificial Neural Networks (ANNs) for their event-driven computing paradigm when deployed on energy-efficient neuromorphic hardware. Recently, deep SNNs have shown breathtaking performance improvement through cutting-edge training strategy and flexible structure, which also scales up the number of parameters and computational burdens in a single network. Inspired by the state transition of dendritic spines in the filopodial model of spinogenesis, we model different states of SNN weights, facilitating weight optimization for pruning. Furthermore, the pruning speed can be regulated by using different functions describing the growing threshold of state transition. We organize these techniques as a dynamic pruning algorithm based on nonlinear reparameterization mapping from spine size to SNN weights. Our approach yields sparse deep networks on the large-scale dataset (SEW ResNet18 on ImageNet) while maintaining state-of-the-art low performance loss ( 3% at 88.8% sparsity) compared to existing pruning methods on directly trained SNNs. Moreover, we find out pruning speed regulation while learning is crucial to avoiding disastrous performance degradation at the final stages of training, which may shed light on future work on SNN pruning.
Yanqi Chen, Zhaofei Yu, Wei Fang 0006, Zhengyu Ma, Tiejun Huang 0001, Yonghong Tian 0001
ICML5
2022 Self-Supervised Mutual Learning for Dynamic Scene Reconstruction of Spiking Camera
abstract
Mimicking the sampling mechanism of the primate fovea, a retina-inspired vision sensor named spiking camera has been developed, which has shown great potential for capturing high-speed dynamic scenes with a sampling rate of 40,000 Hz. Unlike conventional digital cameras, the spiking camera continuously captures photons and outputs asynchronous binary spikes with various inter-spike intervals to record dynamic scenes. However, how to reconstruct dynamic scenes from asynchronous spike streams remains challenging. In this work, we propose a novel pretext task to build a self-supervised reconstruction framework for spiking cameras. Specifically, we utilize the blind-spot network commonly used in self-supervised denoising tasks as our backbone, and perform self-supervised learning by constructing proper pseudo-labels. In addition, in view of the poor scalability and insufficient information utilization of the blind-spot network, we present a mutual learning framework to improve the overall performance of the network through mutual distillation between a non-blind-spot network and a blind-spot network. This also enables the network to bypass constraints of the blind-spot network, allowing state-of-the-art modules to be used to further improve performance. The experimental results demonstrate that our methods evidently outperform previous unsupervised spiking camera reconstruction methods and achieve desirable results compared with supervised methods.
Shiyan Chen, Chaoteng Duan, Zhaofei Yu, Ruiqin Xiong, Tiejun Huang 0001
IJCAI5
2022 SpikingSIM: A Bio-Inspired Spiking Simulator
abstract
Large-scale neuromorphic dataset is costly to construct and difficult to annotate because of the unique high-speed asynchronous imaging principle of bio-inspired cameras. Lacking of large-scale annotated neuromorphic datasets has significantly hindered the applications of bio-inspired cameras in deep neural networks. Synthesizing neuromorphic data from annotated RGB images can be considered to alleviate this challenge. This paper proposes a simulator to generate simulated spiking data from images recorded by frame cameras. To minimize the deviations between synthetic data and real data, the proposed simulator named SpikingSIM considers the sensing principle of spiking cameras, and generates high-quality simulated spiking data, e.g., the noises in real data are also simulated. Experimental results show that, our simulator generates more realistic spiking data than existing methods. We hence train deep neural networks with synthesized spiking data. Experiments show that, the net- work trained by our simulated data generalizes well on real spiking data. The source code of SpikingSIM is available at http://github.com/Evin-X/SpikingSIM.
Junwei Zhao 0003, Shiliang Zhang, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001
ISCAS5
2022 Learning Optical Flow from Continuous Spike Streams
abstract
Spike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting continuous binary spike streams. Optical flow is a key task for spike cameras and their applications. A previous attempt has been made for spike-based optical flow. However, the previous work only focuses on motion between two moments, and it uses graphics-based data for training, whose generalization is limited. In this paper, we propose a tailored network, Spike2Flow that extracts information from binary spikes with temporal-spatial representation based on the differential of spike firing time and spatial information aggregation. The network utilizes continuous motion clues through joint correlation decoding. Besides, a new dataset with real-world scenes is proposed for better generalization. Experimental results show that our approach achieves state-of-the-art performance on existing synthetic datasets and real data captured by spike cameras. The source code and dataset are available at \url{https://github.com/ruizhao26/Spike2Flow}.
Rui Zhao 0010, Ruiqin Xiong, Jing Zhao 0011, Zhaofei Yu, Xiaopeng Fan 0001, Tiejun Huang 0001
NeurIPS6
2022 Oscillatory Tracking of Continuous Attractor Neural Networks Account for Phase Precession and Procession of Hippocampal Place Cells
abstract
Hippocampal place cells of freely moving rodents display an intriguing temporal organization in their responses known as `theta phase precession', in which individual neurons fire at progressively earlier phases in successive theta cycles as the animal traverses the place fields. Recent experimental studies found that in addition to phase precession, many place cells also exhibit accompanied phase procession, but the underlying neural mechanism remains unclear. Here, we propose a neural circuit model to elucidate the generation of both kinds of phase shift in place cells' firing. Specifically, we consider a continuous attractor neural network (CANN) with feedback inhibition, which is inspired by the reciprocal interaction between the hippocampus and the medial septum. The feedback inhibition induces intrinsic mobility of the CANN which competes with the extrinsic mobility arising from the external drive. Their interplay generates an oscillatory tracking state, that is, the network bump state (resembling the decoded virtual position of the animal) sweeps back and forth around the external moving input (resembling the physical position of the animal). We show that this oscillatory tracking naturally explains the forward and backward sweeps of the decoded position during the animal's locomotion. At the single neuron level, the forward and backward sweeps account for, respectively, theta phase precession and procession. Furthermore, by tuning the feedback inhibition strength, we also explain the emergence of bimodal cells and unimodal cells, with the former having co-existed phase precession and procession, and the latter having only significant phase precession. We hope that this study facilitates our understanding of hippocampal temporal coding and lays foundation for unveiling their computational functions.
Tianhao Chu, Zilong Ji, Junfeng Zuo, Wenhao Zhang 0002, Tiejun Huang 0001, Yuanyuan Mi, Si Wu 0001
NeurIPS5
2022 SNN-RAT: Robustness-enhanced Spiking Neural Network through Regularized Adversarial Training
abstract
Spiking neural networks (SNNs) are promising to be widely deployed in real-time and safety-critical applications with the advance of neuromorphic computing. Recent work has demonstrated the insensitivity of SNNs to small random perturbations due to the discrete internal information representation. The variety of training algorithms and the involvement of the temporal dimension pose more threats to the robustness of SNNs than that of typical neural networks. We account for the vulnerability of SNNs by constructing adversaries based on different differentiable approximation techniques. By deriving a Lipschitz constant specifically for the spike representation, we first theoretically answer the question of how much adversarial invulnerability is retained in SNNs. Hence, to defend against the broad attack methods, we propose a regularized adversarial training scheme with low computational overheads. SNNs can benefit from the constraint of the perturbed spike distance's amplification and the generalization on multiple adversarial $\epsilon$-neighbourhoods. Our experiments on the image recognition benchmarks have proven that our training scheme can defend against powerful adversarial attacks crafted from strong differentiable approximations. To be specific, our approach makes the black-box attacks of the Projected Gradient Descent attack nearly ineffective. We believe that our work will facilitate the spread of SNNs for safety-critical applications and help understand the robustness of the human brain.
Jianhao Ding, Tong Bu, Zhaofei Yu, Tiejun Huang 0001, Jian K. Liu
NeurIPS4
2022 Adaptation Accelerating Sampling-based Bayesian Inference in Attractor Neural Networks
abstract
The brain performs probabilistic Bayesian inference to interpret the external world. The sampling-based view assumes that the brain represents the stimulus posterior distribution via samples of stochastic neuronal responses. Although the idea of sampling-based inference is appealing, it faces a critical challenge of whether stochastic sampling is fast enough to match the rapid computation of the brain. In this study, we explore how latent stimulus sampling can be accelerated in neural circuits. Specifically, we consider a canonical neural circuit model called continuous attractor neural networks (CANNs) and investigate how sampling-based inference of latent continuous variables is accelerated in CANNs. Intriguingly, we find that by including noisy adaptation in the neuronal dynamics, the CANN is able to speed up the sampling process significantly. We theoretically derive that the CANN with noisy adaptation implements the efficient sampling method called Hamiltonian dynamics with friction, where noisy adaption effectively plays the role of momentum. We theoretically analyze the sampling performances of the network and derive the condition when the acceleration has the maximum effect. Simulation results confirm our theoretical analyses. We further extend the model to coupled CANNs and demonstrate that noisy adaptation accelerates the sampling of the posterior distribution of multivariate stimuli. We hope that this study enhances our understanding of how Bayesian inference is realized in the brain.
Xingsi Dong, Zilong Ji, Tianhao Chu, Tiejun Huang 0001, Wenhao Zhang 0002, Si Wu 0001
NeurIPS4
2022 Temporal Effective Batch Normalization in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are promising in neuromorphic hardware owing to utilizing spatio-temporal information and sparse event-driven signal processing. However, it is challenging to train SNNs due to the non-differentiable nature of the binary firing function. The surrogate gradients alleviate the training problem and make SNNs obtain comparable performance as Artificial Neural Networks (ANNs) with the same structure. Unfortunately, batch normalization, contributing to the success of ANNs, does not play a prominent role in SNNs because of the additional temporal dimension. To this end, we propose an effective normalization method called temporal effective batch normalization (TEBN). By rescaling the presynaptic inputs with different weights at every time-step, temporal distributions become smoother and uniform. Theoretical analysis shows that TEBN can be viewed as a smoother of SNN's optimization landscape and could help stabilize the gradient norm. Experimental results on both static and neuromorphic datasets show that SNNs with TEBN outperform the state-of-the-art accuracy with fewer time-steps, and achieve better robustness to hyper-parameters than other normalizations.
Chaoteng Duan, Jianhao Ding, Shiyan Chen, Zhaofei Yu, Tiejun Huang 0001
NeurIPS5
2022 Training Spiking Neural Networks with Event-driven Backpropagation
abstract
Spiking Neural networks (SNNs) represent and transmit information by spatiotemporal spike patterns, which bring two major advantages: biological plausibility and suitability for ultralow-power neuromorphic implementation. Despite this, the binary firing characteristic makes training SNNs more challenging. To learn the parameters of deep SNNs in an event-driven fashion as in inference of SNNs, backpropagation with respect to spike timing is proposed. Although this event-driven learning has the advantages of lower computational cost and memory occupation, the accuracy is far below the recurrent neural network-like learning approaches. In this paper, we first analyze the commonly used temporal backpropagation training approach and prove that the sum of gradients remains unchanged between fully-connected and convolutional layers. Secondly, we show that the max pooling layer meets the above invariance rule, while the average pooling layer does not, which will suffer the gradient vanishing problem but can be revised to meet the requirement. Thirdly, we point out the reverse gradient problem for time-based gradients and propose a backward kernel that can solve this problem and keep the property of the invariable sum of gradients. The experimental results show that the proposed approach achieves state-of-the-art performance on CIFAR10 among time-based training methods. Also, this is the first time that the time-based backpropagation approach successfully trains SNN on the CIFAR100 dataset. Our code is available at https://github.com/zhuyaoyu/SNN-event-driven-learning.
Yaoyu Zhu, Zhaofei Yu, Wei Fang 0006, Tiejun Huang 0001, Timothée Masquelier
NeurIPS5
2022 High-Speed Scene Reconstruction from Low-Light Spike Streams
abstract
Benefiting from the high temporal resolution, the spike camera shows promising potential in capturing high-speed scenes via accumulating luminance intensity and firing spikes. Although the spike camera compared to the high-speed camera is quite cost-effective, its performance in capturing low-light scenes is poor. Specifically, it takes more time for the spike camera to accumulate enough light signal for firing a spike in low-light scenes, while the scenes may have already changed because of the high-speed motion. There may be no effective spikes for a long time due to the insufficient incident light, and the signal-to-noise ratio of spike streams in low-light scenes is unsatisfactory. Thus, it's easy to introduce noise and motion blur while reconstructing, especially in rapidly changing scenes. To address this issue, we propose a low-light scene reconstruction method for the spike camera. In particular, we first develop a Brightness-Adaptive Light Inference (BALI) method to preliminarily reconstruct the low-light scene according to the brightness, which utilizes both the spike interval and the spike number. Considering the motion, we then estimate optical flow and filter the preliminary restored frames iteratively to handle the noise via temporal correlation. After that, there is still some noise, and we further handle it by a spatial filter according to the brightness. As a result, we restore a clear image based on both temporal and spatial correlation. The experimental results demonstrate that our method achieves good visual quality in low-light scene reconstruction.
Yanchen Dong 0001, Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001
VCIP4
2022 Spike Signal Reconstruction Based on Inter-Spike Similarity
abstract
Spike camera is a kind of bio-inspired camera which is particularly proposed for capturing dynamic scenes with high speed motion. Spike camera works in a way simulating the retina that it receives incoming photons continuously and fires a spike whenever the accumulated photons reach a threshold. The spike stream can be recorded at an extremely high temporal resolution so that the dynamic process of light-intensity changes may be recovered. This paper addresses the problem of recovering the original visual signal from spikes. To reduce the fluctuation in spike intervals caused by the Poisson effect of photon arrivals and the quantization effect in spike reading, we estimate the real interval from a sequence of temporally neighboring spikes. To avoid mixing the spikes generated from significantly different light intensities, we propose a temporal and spatial weighting method based on the inter-spike similarity. Experimental results demonstrate that the proposed method outperforms the previous light intensity inference methods and achieves better performance under different motion conditions.
Ruiqin Xiong, Tiejun Huang 0001
VCIP3
2022 Evolution of AVS video coding standards: twenty years of innovation and development
Siwei Ma 0001, Li Zhang 0006, Shiqi Wang 0001, Chuanmin Jia, Shanshe Wang, Tiejun Huang 0001, Feng Wu 0001, Wen Gao 0001
Sci. China Inf. Sci.6
2022 Decoding Pixel-Level Image Features From Two-Photon Calcium Signals of Macaque Visual Cortex
abstract
Images of visual scenes comprise essential features important for visual cognition of the brain. The complexity of visual features lies at different levels, from simple artificial patterns to natural images with different scenes. It has been a focus of using stimulus images to predict neural responses. However, it remains unclear how to extract features from neuronal responses. Here we address this question by leveraging two-photon calcium neural data recorded from the visual cortex of awake macaque monkeys. With stimuli including various categories of artificial patterns and diverse scenes of natural images, we employed a deep neural network decoder inspired by image segmentation technique. Consistent with the notation of sparse coding for natural images, a few neurons with stronger responses dominated the decoding performance, whereas decoding of ar tificial patterns needs a large number of neurons. When natural images using the model pretrained on artificial patterns are decoded, salient features of natural scenes can be extracted, as well as the conventional category information. Altogether, our results give a new perspective on studying neural encoding principles using reverse-engineering decoding strategies.
Yijun Zhang 0003, Tong Bu, Jiyuan Zhang 0005, Shiming Tang, Zhaofei Yu, Jian K. Liu, Tiejun Huang 0001
Neural Comput.7
2022 An FPGA Accelerator for High-Speed Moving Objects Detection and Tracking With a Spike Camera
abstract
Ultra-high-speed object detection and tracking are crucial in fields such as fault detection and scientific observation. Existing solutions to this task have deficiencies in processing speeds. To deal with this difficulty, we propose a neural-inspired ultra-high-speed moving object filtering, detection, and tracking scheme, as well as a corresponding accelerator based on a high-speed spike camera. We parallelize the filtering module and divide the detection module to accelerate the algorithm and balance latency among modules for the benefit of the task-level pipeline. To be specific, a block-based parallel computation model is proposed to accelerate the filtering module, and the detection module is accelerated by a parallel connected component labeling algorithm modeling spike sparsity and spatial connectivity of moving objects with a searching tree. The hardware optimizations include processing the LIF layer with a group of multiplexers to reduce ADD operations and replacing expensive exponential operations with multiplications of preprocessed fixed-point values to increase processing speed and minimize resource consumption. We design an accelerator with the above techniques, achieving 19 times acceleration over the serial version after 25-way parallelization. A processing system for the accelerator is also implemented on the Xilinx ZCU-102 board to validate its functionality and performance. Our accelerator can process more than 20,000 spike images with 250 × 400 resolution per second with 1.618 W dynamic power consumption.
Yaoyu Zhu, Tiejun Huang 0001
Neural Comput.4
2022 Neural feedback facilitates rough-to-fine information retrieval
abstract
Categorical relationships between objects are encoded as overlapped neural representations in the brain, where the more similar the objects are, the larger the correlations between their evoked neuronal responses. These representation correlations, however, inevitably incur interference when memories are retrieved. Here, we propose that neural feedback, which is widely observed in the brain but whose function remains largely unknown, contributes to disentangle neural correlations to improve information retrieval. We study a hierarchical neural network storing the hierarchical categorical information of objects, and information retrieval goes from rough-to-fine, aided by the push-pull neural feedback. We elucidate that the push and the pull components of the feedback suppress the interferences due to the representation correlations between objects from different and the same categories, respectively. Our model reproduces the push-pull phenomenon observed in neural data and sheds light on our understanding of the role of feedback in neural information processing.
Xiaolong Zou, Zilong Ji, Gengshuo Tian, Yuanyuan Mi, Tiejun Huang 0001, K. Y. Michael Wong, Si Wu 0001
Neural Networks6
2022 Guided Event Filtering: Synergy Between Intensity Images and Neuromorphic Events for High Performance Imaging
abstract
Many visual and robotics tasks in real-world scenarios rely on robust handling of high speed motion and high dynamic range (HDR) with effectively high spatial resolution and low noise. Such stringent requirements, however, cannot be directly satisfied by a single imager or imaging modality, rather by multi-modal sensors with complementary advantages. In this paper, we address high performance imaging by exploring the synergy between traditional frame-based sensors with high spatial resolution and low sensor noise, and emerging event-based sensors with high speed and high dynamic range. We introduce a novel computational framework, termed Guided Event Filtering (GEF), to process these two streams of input data and output a stream of super-resolved yet noise-reduced events. To generate high quality events, GEF first registers the captured noisy events onto the guidance image plane according to our flow model. it then performs joint image filtering that inherits the mutual structure from both inputs. Lastly, GEF re-distributes the filtered event frame in the space-time volume while preserving the statistical characteristics of the original events. When the guidance images under-perform, GEF incorporates an event self-guiding mechanism that resorts to neighbor events for guidance. We demonstrate the benefits of GEF by applying the output high quality events to existing event-based algorithms across diverse application categories, including high speed object tracking, depth estimation, high frame-rate video synthesis, and super resolution/HDR/color image restoration.
Peiqi Duan 0002, Zihao W. Wang, Boxin Shi, Oliver Cossairt, Tiejun Huang 0001, Aggelos K. Katsaggelos
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 BDCN: Bi-Directional Cascade Network for Perceptual Edge Detection
abstract
Exploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a bi-directional cascade network (BDCN) architecture, where an individual layer is supervised by labeled edges at its specific scale, rather than directly applying the same supervision to different layers. Furthermore, to enrich multi-scale representations learned by each layer of BDCN, we introduce a scale enhancement module (SEM), which utilizes dilated convolution to generate multi-scale features, instead of using deeper CNNs. These new approaches encourage the learning of multi-scale representations in different layers and detect edges that are well delineated by their scales. Learning scale dedicated layers also results in a compact network with a fraction of parameters. We evaluate our method on three datasets, i.e., BSDS500, NYUDv2, and Multicue, and achieve ODS F-measure of 0.832, 2.7 percent higher than current state-of-the-art on the BSDS500 dataset. We also applied our edge detection result to other vision tasks. Experimental results show that, our method further boosts the performance of image segmentation, optical flow estimation, and object proposal generation.
Shiliang Zhang, Ming Yang 0007, Yanhu Shan, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 MRDFlow: Unsupervised Optical Flow Estimation Network With Multi-Scale Recurrent Decoder
abstract
Optical flow estimation is a fundamental task in computer vision and image processing. Due to the difficulty in obtaining the ground truth of flow field, unsupervised learning approaches attract more and more research interests in recent years. However, despite of their good generalization capability, unsupervised optical flow methods suffer in the scenarios with large displacement, small objects, and occlusions. In this work, we propose a novel optical flow network based on decoder with multi-scale kernels. Different from previous U-Net like or pyramidal methods, we design our network based on RAFT architecture that with a 4D correlation layer and recurrent decoder. More importantly, we incorporate three novel ideas with regard to the input, information processing and output of the update units improve the performance. Firstly, we utilize various motion-related information as input to the update units. Secondly, we propose a module of multi-scale update unit. Thirdly, for the final flow up-sampling procedure, we propose an image-guided up-sampling loss to guide the learning of up-sampling masks. Our model is trained by the occlusion-aware photometric loss, edge-aware smoothness loss, self-supervised loss, and image-guided up-sampling loss. Experimental results demonstrate that our model achieves the state-of-the-art performance on both Sintel and KITTI and outperforms other unsupervised optical flow methods remarkably.
Rui Zhao 0010, Ruiqin Xiong, Ziluo Ding, Xiaopeng Fan 0001, Jian Zhang 0018, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2022 Neural System Identification With Spike-Triggered Non-Negative Matrix Factorization
abstract
Neuronal circuits formed in the brain are complex with intricate connection patterns. Such complexity is also observed in the retina with a relatively simple neuronal circuit. A retinal ganglion cell (GC) receives excitatory inputs from neurons in previous layers as driving forces to fire spikes. Analytical methods are required to decipher these components in a systematic manner. Recently a method called spike-triggered non-negative matrix factorization (STNMF) has been proposed for this purpose. In this study, we extend the scope of the STNMF method. By using retinal GCs as a model system, we show that STNMF can detect various computational properties of upstream bipolar cells (BCs), including spatial receptive field, temporal filter, and transfer nonlinearity. In addition, we recover synaptic connection strengths from the weight matrix of STNMF. Furthermore, we show that STNMF can separate spikes of a GC into a few subsets of spikes, where each subset is contributed by one presynaptic BC. Taken together, these results corroborate that STNMF is a useful method for deciphering the structure of neuronal circuits.
Shanshan Jia 0001, Zhaofei Yu, Arno Onken, Yonghong Tian 0001, Tiejun Huang 0001, Jian K. Liu
IEEE Trans. Cybern.5
2022 Revealing Fine Structures of the Retinal Receptive Field by Deep-Learning Networks
abstract
Deep convolutional neural networks (CNNs) have demonstrated impressive performance on many visual tasks. Recently, they became useful models for the visual system in neuroscience. However, it is still not clear what is learned by CNNs in terms of neuronal circuits. When a deep CNN with many layers is used for the visual system, it is not easy to compare the structure components of CNNs with possible neuroscience underpinnings due to highly complex circuits from the retina to the higher visual cortex. Here, we address this issue by focusing on single retinal ganglion cells with biophysical models and recording data from animals. By training CNNs with white noise images to predict neuronal responses, we found that fine structures of the retinal receptive field can be revealed. Specifically, convolutional filters learned are resembling biological components of the retinal circuit. This suggests that a CNN learning from one single retinal cell reveals a minimal neural network carried out in this cell. Furthermore, when CNNs learned from different cells are transferred between cells, there is a diversity of transfer learning performance, which indicates that CNNs are cell specific. Moreover, when CNNs are transferred between different types of input images, here white noise versus natural images, transfer learning shows a good performance, which implies that CNNs indeed capture the full computational ability of a single retinal cell for different inputs. Taken together, these results suggest that CNNs could be used to reveal structure components of neuronal circuits, and provide a powerful model for neural system identification.
Qi Yan 0005, Yajing Zheng, Shanshan Jia 0001, Yichen Zhang 0002, Zhaofei Yu, Feng Chen 0007, Yonghong Tian 0001, Tiejun Huang 0001, Jian K. Liu
IEEE Trans. Cybern.8
2022 Asynchronous Spatio-Temporal Memory Network for Continuous Event-Based Object Detection
abstract
Event cameras, offering extremely high temporal resolution and high dynamic range, have brought a new perspective to addressing common object detection challenges (e.g., motion blur and low light). However, how to learn a better spatio-temporal representation and exploit rich temporal cues from asynchronous events for object detection still remains an open issue. To address this problem, we propose a novel asynchronous spatio-temporal memory network (ASTMNet) that directly consumes asynchronous events instead of event images prior to processing, which can well detect objects in a continuous manner. Technically, ASTMNet learns an asynchronous attention embedding from the continuous event stream by adopting an adaptive temporal sampling strategy and a temporal attention convolutional module. Besides, a spatio-temporal memory module is designed to exploit rich temporal cues via a lightweight yet efficient inter-weaved recurrent-convolutional architecture. Empirically, it shows that our approach outperforms the state-of-the-art methods using the feed-forward frame-based detectors on three datasets by a large margin (i.e., 7.6% in the KITTI Simulated Dataset, 10.8% in the Gen1 Automotive Dataset, and 10.5% in the 1Mpx Detection Dataset). The results demonstrate that event cameras can perform robust object detection even in cases where conventional cameras fail, e.g., fast motion and challenging light conditions.
Jianing Li 0001, Jia Li 0003, Lin Zhu 0012, Xijie Xiang, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Image Process.5
2022 GraphReg: Dynamical Point Cloud Registration With Geometry-Aware Graph Signal Processing
abstract
This study presents a high-accuracy, efficient, and physically induced method for 3D point cloud registration, which is the core of many important 3D vision problems. In contrast to existing physics-based methods that merely consider spatial point information and ignore surface geometry, we explore geometry aware rigid-body dynamics to regulate the particle (point) motion, which results in more precise and robust registration. Our proposed method consists of four major modules. First, we leverage the graph signal processing (GSP) framework to define a new signature, i.e., point response intensity for each point, by which we succeed in describing the local surface variation, resampling keypoints, and distinguishing different particles. Then, to address the shortcomings of current physics-based approaches that are sensitive to outliers, we accommodate the defined point response intensity to median absolute deviation (MAD) in robust statistics and adopt the X84 principle for adaptive outlier depression, ensuring a robust and stable registration. Subsequently, we propose a novel geometric invariant under rigid transformations to incorporate higher-order features of point clouds, which is further embedded for force modeling to guide the correspondence between pairwise scans credibly. Finally, we introduce an adaptive simulated annealing (ASA) method to search for the global optimum and substantially accelerate the registration process. We perform comprehensive experiments to evaluate the proposed method on various datasets captured from range scanners to LiDAR. Results demonstrate that our proposed method outperforms representative state-of-the-art approaches in terms of accuracy and is more suitable for registering large-scale point clouds. Furthermore, it is considerably faster and more robust than most competitors. Our implementation is publicly available at https://github.com/zikai1/GraphReg.
Mingyang Zhao 0001, Lei Ma 0008, Xiaohong Jia 0001, Dong-Ming Yan 0001, Tiejun Huang 0001
IEEE Trans. Image Process.5
2022 Self-Guided Adaptation: Progressive Representation Alignment for Domain Adaptive Object Detection
abstract
Unsupervised domain adaptation (UDA) has achieved unprecedented success in improving the cross-domain robustness of object detection models. However, existing UDA methods largely ignore the instantaneous data distribution and the sampling strategy during model learning, which could deteriorate the feature representation given large domain shift. In this work, we propose a Self-Guided Adaptation (SGA) model, targeting at aligning feature representation and transferring object detection models across domains while considering the instantaneous alignment difficulty. The core of SGA is to calculate “hardness” factors for sample pairs indicating domain distance in a kernel space. With the hardness factor, the proposed SGA adaptively indicates the importance of samples and assigns them different constrains. Indicated by these hardness factors, Self-Guided Progressive Sampling (SPS) is implemented in an “easy-to-hard” way during model adaptation. Using multi-stage convolutional features, SGA is further aggregated to fully align hierarchical representations of detection models. Extensive experiments on commonly-used benchmarks show that SGA improves the state-of-the-art methods with significant margins especially on large domain shift cases.
Zongxian Li, Peixi Peng, Qixiang Ye, Shijian Lu, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Multim.7
2022 Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification
abstract
introduction Share on Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification Authors: Shiliang Zhang Peking University Peking UniversityView Profile , Guorong Li University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Weigang Zhang Harbin Institute of Technology Harbin Institute of TechnologyView Profile , Qingming Huang University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Tiejun Huang Peking University Peking UniversityView Profile , Mubarak Shah University of Central Florida University of Central FloridaView Profile , Nicu Sebe University of Trento University of TrentoView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 18Issue 1sFebruary 2022 Article No.: 24pp 1–3https://doi.org/10.1145/3505280Online:25 January 2022Publication History 0citation169DownloadsMetricsTotal Citations0Total Downloads169Last 12 Months169Last 6 weeks22 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Shiliang Zhang, Guorong Li, Weigang Zhang, Qingming Huang, Tiejun Huang 0001, Mubarak Shah, Nicu Sebe
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Spk2ImgNet: Learning To Reconstruct Dynamic Scene From Continuous Spike Stream
abstract
The recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. Different from the conventional digital cameras that compact the photoelectric information within the exposure interval into a single snapshot, the spike camera produces a continuous spike stream to record the dynamic light intensity variation process. For spike cameras, image reconstruction remains an important and challenging issue. To this end, this paper develops a spike-to-image neural network (Spk2ImgNet) to reconstruct the dynamic scene from the continuous spike stream. In particular, to handle the challenges brought by both noise and high-speed motion, we propose a hierarchical architecture to exploit the temporal correlation of the spike stream progressively. Firstly, a spatially adaptive light inference subnet is proposed to exploit the local temporal correlation, producing basic light intensity estimates of different moments. Then, a pyramid deformable alignment is utilized to align the intermediate features such that the feature fusion module can exploit the long-term temporal correlation, while avoiding undesired motion blur. In addition, to train the network, we simulate the working mechanism of spike camera to generate a large-scale spike dataset composed of spike streams and corresponding ground truth images. Experimental results demonstrate that the proposed network evidently outperforms the state-of-the-art spike camera reconstruction methods.
Jing Zhao 0011, Ruiqin Xiong, Hangfan Liu, Jian Zhang 0018, Tiejun Huang 0001
CVPR5
2021 High-Speed Image Reconstruction Through Short-Term Plasticity for Spiking Cameras
abstract
Fovea, located in the centre of the retina, is specialized for high-acuity vision. Mimicking the sampling mechanism of the fovea, a retina-inspired camera, named spiking camera, is developed to record the external information with a sampling rate of 40,000 Hz, and outputs asynchronous binary spike streams. Although the temporal resolution of visual information is improved, how to reconstruct the scenes is still a challenging problem. In this paper, we present a novel high-speed image reconstruction model through the short-term plasticity (STP) mechanism of the brain. We derive the relationship between postsynaptic potential regulated by STP and the firing frequency of each pixel. By setting up the STP model at each pixel of the spiking camera, we can infer the scene radiance with the temporal regularity of the spike stream. Moreover, we show that STP can be used to distinguish the static and motion areas and further enhance the reconstruction results. The experimental results show that our methods achieve state-of-the-art performance in both image quality and computing time.
Yajing Zheng, Lingxiao Zheng, Zhaofei Yu, Boxin Shi, Yonghong Tian 0001, Tiejun Huang 0001
CVPR6
2021 Allocating DNN Layers Computation Between Front-End Devices and The Cloud Server for Video Big Data Processing
abstract
With the development of intelligent hardware, front-end devices can also perform DNN computation. Moreover, the deep neural network can be divided into several layers. In this way, part of the computation of DNN models can be migrated to the front-end devices, which can alleviate the cloud burden and shorten the processing latency. This paper proposes a computation allocation algorithm of DNN between the front-end devices and the cloud server. In brief, we divide the DNN layers dynamically according to the current and the predicted future status of the processing system, by which we obtain a shorter end-to-end latency. The simulation results reveal that the overall latency reduction is more than 70% compared with traditional cloud-centered processing.
Peiyin Xing, Peixi Peng, Tiejun Huang 0001, Yonghong Tian 0001
ICASSP4
2021 Super Resolve Dynamic Scene from Continuous Spike Streams
abstract
Recently, a novel retina-inspired camera, namely spike camera, has shown great potential for recording high-speed dynamic scenes. Unlike conventional digital cameras that compact the visual information within an exposure interval into a single snapshot, the spike camera continuously outputs binary spike streams to record the dynamic scenes, yielding a very high temporal resolution. Most of the existing reconstruction methods for spike camera focus on reconstructing images with the same resolution as spike camera. However, as a trade-off of high temporal resolution, the spatial resolution of spike camera is limited, resulting in inferior details of the reconstruction. To address this issue, we develop a spike camera super-resolution framework, aiming to super resolve high-resolution intensity images from the low-resolution binary spike streams. Due to the relative motion between the camera and the objects to capture, the spikes fired by the same sensor pixel no longer describes the same points in the external scene. In this paper, we exploit the relative motion and derive the relationship between light intensity and each spike, so as to recover the external scene with both high temporal and high spatial resolution. Experimental results demonstrate that the proposed method can reconstruct pleasant high-resolution images from low- resolution spike streams.
Jing Zhao 0011, Jiyu Xie, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Tiejun Huang 0001
ICCV6
2021 NeuSpike-Net: High Speed Video Reconstruction via Bio-inspired Neuromorphic Cameras
abstract
Neuromorphic vision sensor is a new bio-inspired imaging paradigm that emerged in recent years, which continuously sensing luminance intensity and firing asynchronous spikes (events) with high temporal resolution. Typically, there are two types of neuromorphic vision sensors, namely dynamic vision sensor (DVS) and spike camera. From the perspective of bio-inspired sampling, DVS only perceives movement by imitating the retinal periphery, while the spike camera was developed to perceive fine textures by simulating the fovea. It is meaningful to explore how to combine two types of neuromorphic cameras to reconstruct high quality image like human vision. In this paper, we propose a NeuSpike-Net to learn both the high dynamic range and high motion sensitivity of DVS and the full texture sampling of spike camera to achieve high-speed and high dynamic image reconstruction. We propose a novel representation to effectively extract the temporal information of spike and event data. By introducing the feature fusion module, the two types of neuromorphic data achieve complementary to each other. The experimental results on the simulated and real datasets demonstrate that the proposed approach is effective to reconstruct high-speed and high dynamic range images via the combination of spike and event data.
Lin Zhu 0012, Jianing Li 0001, Xiao Wang 0014, Tiejun Huang 0001, Yonghong Tian 0001
ICCV4
2021 Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have attracted enormous research interest due to temporal information processing capability, low power consumption, and high biological plausibility. However, the formulation of efficient and high-performance learning algorithms for SNNs is still challenging. Most existing learning methods learn weights only, and require manual tuning of the membrane-related parameters that determine the dynamics of a single spiking neuron. These parameters are typically chosen to be the same for all neurons, which limits the diversity of neurons and thus the expressiveness of the resulting SNNs. In this paper, we take inspiration from the observation that membrane-related parameters are different across brain regions, and propose a training algorithm that is capable of learning not only the synaptic weights but also the membrane time constants of SNNs. We show that incorporating learnable membrane time constants can make the network less sensitive to initial values and can speed up learning. In addition, we reevaluate the pooling methods in SNNs and find that max-pooling will not lead to significant information loss and have the advantage of low computation cost and binary compatibility. We evaluate the proposed method for image classification tasks on both traditional static MNIST, Fashion-MNIST, CIFAR-10 datasets, and neuromorphic N-MNIST, CIFAR10-DVS, DVS128 Gesture datasets. The experiment results show that the proposed method outperforms the state-of-the-art accuracy on nearly all datasets, using fewer time-steps. Our codes are available at https://github.com/fangwei123456/Parametric-Leaky-Integrate-and-Fire-Spiking-Neuron.
Wei Fang 0006, Zhaofei Yu, Yanqi Chen, Timothée Masquelier, Tiejun Huang 0001, Yonghong Tian 0001
ICCV5
2021 Recover The Residual Of Residual: Recurrent Residual Refinement Network For Image Super-Resolution
abstract
Benefiting from learning the residual between low resolution (LR) image and high resolution (HR) image, image super-resolution (SR) networks demonstrate superior reconstruction performance in recent studies. However, for the images with rich texture information, the residuals are complex and difficult for networks to learn. To address this problem, we propose a recurrent residual refinement network (RRRN) to gradually refine the residual with a recurrent structure. Instead of directly reconstructing the residual between LR image and HR image, each sub-network in our framework reconstructs the residual between SR image from previous stage and HR image, i.e. recovers the residual of residual (RoR). Considering the domain gap between the image feature and the RoR feature, we introduce a residual projection block to explicitly transform the feature from image domain to RoR domain. The RoR feature is further optimized in an iterative up- and down-sampling manner with a residual learning block. We construct the structure of each block based on the optimization methods of conventional SR and improve our network with dense connections. Experimental results prove that our method improves the quality of super-resolution images on different datasets with variable scenes.
Tianxiao Gao, Ruiqin Xiong, Rui Zhao 0010, Jian Zhang 0018, Shuyuan Zhu, Tiejun Huang 0001
ICIP6
2021 Pruning of Deep Spiking Neural Networks through Gradient Rewiring
abstract
Spiking Neural Networks (SNNs) have been attached great importance due to their biological plausibility and high energy-efficiency on neuromorphic chips. As these chips are usually resource-constrained, the compression of SNNs is thus crucial along the road of practical use of SNNs. Most existing methods directly apply pruning approaches in artificial neural networks (ANNs) to SNNs, which ignore the difference between ANNs and SNNs, thus limiting the performance of the pruned SNNs. Besides, these methods are only suitable for shallow SNNs. In this paper, inspired by synaptogenesis and synapse elimination in the neural system, we propose gradient rewiring (Grad R), a joint learning algorithm of connectivity and weight for SNNs, that enables us to seamlessly optimize network structure without retraining. Our key innovation is to redefine the gradient to a new synaptic parameter, allowing better exploration of network structures by taking full advantage of the competition between pruning and regrowth of connections. The experimental results show that the proposed method achieves minimal loss of SNNs' performance on MNIST and CIFAR-10 datasets so far. Moreover, it reaches a ~3.5% accuracy loss under unprecedented 0.73% connectivity, which reveals remarkable structure refining capability in SNNs. Our work suggests that there exists extremely high redundancy in deep SNNs. Our codes are available at https://github.com/Yanqi-Chen/Gradient-Rewiring.
Yanqi Chen, Zhaofei Yu, Wei Fang 0006, Tiejun Huang 0001, Yonghong Tian 0001
IJCAI4
2021 Optimal ANN-SNN Conversion for Fast and Accurate Inference in Deep Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs), as bio-inspired energy-efficient neural networks, have attracted great attentions from researchers and industry. The most efficient way to train deep SNNs is through ANN-SNN conversion. However, the conversion usually suffers from accuracy loss and long inference time, which impede the practical application of SNN. In this paper, we theoretically analyze ANN-SNN conversion and derive sufficient conditions of the optimal conversion. To better correlate ANN-SNN and get greater accuracy, we propose Rate Norm Layer to replace the ReLU activation function in source ANN training, enabling direct conversion from a trained ANN to an SNN. Moreover, we propose an optimal fit curve to quantify the fit between the activation value of source ANN and the actual firing rate of target SNN. We show that the inference time can be reduced by optimizing the upper bound of the fit curve in the revised ANN to achieve fast inference. Our theory can explain the existing work on fast reasoning and get better results. The experimental results show that the proposed method achieves near loss-less conversion with VGG-16, PreActResNet-18, and deeper structures. Moreover, it can reach 8.6× faster reasoning performance under 0.265× energy consumption of the typical method. The code is available at https://github.com/DingJianhao/OptSNNConvertion-RNL-RIL.
Jianhao Ding, Zhaofei Yu, Yonghong Tian 0001, Tiejun Huang 0001
IJCAI4
2021 Noisy Adaptation Generates Lévy Flights in Attractor Neural Networks
abstract
Lévy flights describe a special class of random walks whose step sizes satisfy a power-law tailed distribution. As being an efficientsearching strategy in unknown environments, Lévy flights are widely observed in animal foraging behaviors. Recent studies further showed that human cognitive functions also exhibit the characteristics of Lévy flights. Despite being a general phenomenon, the neural mechanism at the circuit level for generating Lévy flights remains unresolved. Here, we investigate how Lévy flights can be achieved in attractor neural networks. To elucidate the underlying mechanism clearly, we first study continuous attractor neural networks (CANNs), and find that noisy neural adaptation, exemplified by spike frequency adaptation (SFA) in this work, can generate Lévy flights representing transitions of the network state in the attractor space. Specifically, the strength of SFA defines a travelling wave boundary, below which the network state displays local Brownian motion, and above which the network state displays long-jump motion. Noises in neural adaptation causes the network state to intermittently switch between these two motion modes, manifesting the characteristics of Lévy flights. We further extend the study to a general attractor neural network, and demonstrate that our model can explain the Lévy-flight phenomenon observed during free memory retrieval of humans. We hope that this study will give us insight into understanding the neural mechanism for optimal information processing in the brain.
Xingsi Dong, Tianhao Chu, Tiejun Huang 0001, Zilong Ji, Si Wu 0001
NeurIPS3
2021 Deep Residual Learning in Spiking Neural Networks
abstract
Deep Spiking Neural Networks (SNNs) present optimization difficulties for gradient-based approaches due to discrete binary activation and complex spatial-temporal dynamics. Considering the huge success of ResNet in deep learning, it would be natural to train deep SNNs with residual learning. Previous Spiking ResNet mimics the standard residual block in ANNs and simply replaces ReLU activation layers with spiking neurons, which suffers the degradation problem and can hardly implement residual learning. In this paper, we propose the spike-element-wise (SEW) ResNet to realize residual learning in deep SNNs. We prove that the SEW ResNet can easily implement identity mapping and overcome the vanishing/exploding gradient problems of Spiking ResNet. We evaluate our SEW ResNet on ImageNet, DVS Gesture, and CIFAR10-DVS datasets, and show that SEW ResNet outperforms the state-of-the-art directly trained SNNs in both accuracy and time-steps. Moreover, SEW ResNet can achieve higher performance by simply adding more layers, providing a simple method to train deep SNNs. To our best knowledge, this is the first time that directly training deep SNNs with more than 100 layers becomes possible. Our codes are available at https://github.com/fangwei123456/Spike-Element-Wise-ResNet.
Wei Fang 0006, Zhaofei Yu, Yanqi Chen, Tiejun Huang 0001, Timothée Masquelier, Yonghong Tian 0001
NeurIPS4
2021 A brain-inspired computational model for spatio-temporal information processing
abstract
Spatio-temporal information processing is fundamental in both brain functions and AI applications. Current strategies for spatio-temporal pattern recognition usually involve explicit feature extraction followed by feature aggregation, which requires a large amount of labeled data. In the present study, motivated by the subcortical visual pathway and early stages of the auditory pathway for motion and sound processing, we propose a novel brain-inspired computational model for generic spatio-temporal pattern recognition. The model consists of two modules, a reservoir module and a decision-making module. The former projects complex spatio-temporal patterns into spatially separated neural representations via its recurrent dynamics, the latter reads out neural representations via integrating information over time, and the two modules are linked together using known examples. Using synthetic data, we demonstrate that the model can extract the frequency and order information of temporal inputs. We apply the model to reproduce the looming pattern discrimination behavior as observed in experiments successfully. Furthermore, we apply the model to the gait recognition task, and demonstrate that our model accomplishes the recognition in an event-based manner and outperforms deep learning counterparts when training data is limited.
Xiaohan Lin, Xiaolong Zou, Zilong Ji, Tiejun Huang 0001, Si Wu 0001, Yuanyuan Mi
Neural Networks4
2021 Digital Retina: A Way to Make the City Brain More Efficient by Visual Coding
abstract
The ubiquitous camera networks in the city brain system grow at a rapid pace, creating massive amounts of images and videos at a range of spatial-temporal scales and thereby forming the “biggest” big data. However, the sensing system often lags behind the construction of the fast-growing city brain system, in the sense that such exponentially growing data far exceed today’s sensing capabilities. Therefore, critical issues arise regarding how to better leverage the existing city brain system and significantly improve the city-scale performance in intelligent applications. To tackle the unprecedented challenges, we articulate a vision towards a novel visual computing framework, termed asdigital retina, which aligns high-efficiency sensing models with the emerging Visual Coding for Machine (VCM) paradigm. In particular, digital retina may consist of video coding, feature coding, model coding, as well as their joint optimization. The digital retina is biologically-inspired, rooted on the widely accepted view that the retina encodes the visual information for human perception, and extracts features by the brain downstream areas to disentangle the visual objects. Within the digital retina framework, three streams, i.e., video stream, feature stream, and model stream, work collaboratively over the end-edge-cloud platform. In particular, the compressed video stream serves for human vision, the compact feature stream targets for machine vision, and the model stream incrementally updates deep learning models to improve the performance of human/machine vision tasks. We have developed a prototype to demonstrate the technical advantages of digital retina, and extensive experiments have been conducted to validate that it is able to effectively support the video big data analysis and retrieval in the intelligent city system. In particular, up to$7000\times $compression ratio could be realized for visual data compression while maintaining competitive performance with pristine signal in a series of visual analysis tasks.
Wen Gao 0001, Siwei Ma 0001, Ling-Yu Duan, Yonghong Tian 0001, Peiyin Xing, Yaowei Wang 0001, Shanshe Wang, Huizhu Jia, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.9
2021 Hybrid Coding of Spatiotemporal Spike Data for a Bio-Inspired Camera
abstract
Recently, a novel bio-inspired camera was developed by mimicking the retina fovea to continuously accumulate luminance intensity and then fire spikes once the dispatch threshold is reached. In contrast to the conventional frame-based cameras and the emerging dynamic vision sensors, this spike camera has shown remarkable advantages in capturing fast-moving scenes in a frame-free manner with full texture reconstruction capabilities. However, the ultra-high temporal resolution makes the transmission or storage of the output data of spike camera (referred to as spike data) quite difficult. To address the above challenges, we propose a unified lossy spike coding framework, which exploits the motion patterns hidden in the spike data distribution to design the motion-fidelity coding modes for the first time. We investigate the spatiotemporal distribution of spike data and propose an intensity-based measurement of the spike train distance. Then, the adaptive polyhedron partitioning is proposed to deal with the spike data with different motion characteristics. Finally, the intra-/inter-polyhedron prediction with spike-time and spike-rate modes, transform and multi-layer quantization are proposed and introduced into the codec. We also construct a PKU-Spike dataset captured by the spike camera to evaluate the compression performance. The experimental results on the dataset demonstrate that the proposed approach is effective in compressing such spike data while maintaining the visual fidelity especially for high-speed scenarios.
Lin Zhu 0012, Siwei Dong, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Neuromorphic Camera Guided High Dynamic Range Imaging
abstract
Reconstruction of high dynamic range image from a single low dynamic range image captured by a frame-based conventional camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras are able to record high dynamic range scenes in the form of an intensity map, with much lower spatial resolution, and without color. In this paper, we propose a neuromorphic camera guided high dynamic range imaging pipeline, and a network consisting of specially designed modules according to each step in the pipeline, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images. A hybrid camera system has been built to validate that the proposed method is able to reconstruct quantitatively and qualitatively high-quality high dynamic range images by successfully fusing the images and intensity maps for various real-world scenarios.
Jin Han 0001, Chu Zhou, Peiqi Duan 0002, Yehui Tang 0001, Chang Xu 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
CVPR7
2020 Joint Filtering of Intensity Images and Neuromorphic Events for High-Resolution Noise-Robust Imaging
abstract
We present a novel computational imaging system with high resolution and low noise. Our system consists of a traditional video camera which captures high-resolution intensity images, and an event camera which encodes high-speed motion as a stream of asynchronous binary events. To process the hybrid input, we propose a unifying framework that first bridges the two sensing modalities via a noise-robust motion compensation model, and then performs joint image filtering. The filtered output represents the temporal gradient of the captured space-time volume, which can be viewed as motion-compensated event frames with high resolution and low noise. Therefore, the output can be widely applied to many existing event-based algorithms that are highly dependent on spatial resolution and noise robustness. In experimental results performed on both publicly available datasets as well as our contributing RGB-DAVIS dataset, we show systematic performance improvement in applications such as high frame-rate video synthesis, feature/corner detection and tracking, as well as high dynamic range image reconstruction.
Zihao W. Wang, Peiqi Duan 0002, Oliver Cossairt, Aggelos K. Katsaggelos, Tiejun Huang 0001, Boxin Shi
CVPR5
2020 Retina-Like Visual Image Reconstruction via Spiking Neural Model
abstract
The high-sensitivity vision of primates, including humans, is mediated by a small retinal region called the fovea. As a novel bio-inspired vision sensor, spike camera mimics the fovea to record the nature scenes by continuous-time spikes instead of frame-based manner. However, reconstructing visual images from the spikes remains to be a challenge. In this paper, we design a retina-like visual image reconstruction framework, which is flexible in reconstructing full texture of natural scenes from the totally new spike data. Specifically, the proposed architecture consists of motion local excitation layer, spike refining layer and visual reconstruction layer motivated by bio-realistic leaky integrate and fire (LIF) neurons and synapse connection with spike-timing-dependent plasticity (STDP) rules. This approach may represent a major shift from conventional frame-based vision to the continuous-time retina-like vision, owning to the advantages of high temporal resolution and low power consumption. To test the performance, a spike dataset is constructed which is recorded by the spike camera. The experimental results show that the proposed approach is extremely effective in reconstructing the visual image in both normal and high speed scenes, while achieving high dynamic range and high image quality.
Lin Zhu 0012, Siwei Dong, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001
CVPR4
2020 Binary Representation and High Efficient Compression of 3D CNN Features for Action Recognition
abstract
A common framework of the action recognition is to collect the videos from different cameras into a cloud center firstly, and then perform the 3D CNN on the cloud server. Although directly, this framework will bring a huge burden to the cloud server and video transmission. To handle this challenge, the "front-cloud" collaborative processing architecture can be used. The most import issue is to compress the feature from 3D CNN effectively without significant loss of accuracy. We propose logarithmic quantization with a maximum value threshold and HEVC inter encoding for 3D CNN features. Experimental results on ResNet-50 and InceptionV1 show that the features can be represented by only 1 bit without significant loss of accuracy. The compression ratio of the quantized 1 bit features using HEVC inter coding can reach to 5000 times and the loss of accuracy is less than 1%.
Peiyin Xing, Peixi Peng, Yongsheng Liang 0001, Tiejun Huang 0001, Yonghong Tian 0001
DCC4
2020 Learning Open Set Network with Discriminative Reciprocal Points
Limeng Qiao, Yemin Shi 0001, Peixi Peng, Jia Li 0003, Tiejun Huang 0001, Shiliang Pu, Yonghong Tian 0001
ECCV (3)6
2020 An Attention-Driven Two-Stage Clustering Method for Unsupervised Person Re-identification
Zilong Ji, Xiaolong Zou, Xiaohan Lin, Tiejun Huang 0001, Si Wu 0001
ECCV (28)5
2020 Graph Convolutional Reinforcement Learning
Jiechuan Jiang, Chen Dun, Tiejun Huang 0001, Zongqing Lu 0002
ICLR3
2020 High-Speed Motion Scene Reconstruction for Spike Camera via Motion Aligned Filtering
abstract
A new retina-inspired bionic spike camera has recently shown great potential for capturing high speed movements. Unlike conventional cameras with a fixed low sampling rate, retina-inspired spike camera can well record fast-moving scenes by continuously accumulating luminance intensity and firing spikes. To restore the captured high-speed motion scenes from spike data, several reconstruction methods have been proposed. A typical method utilizes two neighbouring spikes to infer the instantaneous luminance intensity. Although high temporal resolution imaging can be achieved, the signal to noise ratio (SNR) of reconstructions is generally unsatisfactory. For improving the SNR, some methods propose to average the spikes in big time window. However, the reconstructions may suffer from undesired motion blur, especially when there are objects moving very fast in scenes. To address this issue, we develop a new image reconstruction approach for potential retina-inspired spike camera to recover high-speed motion scenes. Specially, we take the motion of objects into consideration and exploit optical flow to align the scenes of different moments. After motion alignment, a filtering along motion trajectory can be employed to the signals to take the advantage of temporal correlations while not introducing undesired motion blur. Experimental results demonstrate that our proposed method achieves better visual quality than previous reconstruction schemes.
Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001
ISCAS3
2020 Learning Individually Inferred Communication for Multi-Agent Cooperation
abstract
Communication lays the foundation for human cooperation. It is also crucial for multi-agent cooperation. However, existing work focuses on broadcast communication, which is not only impractical but also leads to information redundancy that could even impair the learning process. To tackle these difficulties, we propose Individually Inferred Communication (I2C), a simple yet effective model to enable agents to learn a prior for agent-agent communication. The prior knowledge is learned via causal inference and realized by a feed-forward neural network that maps the agent's local observation to a belief about who to communicate with. The influence of one agent on another is inferred via the joint action-value function in multi-agent reinforcement learning and quantified to label the necessity of agent-agent communication. Furthermore, the agent policy is regularized to better exploit communicated messages. Empirically, we show that I2C can not only reduce communication overhead but also improve the performance in a variety of multi-agent cooperative scenarios, comparing to existing methods.
Ziluo Ding, Tiejun Huang 0001, Zongqing Lu 0002
NeurIPS2
2020 UnModNet: Learning to Unwrap a Modulo Image for High Dynamic Range Imaging
abstract
A conventional camera often suffers from over- or under-exposure when recording a real-world scene with a very high dynamic range (HDR). In contrast, a modulo camera with a Markov random field (MRF) based unwrapping algorithm can theoretically accomplish unbounded dynamic range but shows degenerate performances when there are modulus-intensity ambiguity, strong local contrast, and color misalignment. In this paper, we reformulate the modulo image unwrapping problem into a series of binary labeling problems and propose a modulo edge-aware model, named as UnModNet, to iteratively estimate the binary rollover masks of the modulo image for unwrapping. Experimental results show that our approach can generate 12-bit HDR images from 8-bit modulo images reliably, and runs much faster than the previous MRF-based algorithm thanks to the GPU acceleration.
Chu Zhou, Hang Zhao 0021, Jin Han 0001, Chang Xu 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
NeurIPS6
2020 Motion Estimation for Spike Camera Data Sequence via Spike Interval Analysis
abstract
With the development of emerging computer vision applications, there is an increasing demand for capturing the scenes with high-speed motion. Recently, a novel retina-inspired spike camera has shown great potential for recording the dynamic scenes at high temporal resolution. Different from the conventional digital cameras that capture the visual scene by a single snapshot, the spike camera monitors the incoming light persistently, with each pixel producing a continuous stream of spikes. Recovering the motion process from the spike data sequence is an important problem to study for the spike camera, as it is the foundation of many other tasks, such as image reconstruction, object tracking and object detection. In this paper, we carefully analyze the characteristics of spike data and develop a motion estimation algorithm to recover the continuous high-speed motion process from the spike camera data sequence. Based on the assumption that the spike intervals passed by the same motion trajectories usually have the similar spike densities, we establish a data term constraint to model the temporal consistency of spike intervals. In addition, we integrate a local smoothness constraint with the proposed data term constraint to further improve the estimation accuracy. Experimental results demonstrate that our proposed algorithm can recover high-speed motion process from the captured spike data, and the recovered motion information is beneficial for i mage reconstruction.
Jing Zhao 0011, Ruiqin Xiong, Rui Zhao 0010, Jin Wang 0023, Siwei Ma 0001, Tiejun Huang 0001
VCIP6
2020 Optical Flow Estimation Between Images of Different Resolutions via Variational Method
abstract
Traditional optical flow estimation methods mostly focus on images of the same resolution. However, there are some situations requiring optical flow between images of different resolutions, where the traditional approaches suffer from the inequality of spectrum aliasing level. In this paper, we propose a method estimating the flow fields between a clear image and a highly undersampled one. The proposed method simultaneously describes the motion and integral relationship between the images via an integral form image under the assumption of brightness and gradient consistency as well as motion smoothness. We also derive the numerical solution briefly, through which we can solve the equations easily via linearizations. Experimental results on Middlebury and MPI-Sintel datasets demonstrate that our proposed method outperforms traditional methods preprocessing images of different resolutions to be the same size, offering more accurate results.
Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Bing Zeng 0001, Tiejun Huang 0001, Wen Gao 0001
VCIP5
2020 Global Co-occurrence Feature Learning and Active Coordinate System Conversion for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition has attracted more and more attention in recent years. Besides, the rapid development of deep learning has greatly improved the performance. However, the current exploration of action co-occurrence is still not comprehensive enough. Most existing works only mine co-occurrence features from the temporal or spatial domain seperately, and it's common to combine them in the end. Different from previous works, our approach is able to learn temporal and spatial co-occurrence features integratedly and globally, which is called spatio-temporal-unit feature enhancement (STUFE). In order to better align the skeleton data, we introduce a novel method for skeleton data preprocessing called active coordinate system conversion (ACSC). A coordinate system can be learned automatically to transform skeleton samples for alignment. By the way, the proposed methods are compatible with current two types of mainstream models, the CNN-based and GCN-based models. Finally, on the two benchmarks of NTU-RGB+D and SBU Kinect Interaction, we validated our methods based on two mainstream models. The results show that our methods achieve the state-of-the-art.
Tingting Jiang 0001, Tiejun Huang 0001, Yonghong Tian 0001
WACV3
2020 E2BoWs: An end-to-end Bag-of-Words model via deep convolutional neural network for image retrieval
Shiliang Zhang, Tiejun Huang 0001, Qi Tian 0001
Neurocomputing3
2020 Reconstruction of natural visual scenes from neural spikes with deep neural networks
Yichen Zhang 0002, Shanshan Jia 0001, Yajing Zheng, Zhaofei Yu, Yonghong Tian 0001, Siwei Ma 0001, Tiejun Huang 0001, Jian K. Liu
Neural Networks7
2020 Probabilistic inference of binary Markov random fields in spiking neural networks through mean-field approximation
Yajing Zheng, Shanshan Jia 0001, Zhaofei Yu, Tiejun Huang 0001, Jian K. Liu, Yonghong Tian 0001
Neural Networks4
2020 CDbin: Compact Discriminative Binary Descriptor Learned With Efficient Neural Network
abstract
As an important computer vision task, image matching requires efficient and discriminative local descriptors. Most of the existing descriptors like SIFT and ORB are hand-crafted; therefore it is necessary to study more optimized descriptors through end-to-end learning. This paper proposes the compact binary descriptors learned with a lightweight Convolutional Neural Network (CNN), which is efficient for training and testing. Specifically, we propose a CNN with no larger than five layers for descriptor learning. The resulting descriptors, i.e., Compact Discriminative binary descriptors (CDbin) are optimized with four complementary loss functions, i.e., 1) triplet loss to ensure the discriminative power; 2) quantization loss to decrease the quantization error; 3) correlation loss to ensure the feature compactness; and 4) even-distribution loss to enrich the embedded information. The extensive experiments on two image patch datasets and three image retrieval datasets show that the CDbin exhibits competitive performance compared with the existing descriptors. For example, the 64-bit CDbin substantially outperforms the 256-bit ORB and 1024-bit SIFT on Hpatches dataset. Although generated by a shallow CNN, CDbin also outperforms several recent deep descriptors.
Jianming Ye, Shiliang Zhang, Tiejun Huang 0001, Yong Rui
IEEE Trans. Circuits Syst. Video Technol.3
2020 Joint Coding of Local and Global Deep Features in Videos for Visual Search
abstract
Practically, it is more feasible to collect compact visual features rather than the video streams from hundreds of thousands of cameras into the cloud for big data analysis and retrieval. Then the problem becomes which kinds of features should be extracted, compressed and transmitted so as to meet the requirements of various visual tasks. Recently, many studies have indicated that the activations from the convolutional layers in convolutional neural networks (CNNs) can be treated as local deep features describing particular details inside an image region, which are then aggregated (e.g., using Fisher Vectors) as a powerful global descriptor. Combination of local and global features can satisfy those various needs effectively. It has also been validated that, if only local deep features are coded and transmitted to the cloud while the global features are recovered using the decoded local features, the aggregated global features should be lossy and consequently would degrade the overall performance. Therefore, this paper proposes a joint coding framework for local and global deep features (DFJC) extracted from videos. In this framework, we introduce a coding scheme for real-valued local and global deep features with intra-frame lossy coding and inter-frame reference coding. The theoretical analysis is performed to understand how the number of inliers varies with the number of local features. Moreover, the inter-feature correlations are exploited in our framework. That is, local feature coding can be accelerated by making use of the frame types determined with global features, while the lossy global features aggregated with the decoded local features can be used as a reference for global feature coding. Extensive experimental results under three metrics show that our DFJC framework can significantly reduce the bitrate of local and global deep features from videos while maintaining the retrieval performance.
Lin Ding 0002, Yonghong Tian 0001, Hongfei Fan, Changhuai Chen, Tiejun Huang 0001
IEEE Trans. Image Process.5
2020 Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
abstract
Video coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with compactness and efficiency to serve for machine vision, and the other is with full fidelity, bowing to human perception. The recent endeavors in imminent trends of video compression, e.g. deep learning based coding tools and end-to-end image/video coding, and MPEG-7 compact feature descriptor standards, i.e. Compact Descriptors for Visual Search and Compact Descriptors for Video Analysis, promote the sustainable and fast development in their own directions, respectively. In this paper, thanks to booming AI technology, e.g. prediction and generation models, we carry out exploration in the new area, Video Coding for Machines (VCM), arising from the emerging MPEG standardization efforts1. Towards collaborative compression and intelligent analytics, VCM attempts to bridge the gap between feature coding for machine vision and video coding for human vision. Aligning with the rising Analyze then Compress instance Digital Retina, the definition, formulation, and paradigm of VCM are given first. Meanwhile, we systematically review state-of-the-art techniques in video compression and feature compression from the unique perspective of MPEG standardization, which provides the academic and industrial evidence to realize the collaborative compression of video and feature streams in a broad range of AI applications. Finally, we come up with potential VCM solutions, and the preliminary results have demonstrated the performance and efficiency gains. Further direction is discussed as well.
Ling-Yu Duan, Jiaying Liu 0001, Wenhan Yang, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Image Process.4
2020 Perceptual Temporal Incoherence-Guided Stereo Video Retargeting
abstract
Stereo video retargeting aims at minimizing shape and depth distortions with temporal coherence in resizing a stereo video content to a desired size. Existing methods extend stereo image retargeting schemes to stereo video retargeting by adding additional temporal constraints that demand temporal coherence in all corresponding regions. However, such a straightforward extension incurs conflicts among multiple requirements (i.e., shape and depth preservation and their temporal coherence), thus failing to meet one or more of these requirements satisfactorily. To mitigate conflicts among depth, shape, and temporal constraints and avoid degrading temporal coherence perceptually, we relax temporal constraints for non-paired regions at frame boundaries, derive new temporal constraints to improve human viewing experience of a 3D scene, and propose an efficient grid-based implementation for stereo video retargeting. Experimental results demonstrate that our method achieves superior visual quality over existing methods.
Bing Li 0024, Chia-Wen Lin, Shan Liu 0001, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo
IEEE Trans. Image Process.4
2020 Multi-Scale Temporal Cues Learning for Video Person Re-Identification
abstract
Temporal cues embedded in videos provide important clues for person Re-Identification (ReID). To efficiently exploit temporal cues with a compact neural network, this work proposes a novel 3D convolution layer called Multi-scale 3D (M3D) convolution layer. The M3D layer is easy to implement and could be inserted into traditional 2D convolution networks to learn multi-scale temporal cues by end-to-end training. According to its inserted location, the M3D layer has two variants, i.e., local M3D layer and global M3D layer, respectively. The local M3D layer is inserted between 2D convolution layers to learn spatial-temporal cues among adjacent 2D feature maps. The global M3D layer is computed on adjacent frame feature vectors to learn their global temporal relations. The local and global M3D layers hence learn complementary temporal cues. Their combination introduces a fraction of parameters to traditional 2D CNN, but leads to the strong multi-scale temporal feature learning capability. The learned temporal feature is fused with a spatial feature to compose the final spatial-temporal representation for video person ReID. Evaluations on four widely used video person ReID datasets, i.e., MARS, DukeMTMC-VideoReID, PRID2011, and iLIDS-VID demonstrate the substantial advantages of our method over the state-of-the art. For example, it achieves rank1 accuracy of 88.63% on MARS without re-ranking. Our method also achieves a reasonable trade-off between ReID accuracy and model size, e.g., it saves about 40% parameters of I3D CNN.
Jianing Li 0001, Shiliang Zhang, Tiejun Huang 0001
IEEE Trans. Image Process.3
2020 Compressed Image Restoration via Artifacts-Free PCA Basis Learning and Adaptive Sparse Modeling
abstract
Visually unpleasant compression artifacts frequently appear in block-based transform coding, especially at low bit rates. This paper presents a new artifact reduction scheme based on Bayesian sparse modeling and artifacts-free PCA basis learning. To avoid the effect of blocking artifacts, we propose to learn artifacts-free PCA basis from clean images. We concatenate the clean patches and their compressed counterparts to learn paired distribution prior via the Gaussian Mixture Model (GMM). By this way, the GMM characterizes the mapping between the clean image and its compressed version. To restore a compressed patch, the best matched GMM component is assigned using the patch in the compressed image subspace. The artifacts-free PCA basis is obtained according to the mapping learned by the paired GMM. In practice, the statistical distributions of different sparse coefficients in different patches may dramatically vary with image contents. Instead of using a global zero-mean distribution for all coefficients, we propose to adaptively model the prior of each band in a Bayesian framework. The expectation and variance of each band are adaptively learned from the similar patches within the image. Thus, different transform bands are regularized unequally according to the learned priors. Experimental results show that the proposed scheme outperforms most of the compared schemes in terms of both objective quality and perceptual quality.
Ruiqin Xiong, Xiaopeng Fan 0001, Dong Liu 0002, Feng Wu 0001, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Image Process.6
2019 Multi-Scale 3D Convolution Network for Video Based Person Re-Identification
abstract
This paper proposes a two-stream convolution network to extract spatial and temporal cues for video based person ReIdentification (ReID). A temporal stream in this network is constructed by inserting several Multi-scale 3D (M3D) convolution layers into a 2D CNN network. The resulting M3D convolution network introduces a fraction of parameters into the 2D CNN, but gains the ability of multi-scale temporal feature learning. With this compact architecture, M3D convolution network is also more efficient and easier to optimize than existing 3D convolution networks. The temporal stream further involves Residual Attention Layers (RAL) to refine the temporal features. By jointly learning spatial-temporal attention masks in a residual manner, RAL identifies the discriminative spatial regions and temporal cues. The other stream in our network is implemented with a 2D CNN for spatial feature extraction. The spatial and temporal features from two streams are finally fused for the video based person ReID. Evaluations on three widely used benchmarks datasets, i.e.,MARS, PRID2011, and iLIDS-VID demonstrate the substantial advantages of our method over existing 3D convolution networks and state-of-art methods.
Jianing Li 0001, Shiliang Zhang, Tiejun Huang 0001
AAAI3
2019 Bi-Directional Cascade Network for Perceptual Edge Detection
abstract
Exploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a Bi-Directional Cascade Network (BDCN) structure, where an individual layer is supervised by labeled edges at its specific scale, rather than directly applying the same supervision to all CNN outputs. Furthermore, to enrich multi-scale representations learned by BDCN, we introduce a Scale Enhancement Module (SEM) which utilizes dilated convolution to generate multi-scale features, instead of using deeper CNNs or explicitly fusing multi-scale edge maps. These new approaches encourage the learning of multi-scale representations in different layers and detect edges that are well delineated by their scales. Learning scale dedicated layers also results in compact network with a fraction of parameters. We evaluate our method on three datasets, i.e., BSDS500, NYUDv2, and Multicue, and achieve ODS Fmeasure of 0.828, 1.3% higher than current state-of-the art on BSDS500.
Shiliang Zhang, Ming Yang 0007, Yanhu Shan, Tiejun Huang 0001
CVPR5
2019 An Efficient Coding Method for Spike Camera Using Inter-Spike Intervals
abstract
Recently, a novel bio-inspired spike camera has been proposed, which continuously accumulates luminance intensity and fires spikes once the dispatch threshold is reached. It has shown great advantages in capturing fast-moving scene in a frame-free manner with full texture reconstruction capabilities. However, it is difficult to transmit or store the large amount of spike data. By investigating the spatiotemporal distribution of the spikes, we propose an intensity-based measurement for spike train distance and design an efficient coding method to meet the challenge. First, the spike train is transformed into inter-spike intervals (ISIs), and ISIs are adaptively partitioned into multiple segments in temporal. Then, intra-and inter-pixel prediction are performed to find the best reference candidate. The prediction residuals are quantized to achieve lossy compression. Finally, the quantized residuals are fed into an adaptive context-based entropy coder. Overall, to achieve the best performance, each prediction mode will be tried and the one with minimum rate-distortion cost is chosen.
Siwei Dong, Lin Zhu 0012, Daoyuan Xu, Yonghong Tian 0001, Tiejun Huang 0001
DCC5
2019 Spike Coding: Towards Lossy Compression for Dynamic Vision Sensor
abstract
Dynamic vision sensor (DVS) as a bio-inspired camera, has shown great advantages in high dynamic range (HDR) and high temporal resolution (us) in vision tasks. However, how to lossy compress asynchronous spikes for meeting the demand of large-scale transmission and storage meanwhile maintaining the analysis performance still remains open. Towards this end, this paper proposes a lossy spike coding framework for DVS.
Yihua Fu, Jianing Li 0001, Siwei Dong, Yonghong Tian 0001, Tiejun Huang 0001
DCC5
2019 Transductive Episodic-Wise Adaptive Metric for Few-Shot Learning
abstract
Few-shot learning, which aims at extracting new concepts rapidly from extremely few examples of novel classes, has been featured into the meta-learning paradigm recently. Yet, the key challenge of how to learn a generalizable classifier with the capability of adapting to specific tasks with severely limited data still remains in this domain. To this end, we propose a Transductive Episodic-wise Adaptive Metric (TEAM) framework for few-shot learning, by integrating the meta-learning paradigm with both deep metric learning and transductive inference. With exploring the pairwise constraints and regularization prior within each task, we explicitly formulate the adaptation procedure into a standard semi-definite programming problem. By solving the problem with its closed-form solution on the fly with the setup of transduction, our approach efficiently tailors an episodic-wise metric for each task to adapt all features from a shared task-agnostic embedding space into a more discriminative task-specific metric space. Moreover, we further leverage an attention-based bi-directional similarity strategy for extracting the more robust relationship between queries and prototypes. Extensive experiments on three benchmark datasets show that our framework is superior to other existing approaches and achieves the state-of-the-art performance in the few-shot literature.
Limeng Qiao, Yemin Shi 0001, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Yaowei Wang 0001
ICCV5
2019 Event-Based Vision Enhanced: A Joint Detection Framework in Autonomous Driving
abstract
Due to the high-speed motion blur and low dynamic range, conventional frame-based cameras have encountered an important challenge in object detection, especially in autonomous driving. Event-based cameras, by taking the advantages of high temporal resolution and high dynamic range, have brought a new perspective to address the challenge. Motivated by this fact, this paper proposes a joint framework combining event-based and frame-based vision for vehicle detection. Specially, two separate event-based and frame-based streams are incorporated into a convolutional neural network (CNN). Besides, to accommodate the asynchronous events from event-based cameras, a convolutional spiking neural network (SNN) is utilized to generate visual attention maps so that two streams can be synchronized. Moreover, Dempster-Shafer theory is introduced to merge two outputs from CNN in a joint decision model. The experimental results show that the proposed approach outperforms the state-of-the-art methods only using frame-based information, especially in fast motion and challenging illumination conditions.
Jianing Li 0001, Siwei Dong, Zhaofei Yu, Yonghong Tian 0001, Tiejun Huang 0001
ICME5
2019 Learning a Deep Convolutional Network for Subband Image Denoising
abstract
Due to the fast inference and excellent learning capability, deep learning has become an effective means for image denoising and attracted considerable attention recently. However, for the images with rich textures and structures, the performance of deep learning approaches is still unsatisfactory. To address this issue, we develop a new convolutional neural network (CNN) for subband image denoising and name it SDCNN. In the proposed approach, we first decompose images into transform domain and denoise the coefficients of various subbands. By incorporating frequency information with spatial context, SDCNN is more effective in recovering image details. In particular, the introduced procedure of subband transform also plays the role of downsampling and enlarges the receptive field without increasing depth or sacrificing efficiency of network. Experimental results show that the SDCNN achieves promising results in terms of both objective and subjective performance.
Jing Zhao 0011, Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Tiejun Huang 0001
ICME5
2019 A Retina-Inspired Sampling Method for Visual Texture Reconstruction
abstract
Conventional frame-based camera is not able to meet the demand of rapid reaction for real-time applications, while the emerging dynamic vision sensor (DVS) can realize high speed capturing for moving objects. However, to achieve visual texture reconstruction, DVS need extra information apart from the output spikes. This paper introduces a fovea-like sampling method inspired by the neuron signal processing in retina, which aims at visual texture reconstruction only taking advantage of the properties of spikes. In the proposed method, the pixels independently respond to the luminance changes with temporal asynchronous spikes. Analyzing the arrivals of spikes makes it possible to restore the luminance information, enabling reconstructing the natural scene for visualization. Three decoding methods of spike stream for texture reconstruction are proposed for high-speed motion and stationary scenes. Compared to conventional frame-based camera and DVS, our model can achieve better image quality and higher flexibility, which is capable of changing the way that demanding machine vision applications are built.
Lin Zhu 0012, Siwei Dong, Tiejun Huang 0001, Yonghong Tian 0001
ICME3
2019 Push-pull Feedback Implements Hierarchical Information Retrieval Efficiently
abstract
Experimental data has revealed that in addition to feedforward connections, there exist abundant feedback connections in a neural pathway. Although the importance of feedback in neural information processing has been widely recognized in the field, the detailed mechanism of how it works remains largely unknown. Here, we investigate the role of feedback in hierarchical information retrieval. Specifically, we consider a hierarchical network storing the hierarchical categorical information of objects, and information retrieval goes from rough to fine, aided by dynamical push-pull feedback from higher to lower layers. We elucidate that the push (positive) and pull (negative) feedbacks suppress the interferences due to neural correlations between different and the same categories, respectively, and their joint effect improves retrieval performance significantly. Our model agrees with the push-pull phenomenon observed in neural data and sheds light on our understanding of the role of feedback in neural information processing.
Xiaolong Zou, Zilong Ji, Gengshuo Tian, Yuanyuan Mi, Tiejun Huang 0001, K. Y. Michael Wong, Si Wu 0001
NeurIPS6
2019 Excitation-Inhibition Balanced Spiking Neural Networks for Fast Information Processing
abstract
The balance of excitation and inhibition is a fundamental property of neural systems. The present study investigates an excitation and inhibition (E-I) balanced spiking neural network model for neuromorphic computing, in particular, to track rapid changes of external inputs. We analyze the working mechanism of an E-I balanced network and find that the network generates internal noises of a nearly optimal structure which enables neural population firing rates to track input changes almost instantly. Moreover, we extend the network model from homogenous connectivity to local connectivity, so that the network can remain balanced under spatially heterogeneous inputs and retain spatial information. Simulation results confirm that the model works well. This model may serve as a fast responding module for general neuromorphic computing systems.
Gengshuo Tian, Tiejun Huang 0001, Si Wu 0001
SMC2
2019 3D Human Skeleton Data Compression for Action Recognition
abstract
Skeleton-based action recognition continues to open up new application scenarios with the popularity of acquisition devices. This also leads to a rapid increase in the amount of human skeleton data. Currently, there is no skeleton data compression algorithm for the task of action recognition. In order to solve this problem, we propose the first skeleton data compression algorithm, which can compress the skeleton data stream to a small bandwidth while keeping the accuracy of action recognition as high as possible. The proposed compression algorithm is called Motion-based Joints Selection (MJS). It performs compression based on the amount of movement of different joints. In addition, we also explored the combination of MJS and existing lossless compression methods, and found the most suitable one. In the end, we verify that our compression method MJS can achieve promising results on the large dataset NTU-RGB+D.
Tingting Jiang 0001, Yonghong Tian 0001, Tiejun Huang 0001
VCIP4
2019 Spike Coding for Dynamic Vision Sensor in Intelligent Driving
abstract
Dynamic vision sensor (DVS) as a bio-inspired camera, has shown great advantages in wide dynamic range and high temporal resolution imaging in contrast to conventional frame-based cameras. Its ability to capture high speed moving objects enables fast and accurate detection which plays a significant role in the emerging intelligent driving applications. The pixels in DVS independently respond to the luminance changes with output spikes. Thus, the spike stream conveying the x -, y -addresses, the firing time, and the polarity (ON/OFF), is quite different from conventional video frames. How to compress this kind of new data for efficient transmission and storage remains a big challenge, especially for on-board detection, monitoring and recording. To address this challenge, this paper first analyzes the spike firing mechanism and the spatiotemporal characteristics of the spike data, then introduces a cube-based spike coding framework for DVS. In the framework, an octree-based structure is proposed to adaptively partition the spike stream into coding cubes in both spatial and temporal dimensions, then several prediction modes are designed to exploit the spatial and temporal characteristics of spikes for compression, including address-prior mode and time-prior mode. To explore more flexibility, the intercube prediction is discussed extensively involving motion estimation and motion compensation. Finally, the experimental results demonstrate that our approach achieves an impressive coding performance with the average compression ratio of 2.6536 against the raw spike data, which is much higher than the results of conventional lossless coding algorithms.
Siwei Dong, Zhichao Bi, Yonghong Tian 0001, Tiejun Huang 0001
IEEE Internet Things J.4
2019 Guest Editorial Special Issue on IoT on the Move: Enabling Technologies and Driving Applications for Internet of Intelligent Vehicles (IoIV)
abstract
The new era of the Internet of Things (IoT) is prompting the evolution of conventional vehicle ad-hoc networks (VANETs) into the Internet of Intelligent Vehicles (IoIV). Different from VANETs, where a vehicle is essentially considered as a node disseminating messages, the emerging IoIV paradigm is expected to regard each vehicle as a smart object equipped with a powerful multisensor platform, unprecedented communication capability, computing units, and Internet protocol (IP)-based connectivity. As such, the vehicles in IoIV are highly efficient in a broad array of vehicular and transportation applications. As a unique subset of general purpose IoT, IoIV can benefit from the existing research on VANET, which lays the foundation toward a more pervasive and ubiquitous communications and networking core that is essential for IoIV. Nevertheless, research in many aspects of IoIV, especially those that are application-driven and data-oriented ones, is still at its infancy.
Liuqing Yang 0001, Xiang Cheng 0001, Mounir Ghogho, Ender Ayanoglu, Tiejun Huang 0001, Nanning Zheng 0001
IEEE Internet Things J.5
2019 Toward Knowledge as a Service Over Networks: A Deep Learning Model Communication Paradigm
abstract
The advent of artificial intelligence and Internet of Things has led to the seamless transition turning the big data into the big knowledge. The deep learning models, which assimilate knowledge from large-scale data, can be regarded as an alternative but promising modality of knowledge for artificial intelligence services. Yet, the compression, storage, and communication of the deep learning models towards better knowledge services, especially over networks, pose a set of challenging problems on both industrial and academic realms. This paper presents the deep learning model communication paradigm based on multiple model compression, which greatly exploits the redundancy among multiple deep learning models in different application scenarios. We analyze the potential and demonstrate the promise of the compression strategy for deep learning model communication through a set of experiments. Moreover, the interoperability in deep learning model communication, which is enabled based on the standardization of compact deep learning model representation, is also discussed and envisioned.
Ziqian Chen, Ling-Yu Duan, Shiqi Wang 0001, Yihang Lou, Tiejun Huang 0001, Dapeng Oliver Wu, Wen Gao 0001
IEEE J. Sel. Areas Commun.5
2019 Robust estimation for image noise based on eigenvalue distributions of large sample covariance matrices
Rui Chen 0006, Changshui Yang, Yuan Li 0014, Tiejun Huang 0001
J. Vis. Commun. Image Represent.5
2019 Multiscale video sequence matching for near-duplicate detection and retrieval
Yonghong Tian 0001, Tiejun Huang 0001
Multim. Tools Appl.3
2018 SAP: Self-Adaptive Proposal Model for Temporal Action Detection Based on Reinforcement Learning
abstract
Existing action detection algorithms usually generate action proposals through an extensive search over the video at multiple temporal scales, which brings about huge computational overhead and deviates from the human perception procedure. We argue that the process of detecting actions should be naturally one of observation and refinement: observe the current window and refine the span of attended window to cover true action regions. In this paper, we propose a Self-Adaptive Proposal (SAP) model that learns to find actions through continuously adjusting the temporal bounds in a self-adaptive way. The whole process can be deemed as an agent, which is firstly placed at the beginning of the video and traverse the whole video by adopting a sequence of transformations on the current attended region to discover actions according to a learned policy. We utilize reinforcement learning, especially the Deep Q-learning algorithm to learn the agent’s decision policy. In addition, we use temporal pooling operation to extract more effective feature representation for the long temporal window, and design a regression network to adjust the position offsets between predicted results and the ground truth. Experiment results on THUMOS’14 validate the effectiveness of SAP, which can achieve competitive performance with current action detection algorithms via much fewer proposals.
Jingjia Huang, Nannan Li 0001, Tao Zhang 0069, Ge Li 0002, Tiejun Huang 0001, Wen Gao 0001
AAAI5
2018 Depth-Aware Stereo Video Retargeting
abstract
As compared with traditional video retargeting, stereo video retargeting poses new challenges because stereo video contains the depth information of salient objects and its time dynamics. In this work, we propose a depth-aware stereo video retargeting method by imposing the depth fidelity constraint. The proposed depth-aware retargeting method reconstructs the 3D scene to obtain the depth information of salient objects. We cast it as a constrained optimization problem, where the total cost function includes the shape, temporal and depth distortions of salient objects. As a result, the solution can preserve the shape, temporal and depth fidelity of salient objects simultaneously. It is demonstrated by experimental results that the depth-aware retargeting method achieves higher retargeting quality and provides better user experience.
Bing Li 0024, Chia-Wen Lin, Boxin Shi, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo
CVPR4
2018 Spike Coding for Dynamic Vision Sensors
abstract
As an emerging kind of retinomorphic camera, the dynamic vision sensors (DVS) have shown great advantages in wide dynamic range and high temporal resolution in various applications such as autonomous driving and high-speed motion photography. However, how to compress the output spike data of DVS still remains a big challenge. To address this challenge, this paper firstly analyzes the spike firing mechanism and the redundancies of the spike data generated from DVS, and then introduces an efficient cube-based coding framework. Typically, a spike in DVS contains the location (the x-, y- addresses, the timestamp) and the polarity (On/Off). Three key strategies are designed to exploit the spatial and temporal characteristics of the spike location information for compression, including the adaptive macro-cube partitioning structure, the address-prior mode and the time-prior mode. Finally, the experimental results demonstrate that our approach achieves an impressive coding performance, with the average compression ratio of 19.519 over the original spike data, which is much higher than the results of conventional lossless coding algorithms.
Zhichao Bi, Siwei Dong, Yonghong Tian 0001, Tiejun Huang 0001
DCC4
2018 Compressed Image Restoration via External-Image Assisted Band Adaptive PCA Model Learning
abstract
Visually annoying compression artifacts frequently appear in block-based transform coding at low bit rates, due to coarse and independent quantization of transform coefficients in coding blocks. This paper presents a subband adaptive modeling framework for reducing quantization artifacts. In this framework, each patch is jointly regularized by bandwise distribution priors adaptively learned in its PCA transform domain together with a quantization constraint prior in the DCT domain. Since the compression artifacts influence the covariance statistics of coded image patches remarkably, external images are utilized to provide more robust PCA domains for patch sparse modeling. Instead of using a global distribution model for all patches, the distribution prior of each patch is adaptively learned from similar patches within the compressed image itself to address the non-stationarity of image signals. The coefficients in different PCA bands are regularized unequally according to the learned priors. Experimental results show that the proposed scheme outperforms existing schemes in terms of both the objective and the perceptual qualities.
Ruiqin Xiong, Xiaopeng Fan 0001, Xianming Liu 0005, Tiejun Huang 0001, Wen Gao 0001
DCC5
2018 Temporal Attentive Network for Action Recognition
abstract
In action recognition, one of the most important challenges is to jointly utilize the texture and motion information as well as capturing the long-term dependence of various common and action-specific postures. Motivated by this fact, this paper proposes Temporal Attentive Network (TAN) for action recognition. The key idea in TAN is that not all postures, each of which represented by a small collection of consecutive frames, contribute equally to the successful recognition of an action. As a result, TAN incorporates two separate spatial and temporal streams into one network. Information in the two streams is partially shared so that discriminative spatiotemporal features can be extracted to characterize various postures in an action. Moreover, a temporal attention mechanism is introduced in the form of Long-Short Term Memory (LSTM) network. With this mechanism, features from the action-specific postures can be emphasized, while common postures shared by many different actions will be ignored to some extent. By jointly using such spatial and temporal information as well as attentive cues in a single network, TAN achieves impressive performance on two public datasets, HMDB51 and UCF101, with accuracy scores of 72.5% and 94.1 %, respectively.
Yemin Shi 0001, Yonghong Tian 0001, Tiejun Huang 0001, Yaowei Wang 0001
ICME3
2018 Neural Information Processing in Hierarchical Prototypical Networks
Zilong Ji, Xiaolong Zou, Tiejun Huang 0001, Yuanyuan Mi, Si Wu 0001
ICONIP (3)4
2018 Learning, Storing, and Disentangling Correlated Patterns in Neural Networks
Xiaolong Zou, Zilong Ji, Tiejun Huang 0001, Yuanyuan Mi, Dahui Wang, Si Wu 0001
ICONIP (3)4
2018 From Data to Knowledge: Deep Learning Model Compression, Transmission and Communication
abstract
With the advances of artificial intelligence, recent years have witnessed a gradual transition from the big data to the big knowledge. Based on the knowledge-powered deep learning models, the big data such as the vast text, images and videos can be efficiently analyzed. As such, in addition to data, the communication of knowledge implied in the deep learning models is also strongly desired. As a specific example regarding the concept of knowledge creation and communication in the context of Knowledge Centric Networking (KCN), we investigate the deep learning model compression and demonstrate its promise use through a set of experiments. In particular, towards future KCN, we introduce efficient transmission of deep learning models in terms of both single model compression and multiple model prediction. The necessity, importance and open problems regarding the standardization of deep learning models, which enables the interoperability with the standardized compact model representation bitstream syntax, are also discussed.
Ziqian Chen, Shiqi Wang 0001, Dapeng Oliver Wu, Tiejun Huang 0001, Ling-Yu Duan
ACM Multimedia4
2018 Perceptual Temporal Incoherence Aware Stereo Video Retargeting
abstract
Stereo video retargeting aims to avoid shape and depth distortions while maintaining temporal coherence of shape and depth while resizing a stereo video to a desired size. Existing methods resort to extending stereo image retargeting schemes to stereo video retargeting by imposing temporal constraints to consistently resize all corresponding regions so as to maintain temporal coherence. However, such a direct extension often incurs conflicts among the requirements for preserving shape information and depth information and maintaining their temporal coherence, thereby failing to meet one or more of these requirements. We find that properly relaxing temporal constraints for non-paired regions at frame boundaries can effectively mitigate conflicts among depth, shape, and temporal constraints without severely degrading temporal coherence perceptually. Based on this new finding, we derive effective temporal constraints to improve the viewing experience of a 3D scene for stereo video retargeting. Accordingly, we propose an efficient grid-based implementation for our method. Experimental results show that our method achieves superior visual quality over existing methods.
Bing Li 0024, Chia-Wen Lin, Shan Liu 0001, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo
ACM Multimedia4
2018 Cross-Domain Adversarial Feature Learning for Sketch Re-identification
abstract
Under person re-identification (Re-ID), a query photo of the target person is often required for retrieval. However, one is not always guaranteed to have such a photo readily available under a practical forensic setting. In this paper, we define the problem of Sketch Re-ID, which instead of using a photo as input, it initiates the query process using a professional sketch of the target person. This is akin to the traditional problem of forensic facial sketch recognition, yet with the major difference that our sketches are whole-body other than just the face. This problem is challenging because sketches and photos are in two distinct domains. Specifically, a sketch is the abstract description of a person. Besides, person appearance in photos is variational due to camera viewpoint, human pose and occlusion. We address the Sketch Re-ID problem by proposing a cross-domain adversarial feature learning approach to jointly learn the identity features and domain-invariant features. We employ adversarial feature learning to filter low-level interfering features and remain high-level semantic information. We also contribute to the community the first Sketch Re-ID dataset with 200 persons, where each person has one sketch and two photos from different cameras associated. Extensive experiments have been performed on the proposed dataset and other common sketch datasets including CUFSF and QUML-shoe. Results show that the proposed method outperforms the state-of-the-arts.
Lu Pang 0001, Yaowei Wang 0001, Yi-Zhe Song, Tiejun Huang 0001, Yonghong Tian 0001
ACM Multimedia4
2018 Implementation of Bayesian Inference In Distributed Neural Networks
abstract
Numerous neuroscience experiments have suggested that the cognitive process of human brain is realized as probability reasoning and further modeled as Bayesian inference. It is still unclear how Bayesian inference could be implemented by neural underpinnings in the brain. Here we present a novel Bayesian inference algorithm based on importance sampling. By distributed sampling through a deep tree structure with simple and stackable basic motifs for any given neural circuit, one can perform local inference while guaranteeing the accuracy of global inference. We show that these task-independent motifs can be used in parallel for fast inference without iteration and scale-limitation. Furthermore, experimental simulations with a small-scale neural network demonstrate that our distributed sampling-based algorithm, consisting with our theoretical analysis, can approximate Bayesian inference. Taken all together, we provide a proofof- principle to use distributed neural networks to implement Bayesian inference, which gives a road-map for large-scale Bayesian network implementation based on spiking neural networks with computer hardwares, including neuromorphic chips.
Zhaofei Yu, Tiejun Huang 0001, Jian K. Liu
PDP2
2018 PA-Search: Predicting units adaptive motion search for surveillance video coding
Yonghong Tian 0001, Jiaying Yan, Siwei Dong, Tiejun Huang 0001
Comput. Vis. Image Underst.4
2018 Joint Semantic and Latent Attribute Modelling for Cross-Class Transfer Learning
abstract
A number of vision problems such as zero-shot learning and person re-identification can be considered as cross-class transfer learning problems. As mid-level semantic properties shared cross different object classes, attributes have been studied extensively for knowledge transfer across classes. Most previous attribute learning methods focus only on human-defined/nameable semantic attributes, whilst ignoring the fact there also exist undefined/latent shareable visual properties, or latent attributes. These latent attributes can be either discriminative or non-discriminative parts depending on whether they can contribute to an object recognition task. In this work, we argue that learning the latent attributes jointly with user-defined semantic attributes not only leads to better representation but also helps semantic attribute prediction. A novel dictionary learning model is proposed which decomposes the dictionary space into three parts corresponding to semantic, latent discriminative and latent background attributes respectively. Such a joint attribute learning model is then extended by following a multi-task transfer learning framework to address a more challenging unsupervised domain adaptation problem, where annotations are only available on an auxiliary dataset and the target dataset is completely unlabelled. Extensive experiments show that the proposed models, though being linear and thus extremely efficient to compute, produce state-of-the-art results on both zero-shot learning and person re-identification.
Peixi Peng, Yonghong Tian 0001, Tao Xiang 0002, Yaowei Wang 0001, Massimiliano Pontil, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2018 MPEG Internet Video Coding Standard and Its Performance Evaluation
abstract
MPEG has produced standards that have provided the industry with the best video compression technologies. To address diverse Internet needs, MPEG issued a Call for Proposals (CfP) for Internet video coding (IVC) in July, 2011. The anticipation is that any patent declaration associated with the baseline profile of this standard will indicate that the patent owner is prepared to grant a free of charge license to an unrestricted number of applicants worldwide. Three codecs have responded to the CfP: Web video coding (WVC), video coding for browsers (VCB), and IVC. WVC is in fact the AVC baseline, and VCB uses the same coding tools as VP8. IVC has been developed in MPEG from scratch by combining well-known existing technology elements and new coding tools with royalty-free declarations. In June 2015, the IVC project was approved as ISO/IEC 14496-33 (MPEG-4 IVC). This standard can be highly beneficial for video services in the Internet domain. This paper describes the main coding tools used in IVC, and evaluates its objective and subjective performances compared with WVC, VCB, and AVC high profile (AVC HP). The experimental results show that IVC's compression performance is approximately equal to that of the AVC HP for typical operational settings, both for streaming and low-delay applications, and is superior to WVC and VCB.
Ronggang Wang, Zhenyu Wang 0002, Kui Fan, Tiejun Huang 0001, Wenmin Wang 0001, Ge Li 0002, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Learning Affective Features With a Hybrid Deep Model for Audio-Visual Emotion Recognition
abstract
Emotion recognition is challenging due to the emotional gap between emotions and audio-visual features. Motivated by the powerful feature learning ability of deep neural networks, this paper proposes to bridge the emotional gap by using a hybrid deep model, which first produces audio-visual segment features with Convolutional Neural Networks (CNNs) and 3D-CNN, then fuses audio-visual segment features in a Deep Belief Networks (DBNs). The proposed method is trained in two stages. First, CNN and 3D-CNN models pre-trained on corresponding large-scale image and video classification tasks are fine-tuned on emotion recognition tasks to learn audio and visual segment features, respectively. Second, the outputs of CNN and 3D-CNN models are combined into a fusion network built with a DBN model. The fusion network is trained to jointly learn a discriminative audio-visual segment feature representation. After average-pooling segment features learned by DBN to form a fixed-length global video feature, a linear Support Vector Machine is used for video emotion classification. Experimental results on three public audio-visual emotional databases, including the acted RML database, the acted eNTERFACE05 database, and the spontaneous BAUM-1s database, demonstrate the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues with CNN, 3D-CNN, and DBN for audio-visual emotion recognition.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Fast MPEG-CDVS Encoder With GPU-CPU Hybrid Computing
abstract
The compact descriptors for visual search (CDVS) standard from ISO/IEC moving pictures experts group has succeeded in enabling the interoperability for efficient and effective image retrieval by standardizing the bitstream syntax of compact feature descriptors. However, the intensive computation of a CDVS encoder unfortunately hinders its widely deployment in industry for large-scale visual search. In this paper, we revisit the merits of low complexity design of CDVS core techniques and present a very fast CDVS encoder by leveraging the massive parallel execution resources of graphics processing unit (GPU). We elegantly shift the computation-intensive and parallel-friendly modules to the state-of-the-arts GPU platforms, in which the thread block allocation as well as the memory access mechanism are jointly optimized to eliminate performance loss. In addition, those operations with heavy data dependence are allocated to CPU for resolving the extra but non-necessary computation burden for GPU. Furthermore, we have demonstrated the proposed fast CDVS encoder can work well with those convolution neural network approaches which enables to leverage the advantages of GPU platforms harmoniously, and yield significant performance improvements. Comprehensive experimental results over benchmarks are evaluated, which has shown that the fast CDVS encoder using GPU-CPU hybrid computing is promising for scalable visual search.
Ling-Yu Duan, Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Jianxiong Yin, Simon See, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
IEEE Trans. Image Process.8
2018 Speech Emotion Recognition Using Deep Convolutional Neural Network and Discriminant Temporal Pyramid Matching
abstract
Speech emotion recognition is challenging because of the affective gap between the subjective emotions and low-level features. Integrating multilevel feature learning and model training, deep convolutional neural networks (DCNN) has exhibited remarkable success in bridging the semantic gap in visual tasks like image classification, object detection. This paper explores how to utilize a DCNN to bridge the affective gap in speech signals. To this end, we first extract three channels of log Mel-spectrograms (static, delta, and delta delta) similar to the red, green, blue (RGB) image representation as the DCNN input. Then, the AlexNet DCNN model pretrained on the large ImageNet dataset is employed to learn high-level feature representations on each segment divided from an utterance. The learned segment-level features are aggregated by a discriminant temporal pyramid matching (DTPM) strategy. DTPM combines temporal pyramid matching and optimal Lp-norm pooling to form a global utterance-level feature representation, followed by the linear support vector machines for emotion classification. Experimental results on four public datasets, that is, EMO-DB, RML, eNTERFACE05, and BAUM-1s, show the promising performance of our DCNN model and the DTPM strategy. Another interesting finding is that the DCNN model pretrained for image applications performs reasonably good in affective speech feature extraction. Further fine tuning on the target emotional speech datasets substantially promotes recognition performance.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Multim.3
2017 Spike Camera and Its Coding Methods
abstract
Summary form only given. This paper introduces a spike camera with a distinct video capture scheme and proposes two methods of decoding the spike stream for texture reconstruction. The spike camera captures light and accumulates the converted luminance intensity at each pixel. A spike is fired when the accumulated intensity exceeds the dispatch threshold. The spike stream generated by the camera indicates the luminance variation. Analyzing the patterns of the spike stream makes it possible to reconstruct the picture of any moment which enables the playback of high speed movement.
Siwei Dong, Tiejun Huang 0001, Yonghong Tian 0001
DCC2
2017 Compact Deep Invariant Descriptors for Video Retrieval
abstract
With emerging demand for large-scale video analysis, the Motion Picture Experts Group (MPEG) initiated the Compact Descriptor for Video Analysis (CDVA) standardization in 2014. In this work, we develop novel deep-learning features and incorporate them into the well-established CDVA evaluation framework to study its effectiveness in video analysis. In particular, we propose a Nested Invariance Pooling (NIP) method to obtain compact and robust Convolutional Neural Network (CNNs) descriptors. The CNNs descriptors are generated by applying three different pooling operations to the feature maps of CNNs in a nested way towards rotation and scale invariant feature representation. In particular, the rational, advantages and performance on the combination of CNNs and handcrafted descriptors are provided to better investigate the complementary effects of deep learnt and handcrafted features. Extensive experimental results show that the proposed CNNs descriptors outperform both state-of-the-art CNNs descriptors and canonical handcrafted descriptors adopted in CDVA Experimental Model (CXM) with significant mAP gains of 11.3% and 4.7%, respectively. Moreover, the combination of NIP derived deep invariant descriptors and handcrafted descriptors not only fulfills the lowest bitrate budget of CDVA, but also significantly advances the performance of CDVA core techniques.
Yihang Lou, Jie Lin 0001, Shiqi Wang 0001, Jie Chen 0006, Vijay Chandrasekhar 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
DCC8
2017 Learning Long-Term Dependencies for Action Recognition with a Biologically-Inspired Deep Network
abstract
Despite a lot of research efforts devoted in recent years, how to efficiently learn long-term dependencies from sequences still remains a pretty challenging task. As one of the key models for sequence learning, recurrent neural network (RNN) and its variants such as long short term memory (LSTM) and gated recurrent unit (GRU) are still not powerful enough in practice. One possible reason is that they have only feedforward connections, which is different from the biological neural system that is typically composed of both feedforward and feedback connections. To address this problem, this paper proposes a biologicallyinspired deep network, called shuttleNet. Technologically, the shuttleNet consists of several processors, each of which is a GRU while associated with multiple groups of hidden states. Unlike traditional RNNs, all processors inside shuttleNet are loop connected to mimic the brain's feedforward and feedback connections, in which they are shared across multiple pathways in the loop connection. Attention mechanism is then employed to select the best information flow pathway. Extensive experiments conducted on two benchmark datasets (i.e UCF101 and HMDB51) show that we can beat state-of-the-art methods by simply embedding shuttleNet into a CNN-RNN framework.
Yemin Shi 0001, Yonghong Tian 0001, Yaowei Wang 0001, Wei Zeng 0006, Tiejun Huang 0001
ICCV5
2017 Exploiting Multi-grain Ranking Constraints for Precisely Searching Visually-similar Vehicles
abstract
Precise search of visually-similar vehicles poses a great challenge in computer vision, which needs to find exactly the same vehicle among a massive vehicles with visually similar appearances for a given query image. In this paper, we model the relationship of vehicle images as multiple grains. Following this, we propose two approaches to alleviate the precise vehicle search problem by exploiting multi-grain ranking constraints. One is Generalized Pairwise Ranking, which generalizes the conventional pairwise from considering only binary similar/dissimilar relations to multiple relations. The other is Multi-Grain based List Ranking, which introduces permutation probability to score a permutation of a multi-grain list, and further optimizes the ranking by the likelihood loss function. We implement the two approaches with multi-attribute classification in a multi-task deep learning framework. To further facilitate the research on precise vehicle search, we also contribute two high-quality and well-annotated vehicle datasets, named VD1 and VD2, which are collected from two different cities with diverse annotated attributes. As two of the largest publicly available precise vehicle search datasets, they contain 1,097,649 and 807,260 vehicle images respectively. Experimental results show that our approaches achieve the state-of-the-art performance on both datasets.
Ke Yan 0007, Yonghong Tian 0001, Yaowei Wang 0001, Wei Zeng 0006, Tiejun Huang 0001
ICCV5
2017 Deep regional feature pooling for video matching
abstract
In this work, we study the problem of deep global descriptors for video matching with regional feature pooling. We aim to analyze the joint effect of ROI (Region of Interest) size and pooling moment on video matching performance. To this end, we propose to mathematically model the distribution of video matching function with a pooling function nested in. Matching performance can be estimated by the separability of these class-conditional distributions between matching and non-matching pairs. Empirical studies on the challenging MPEG CDVA dataset demonstrate that performance trends are consistent with the estimation and experimental results, though the theoretical model is largely simplified compared to video matching and retrieval in practice.
Jie Lin 0001, Vijay Chandrasekhar 0001, Yihang Lou, Shiqi Wang 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot
ICIP7
2017 A Multi-Block N-ary trie structure for exact r-neighbour search in hamming space
abstract
This paper proposes a novel algorithm to solve the exact r-neighbour search problem in Hamming space. Existing r-neighbour search methods typically adopt hash table to index binary codes. Given a query, existing approaches search the nearest neighbours by checking all buckets of a Hamming ball centered at the query. The problem is these methods spend most of search time visiting empty buckets (lookup misses). In this paper, we adopt trie structure to index binary codes. We consider several continuous bits of a binary string as a block and use it as an atomic indexing element in trie structure, which is efficient in access speed and memory usage. Our method searches the nearest neighbours of a query by utilizing the records of nodes in trie structure to avoid lookup misses. We name the proposed indexing structure as Multi-Block N-ary Trie (MBNT). A theoretical analysis is given to prove that MBNT has less time cost than other hash-based methods. Extensive results show that MBNT outperforms state-of-the-art algorithms on several large scale benchmarks.
Ling-Yu Duan, Zhe Wang 0019, Jie Lin 0001, Vijay Chandrasekhar 0001, Tiejun Huang 0001
ICIP6
2017 Incorporating intra-class variance to fine-grained visual recognition
abstract
Fine-grained visual recognition aims to capture discriminative characteristics amongst visually similar categories. The state-of-the-art research work has significantly improved the fine-grained recognition performance by deep metric learning using triplet network. However, the impact of intra-category variance on the performance of recognition and robust feature representation has not been well studied. In this paper, we propose to leverage intra-class variance in metric learning of triplet network to improve the performance of fine-grained recognition. Through partitioning training images within each category into a few groups, we form the triplet samples across different categories as well as different groups, which is called Group Sensitive TRiplet Sampling (GS-TRS). Accordingly, the triplet loss function is strengthened by incorporating intra-class variance with GS-TRS, which may contribute to the optimization objective of triplet network. Extensive experiments over benchmark datasets CompCar and VehicleID show that the proposed GS-TRS has significantly outperformed state-of-the-art approaches in both classification and retrieval tasks.
Yan Em, Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan
ICME5
2017 Improving object detection with region similarity learning
abstract
Object detection aims to identify instances of semantic objects of a certain class in images or videos. The success of state-of-the-art approaches is attributed to the significant progress of object proposal and convolutional neural networks (CNNs). Most promising detectors involve multi-task learning with an optimization objective of softmax loss and regression loss. The first is for multi-class categorization, while the latter is for improving localization accuracy. However, few of them attempt to further investigate the hardness of distinguishing different sorts of distracting background regions (i.e., negatives) from true object regions (i.e., positives). To improve the performance of classifying positive object regions vs. a variety of negative background regions, we propose to incorporate triplet embedding into learning objective. The triplet units are formed by assigning each negative region to a meaningful object class and establishing class-specific negatives, followed by triplets construction. Over the benchmark PASCAL VOC 2007, the proposed triplet embedding has improved the performance of well-known Fas-tRCNN model with a mAP gain of 2.1%. In particular, the state-of-the-art approach OHEM can benefit from the triplet embedding and has achieved a mAP improvement of 1.2%.
Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan
ICME5
2017 From Part to Whole: Who is Behind the Painting?
abstract
Compared with normal modalities, the representations of paintings are much more complex due to its large intra-class and small inter-class variation. This poses more difficulties in the task of authorship identification. In this paper, we propose a multi-task multi-range (MTMR) representation framework and try to resolve this issue in two ways. First, we investigate how to improve the representation through multi-task learning. Specifically, we attempt to optimize authorship identification with subtly correlated identification tasks such as style, genre and date. Second, in order to make the representation more comprehensive and reduce the information loss from image scaling, we propose a multi-range structure which is composed of local, regional and global representations. Experiments on the two most representative large-scale painting datasets, Rijksmuseum Challenge and Wikiart, have shown that our method significantly outperforms the existing methods. To give better understanding and provide more effective predictions, we utilize random forest as the feature ranking method to analyze the importance of different features and apply external knowledge matching to further examine the predictions. Moreover, the framework's effects of identifying the authorship are visualized on the paintings' artist-characteristic regions and t-SNE is further applied to perform artist-based cluster analysis. Extensive validation has demonstrated that the proposed framework yields superior performance in the chanllenging task of painting authorship identification.
Daiqian Ma, Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan
ACM Multimedia6
2017 Cross-media analysis and reasoning: advances and directions
abstract
Cross-media analysis and reasoning is an active research area in computer science, and a promising direction for artificial intelligence. However, to the best of our knowledge, no existing work has summarized the state-of-the-art methods for cross-media analysis and reasoning or presented advances, challenges, and future directions for the field. To address these issues, we provide an overview as follows: (1) theory and model for cross-media uniform representation; (2) cross-media correlation understanding and deep mining; (3) cross-media knowledge graph construction and learning methodologies; (4) cross-media knowledge evolution and reasoning; (5) cross-media description and generation; (6) cross-media intelligent engines; and (7) cross-media intelligent applications. By presenting approaches, advances, and future directions in cross-media analysis and reasoning, our goal is not only to draw more attention to the state-of-the-art advances in the field, but also to provide technical insights by discussing the challenges and research directions in these areas.
Yuxin Peng 0001, Wenwu Zhu 0001, Yao Zhao 0001, Changsheng Xu, Qingming Huang, Hanqing Lu, Tiejun Huang 0001, Wen Gao 0001
Frontiers Inf. Technol. Electron. Eng.8
2017 Towards human-like and transhuman perception in AI 2.0: a review
abstract
Perception is the interaction interface between an intelligent system and the real world. Without sophisticated and flexible perceptual capabilities, it is impossible to create advanced artificial intelligence (AI) systems. For the next-generation AI, called ‘AI 2.0’, one of the most significant features will be that AI is empowered with intelligent perceptual capabilities, which can simulate human brain’s mechanisms and are likely to surpass human brain in terms of performance. In this paper, we briefly review the state-of-the-art advances across different areas of perception, including visual perception, auditory perception, speech perception, and perceptual information processing and learning engines. On this basis, we envision several R&D trends in intelligent perception for the forthcoming era of AI 2.0, including: (1) human-like and transhuman active vision; (2) auditory perception and computation in an actual auditory setting; (3) speech perception and computation in a natural interaction setting; (4) autonomous learning of perceptual information; (5) large-scale perceptual information processing and learning platforms; and (6) urban omnidirectional intelligent perception and reasoning engines. We believe these research directions should be highlighted in the future plans for AI 2.0.
Yonghong Tian 0001, Xilin Chen 0001, Hongkai Xiong, Li-Rong Dai 0001, Jing Chen 0002, Junliang Xing, Jing Chen 0003, Xihong Wu, Weiming Hu 0004, Yu Hu 0003, Tiejun Huang 0001, Wen Gao 0001
Frontiers Inf. Technol. Electron. Eng.12
2017 Rate-Performance-Loss Optimization for Inter-Frame Deep Feature Coding From Videos
abstract
With the explosion in the use of cameras in mobile phones or video surveillance systems, it is impossible to transmit a large amount of videos captured from a wide area into a cloud for big data analysis and retrieval. Instead, a feasible solution is to extract and compress features from videos and then transmit the compact features to the cloud. Meanwhile, many recent studies also indicate that the features extracted from the deep convolutional neural networks will lead to high performance for various analysis and recognition tasks. However, how to compress video deep features meanwhile maintaining the analysis or retrieval performance still remains open. To address this problem, we propose a high-efficiency deep feature coding (DFC) framework in this paper. In the DFC framework, we define three types of features in a group-of-features (GOFs) according to their coding modes (i.e., I-feature, P-feature, and S-feature). We then design two prediction structures for these features in a GOF, including a sequential prediction structure and an adaptive prediction structure. Similar to video coding, it is important for P-feature residual coding optimization to make a tradeoff between feature bitrate and analysis/retrieval performance when encoding residuals. To do so, we propose a rate-performance-loss optimization model. To evaluate various feature coding methods for large-scale video retrieval, we construct a video feature coding data set, called VFC-1M, which consists of uncompressed videos from different scenarios captured from real-world surveillance cameras, with totally 1M visual objects. Extensive experiments show that the proposed DFC can significantly reduce the bitrate of deep features in the videos while maintaining the retrieval accuracy.
Lin Ding 0002, Yonghong Tian 0001, Hongfei Fan, Yaowei Wang 0001, Tiejun Huang 0001
IEEE Trans. Image Process.5
2017 Active Sampling Exploiting Reliable Informativeness for Subjective Image Quality Assessment Based on Pairwise Comparison
abstract
Subjective image quality assessment (IQA) based on pairwise comparison (PC) overcome the shortcomings of IQA based on category rating, such as an ambiguous scale definition. However, the testing scale of PC tests can be very large, as the number of image pairs for comparison is a quadratic form of the number of images. To conduct PC tests on a large-scale image set with limited budget, an active sampling strategy to reduce testing scale is required. The conventional active sampling strategies usually select the most informative sample and assume that any image pair's correct label can be obtained from any subjects who are attentive. However, this is not true for IQA, because of human visual system's limitation. If two images are similar, their difference can be too subtle for some subjects to perceive. It means that it takes subjects more effort to obtain correct preference labels of two similar images, and that it is even impossible to obtain the correct preference labels of two images that are too similar. To address this issue, we study the reliability of preference labels. Based on the combination of reliability and informativeness, we design a new active sampling framework. It not only considers the informativeness, but also adjusts the effort spent on an image pair according to its ambiguity. Experiments show that this adjustment can effectively improve the performance of sampling strategies only based on informativeness. Besides, the proposed method is expected to be applied to more general subjective tests based on PC beyond IQA.
Tingting Jiang 0001, Tiejun Huang 0001
IEEE Trans. Multim.3
2017 HNIP: Compact Deep Invariant Representations for Video Matching, Localization, and Retrieval
abstract
With emerging demand for large-scale video analysis, MPEG initiated the compact descriptor for video analysis (CDVA) standardization in 2014. Beyond handcrafted descriptors adopted by the current MPEG-CDVA reference model, we study the problem of deep learned global descriptors for video matching, localization, and retrieval. First, inspired by a recent invariance theory, we propose a nested invariance pooling (NIP) method to derive compact deep global descriptors from convolutional neural networks (CNNs), by progressively encoding translation, scale, and rotation invariances into the pooled descriptors. Second, our empirical studies have shown that a sequence of well designed pooling moments (e.g., max or average) may drastically impact video matching performance, which motivates us to design hybrid pooling operations via NIP (HNIP). HNIP has further improved the discriminability of deep global descriptors. Third, the technical merits and performance improvements by combining deep and handcrafted descriptors are provided to better investigate the complementary effects. We evaluate the effectiveness of HNIP within the well-established MPEG-CDVA evaluation framework. The extensive experiments have demonstrated that HNIP outperforms the state-of-the-art deep and canonical handcrafted descriptors with significant mAP gains of 5.5% and 4.7%, respectively. In particular the combination of HNIP incorporated and handcrafted global descriptors has significantly boosted the performance of CDVA core techniques with comparable descriptor size.
Jie Lin 0001, Ling-Yu Duan, Shiqi Wang 0001, Yihang Lou, Vijay Chandrasekhar 0001, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
IEEE Trans. Multim.7
2017 Sequential Deep Trajectory Descriptor for Action Recognition With Three-Stream CNN
abstract
Learning the spatial-temporal representation of motion information is crucial to human action recognition. Nevertheless, most of the existing features or descriptors cannot capture motion information effectively, especially for long-term motion. To address this problem, this paper proposes a long-term motion descriptor called sequential deep trajectory descriptor (sDTD). Specifically, we project dense trajectories into two-dimensional planes, and subsequently a CNN-RNN network is employed to learn an effective representation for long-term motion. Unlike the popular two-stream ConvNets, the sDTD stream is introduced into a three-stream framework so as to identify actions from a video sequence. Consequently, this three-stream framework can simultaneously capture static spatial features, short-term motion, and long-term motion in the video. Extensive experiments were conducted on three challenging datasets: KTH, HMDB51, and UCF101. Experimental results show that our method achieves state-of-the-art performance on the KTH and UCF101 datasets, and is comparable to the state-of-the-art methods on the HMDB51 dataset.
Yemin Shi 0001, Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
IEEE Trans. Multim.4
2017 Learning Discriminative Subspaces on Random Contrasts for Image Saliency Analysis
abstract
In visual saliency estimation, one of the most challenging tasks is to distinguish targets and distractors that share certain visual attributes. With the observation that such targets and distractors can sometimes be easily separated when projected to specific subspaces, we propose to estimate image saliency by learning a set of discriminative subspaces that perform the best in popping out targets and suppressing distractors. Toward this end, we first conduct principal component analysis on massive randomly selected image patches. The principal components, which correspond to the largest eigenvalues, are selected to construct candidate subspaces since they often demonstrate impressive abilities to separate targets and distractors. By projecting images onto various subspaces, we further characterize each image patch by its contrasts against randomly selected neighboring and peripheral regions. In this manner, the probable targets often have the highest responses, while the responses at background regions become very low. Based on such random contrasts, an optimization framework with pairwise binary terms is adopted to learn the saliency model that best separates salient targets and distractors by optimally integrating the cues from various subspaces. Experimental results on two public benchmarks show that the proposed approach outperforms 16 state-of-the-art methods in human fixation prediction.
Shu Fang, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Xiaowu Chen 0001
IEEE Trans. Neural Networks Learn. Syst.4
2016 Affinity Preserving Quantization for Hashing: A Vector Quantization Approach to Compact Learn Binary Codes
abstract
Hashing techniques are powerful for approximate nearest neighbour (ANN) search.Existing quantization methods in hashing are all focused on scalar quantization (SQ) which is inferior in utilizing the inherent data distribution.In this paper, we propose a novel vector quantization (VQ) method named affinity preserving quantization (APQ) to improve the quantization quality of projection values, which has significantly boosted the performance of state-of-the-art hashing techniques.In particular, our method incorporates the neighbourhood structure in the pre- and post-projection data space into vector quantization.APQ minimizes the quantization errors of projection values as well as the loss of affinity property of original space.An effective algorithm has been proposed to solve the joint optimization problem in APQ, and the extension to larger binary codes has been resolved by applying product quantization to APQ.Extensive experiments have shown that APQ consistently outperforms the state-of-the-art quantization methods, and has significantly improved the performance of various hashing techniques.
Zhe Wang 0019, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
AAAI3
2016 Deep Relative Distance Learning: Tell the Difference between Similar Vehicles
abstract
The growing explosion in the use of surveillance cameras in public security highlights the importance of vehicle search from a large-scale image or video database. However, compared with person re-identification or face recognition, vehicle search problem has long been neglected by researchers in vision community. This paper focuses on an interesting but challenging problem, vehicle re-identification (a.k.a precise vehicle search). We propose a Deep Relative Distance Learning (DRDL) method which exploits a two-branch deep convolutional network to project raw vehicle images into an Euclidean space where distance can be directly used to measure the similarity of arbitrary two vehicles. To further facilitate the future research on this problem, we also present a carefully-organized largescale image database "VehicleID", which includes multiple images of the same vehicle captured by different realworld cameras in a city. We evaluate our DRDL method on our VehicleID dataset and another recently-released vehicle model classification dataset "CompCars" in three sets of experiments: vehicle re-identification, vehicle model verification and vehicle retrieval. Experimental results show that our method can achieve promising results and outperforms several state-of-the-art approaches.
Hongye Liu, Yonghong Tian 0001, Yaowei Wang 0001, Lu Pang 0001, Tiejun Huang 0001
CVPR5
2016 Unsupervised Cross-Dataset Transfer Learning for Person Re-identification
abstract
Most existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. This severely limits their scalability in realworld applications. To overcome this limitation, we develop a novel cross-dataset transfer learning approach to learn a discriminative representation. It is unsupervised in the sense that the target dataset is completely unlabelled. Specifically, we present an multi-task dictionary learning method which is able to learn a dataset-shared but target-data-biased representation. Experimental results on five benchmark datasets demonstrate that the method significantly outperforms the state-of-the-art.
Peixi Peng, Tao Xiang 0002, Yaowei Wang 0001, Massimiliano Pontil, Shaogang Gong, Tiejun Huang 0001, Yonghong Tian 0001
CVPR6
2016 Joint Learning of Semantic and Latent Attributes
Peixi Peng, Yonghong Tian 0001, Tao Xiang 0002, Yaowei Wang 0001, Tiejun Huang 0001
ECCV (4)5
2016 Depth-based local feature selection for mobile visual search
abstract
Selecting local features is crucial in generating robust compact descriptors for mobile visual search. The state-of-the-art MPEG Compact Descriptors for Visual Search (CDVS) standard has utilized the intrinsic characteristics (e.g., scale, orientation, peak, center distance, etc.) of interest points to select salient local features for selective aggregation and compression of local feature descriptors at different bit rates. In particular, the statistics of center distance was considered as an important attribute to select features in mobile visual search, which heavily relies on the assumption of a centralized object in a 2-dimensional query image. However, the ad-hoc assumption would probably fail to delineate query objects in a cluttered scene. In this paper, we propose to incorporate the depth cue to select local features. As most mobile phones are not yet equipped with depth sensor, we recover the disparity of local features through an auxiliary image to fast estimate the depth of a query image. The experiments have shown that, the incorporation of depth cue into feature selection can significantly improve the retrieval performance of the state-of-the-art CDVS compact descriptors at lower bit rates. For example, the mAP is improved from 84.5% to 88.6% at 512 bytes.
Zhaoliang Liu, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001
ICIP4
2016 Two-stage pooling of deep convolutional features for image retrieval
abstract
Convolutional Neural Network (CNN) based image representations have achieved high performance in image retrieval tasks. However, traditional CNN based global representations either provide high-dimensional features, which incurs large memory consumption and computing cost, or inadequately capture discriminative information in images, which degenerates the functionality of CNN features. To address those issues, we propose a two-stage partial mean pooling (PMP) approach to construct compact and discriminative global feature representations. The proposed PMP is meant to tackle the limits of traditional max pooling and mean (or average) pooling. By injecting the PMP pooling strategy into the CNN based patch-level mid-level feature extraction and representation, we have significantly improved the state-of-the-art retrieval performance over several common benchmark datasets.
Tiancheng Zhi, Ling-Yu Duan, Tiejun Huang 0001
ICIP4
2016 To Project More or to Quantize More: Minimize Reconstruction Bias for Learning Compact Binary Codes
Zhe Wang 0019, Ling-Yu Duan, Junsong Yuan 0001, Tiejun Huang 0001, Wen Gao 0001
IJCAI4
2016 Multimodal Deep Convolutional Neural Network for Audio-Visual Emotion Recognition
abstract
Emotion recognition is a challenging task because of the emotional gap between subjective emotion and the low-level audio-visual features. Inspired by the recent success of deep learning in bridging the semantic gap, this paper proposes to bridge the emotional gap based on a multimodal Deep Convolution Neural Network (DCNN), which fuses the audio and visual cues in a deep model. This multimodal DCNN is trained with two stages. First, two DCNN models pre-trained on large-scale image data are fine-tuned to perform audio and visual emotion recognition tasks respectively on the corresponding labeled speech and face data. Second, the outputs of these two DCNNs are integrated in a fusion network constructed by a number of fully-connected layers. The fusion network is trained to obtain a joint audio-visual feature representation for emotion recognition. Experimental results on the RML audio-visual database demonstrates the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues in DCNN for emotion recognition. Its success guarantees further research in this direction.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001
ICMR3
2016 CNN vs. SIFT for Image Retrieval: Alternative or Complementary?
abstract
In the past decade, SIFT is widely used in most vision tasks such as image retrieval. While in recent several years, deep convolutional neural networks (CNN) features achieve the state-of-the-art performance in several tasks such as image classification and object detection. Thus a natural question arises: for the image retrieval task, can CNN features substitute for SIFT? In this paper, we experimentally demonstrate that the two kinds of features are highly complementary. Following this fact, we propose an image representation model, complementary CNN and SIFT (CCS), to fuse CNN and SIFT in a multi-level and complementary way. In particular, it can be used to simultaneously describe scene-level, object-level and point-level contents in images. Extensive experiments are conducted on four image retrieval benchmarks, and the experimental results show that our CCS achieves state-of-the-art retrieval results.
Ke Yan 0007, Yaowei Wang 0001, Dawei Liang, Tiejun Huang 0001, Yonghong Tian 0001
ACM Multimedia4
2016 Rate control for consistent video quality with inter-dependent distortion model for HEVC
abstract
Consistent video quality is important for video coding applications, which is also a popular optimization target for rate control. In this paper, a rate control scheme is proposed to reduce the fluctuation of video quality. First, we set up the optimization formulation of distortion and derive the distortion model by analyzing quadtree-based coding unit (CU) structure in high efficiency video coding (HEVC). Then the frame bit allocation algorithm is proposed by considering the inter-dependency. After the frame basic quantization parameter (QP) is obtained, the solution of optimization formulation finally regulates QP to maintain the consistent video quality. Experimental results show that the proposed rate control scheme can reduce the fluctuation of video quality up to 92.9% than benchmark, where the average reduction is 74.1%.
Yuan Li 0014, Huizhu Jia, Tiejun Huang 0001
VCIP4
2016 Fixed-point Gaussian Mixture Model for analysis-friendly surveillance video coding
Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
Comput. Vis. Image Underst.4
2016 Measuring Visual Surprise Jointly from Intrinsic and Extrinsic Contexts for Image Saliency Estimation
Jia Li 0003, Yonghong Tian 0001, Xiaowu Chen 0001, Tiejun Huang 0001
Int. J. Comput. Vis.4
2016 Overview of the MPEG-CDVS Standard
abstract
Compact descriptors for visual search (CDVS) is a recently completed standard from the ISO/IEC moving pictures experts group (MPEG). The primary goal of this standard is to provide a standardized bitstream syntax to enable interoperability in the context of image retrieval applications. Over the course of the standardization process, remarkable improvements were achieved in reducing the size of image feature data and in reducing the computation and memory footprint in the feature extraction process. This paper provides an overview of the technical features of the MPEG-CDVS standard and summarizes its evolution.
Ling-Yu Duan, Vijay Chandrasekhar 0001, Jie Chen 0006, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001, Bernd Girod, Wen Gao 0001
IEEE Trans. Image Process.6
2015 Swiss-System Based Cascade Ranking for Gait-Based Person Re-Identification
abstract
Human gait has been shown to be an efficient biometric measure for person identification at a distance. However, it often needs different gait features to handle various covariate conditions including viewing angles, walking speed, carrying an object and wearing different types of shoes. In order to improve the robustness of gait-based person re-identification on such multi-covariate conditions, a novel Swiss-system based cascade ranking model is proposed in this paper. Since the ranking model is able to learn a subspace where the potential true match is given the highest ranking, we formulate the gait-based person re-identification as a bipartite ranking problem and utilize it as an effective way for multi-feature ensemble learning. Then a Swiss multi-round competition system is developed for the cascade ranking model to optimize its effectiveness and efficiency. Extensive experiments on three indoor and outdoor public datasets demonstrate that our model outperforms several state-of-the-art methods remarkably.
Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
AAAI4
2015 Overview of the MPEG CDVS Standard
abstract
Towards mobile visual search, compact visual descriptors have been well advocated in both academic and industry endeavors. Moving Picture Experts Group (MPEG) initiated the remarkable Compact Descriptors for Visual Search (CDVS) standard activity in Jan. 2010 to push forward the frontiers of compact descriptors in mobile internet industry. In Oct. 2014, MPEG CDVS successfully entered the Final Draft of International Standard. CDVS made a series of significant breakthroughs in high performance and low complexity compact descriptors. In this paper, we give an overview of the MPEG CDVS standard, with emphasis on the development of the core techniques and their technical merits.
Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
DCC2
2015 Optimizing Binary Fisher Codes for Visual Search
abstract
Fisher vectors (FV) aggregated from local invariant features (e.g., SIFT) is one of the state-of-the-art descriptors for visual search, due to high discriminability but small visual vocabulary. Nevertheless, a high-dimensional FV needs to be compressed into a compact descriptor for light storage and high matching eficiency. In this paper, we formulate the FV compression as a resource-constrained optimization problem. Our goal is to maximize search performance subject to the constraints of descriptor compactness, compression complexity in terms of memory usage and time cost. Accordingly, we present a selective binary Fisher codes (SBFC) to compress the raw FV. Firstly, to fulfill the constraint of compression complexity, we binarize the FV by a sign function, Secondly, we propose to select discriminative bits from the binarized FV (BFC) to maximize search performance, subject to the constraint of descriptor compactness. Extensive experiments over MPEG Compact Descriptor for Visual Search (CDVS) benchmark datasets have shown that S-BFC significantly improves search performance at a smaller descriptor size as well as much lower complexity, compared with the state-of-the-art FV compression algorithms like Hashing and Product Quantziation (PQ). A simplified version of SBFC, SBFC LS has been adopted by the MPEG CDVS standard. In the CDVS evaluation framework, SBFC LS has achieved promising performance mean Average Precision (mAP) 83% on average at much lower memory cost of 40KB.
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001
DCC5
2015 An efficient coding framework for compact descriptors extracted from video sequence
abstract
Towards effective and efficient image matching or retrieval tasks, the emerging MPEG standard, named Compact Descriptors for Visual Search (CDVS), has fulfilled compact descriptors for still images, consisting of compressed local and global descriptor. Nevertheless, the frame-level coding of CDVS descriptors from a video sequence does not address the inter-frame redundancy issue, which may consume considerable bandwidth and storage resources. In this work, we propose an efficient coding framework of CDVS descriptors to generate compact descriptors for video sequences. For local descriptors, we propose a multiple reference predictive technique to exploit the temporal correlation of local descriptors and location coordinates over a sequence of frames. To further improve the prediction performance, keypoint tracking is applied to identify temporally repeated keypoints. For global descriptors, a propagation coding way is employed to compress the global descriptors of adjacent frames. The empirical evaluation has shown that the proposed coding approach has yielded a low bit rate of less than 40kbps on average, while maintaining comparable matching and retrieval performance. Compared to the sequence of original frame-level CDVS descriptors, the proposed approach has achieved over 25× bit rate reduction.
Zhangshuai Huang, Ling-Yu Duan, Jie Lin 0001, Shiqi Wang 0001, Siwei Ma 0001, Tiejun Huang 0001
ICIP6
2015 Hierarchical multi-VLAD for image retrieval
abstract
Constructing discriminative feature descriptors is crucial towards effective image retrieval. The state-of-the-art powerful global descriptor for this purpose is Vector of Locally Aggregated Descriptors (VLAD). Given a set of local features (say, SIFT) extracted from an image, the VLAD is generated by quantizing local features with a small visual vocabulary (64 to 512 centroids), aggregating the residual statistics of quantized features for each centroid and concatenating the aggregated residual vectors from each centroid. One can increase the search accuracy by increasing the size of vocabulary (from hundreds to hundreds of thousands), which, however, it leads to heavy computation cost with flat quantization. In this paper, we propose a hierarchical multi-VLAD to seek the tradeoff between descriptor discriminability and computation complexity. We build up a tree-structured hierarchical quantization (TSHQ) to accelerate the VLAD computation with a large vocabulary. As quantization error may propagate from root to leaf node (centroid) with TSHQ, we introduce multi-VLAD, which constructing a VLAD descriptor for each level of the vocabulary tree, so as to compensate for the quantization error at that level. Extensive evaluation over benchmark datasets has shown that the proposed approach outperforms state-of-the-art in terms of retrieval accuracy, fast extraction, as well as light memory cost.
Ling-Yu Duan, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001
ICIP5
2015 Detecting abnormal behaviors in surveillance videos based on fuzzy clustering and multiple Auto-Encoders
abstract
In this paper, we present a novel framework to detect abnormal behaviors in surveillance videos by using fuzzy clustering and multiple Auto-Encoders (FMAE). As detecting abnormal behaviors is often treated as an unsupervised task, how to describe normal patterns becomes the key point. Considering there are many types of normal behaviors in the daily life, we use the fuzzy clustering technique to roughly divide the training samples into several clusters so that each cluster stands for a normal pattern. Then we deploy multiple Auto-Encoders to estimate these different types of normal behaviors from weighted samples. When testing on an unknown video, our framework can predict whether it contains abnormal behaviors or not by summarizing the reconstruction cost through each Auto-Encoder. Since there are always lots of redundancies in the surveillance video, Auto-Encoder is a pretty good tool to capture common structures of normal video sequences automatically as well as estimate normal patterns. The experimental results show that our approach achieves good performance on three public video analysis datasets and statistically outperforms the state-of-the-art approaches under some scenes.
Zhengying Chen, Yonghong Tian 0001, Wei Zeng 0006, Tiejun Huang 0001
ICME4
2015 Learning Deep Trajectory Descriptor for action recognition in videos using deep neural networks
abstract
Human action recognition is widely recognized as a challenging task due to the difficulty of effectively characterizing human action in a complex scene. Recent studies have shown that the dense-trajectory-based methods can achieve state-of-the-art recognition results on some challenging datasets. However, in these methods, each dense trajectory is often represented as a vector of coordinates, consequently losing the structural relationship between different trajectories. To address the problem, this paper proposes a novel Deep Trajectory Descriptor (DTD) for action recognition. First, we extract dense trajectories from multiple consecutive frames and then project them onto a canvas. This will result in a “trajectory texture” image which can effectively characterize the relative motion in these frames. Based on these trajectory texture images, a deep neural network (DNN) is utilized to learn a more compact and powerful representation of dense trajectories. In the action recognition system, the DTD descriptor, together with other non-trajectory features such as HOG, HOF and MBH, can provide an effective way to characterize human action from various aspects. Experimental results show that our system can statistically outperform several state-of-the-art approaches, with an average accuracy of 95:6% on KTH and an accuracy of 92.14% on UCF50.
Yemin Shi 0001, Wei Zeng 0006, Tiejun Huang 0001, Yaowei Wang 0001
ICME3
2015 Hamming Compatible Quantization for Hashing
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001
IJCAI5
2015 Opportunities and Challenges of Global Network Cameras
abstract
Since the introduction of consumer digital cameras, user-created multimedia content has become increasingly popular. Digital cameras, together with inexpensive editing tools, and free hosting sites have made multimedia an integral part of everyday life. Today, hundreds of hours video are uploaded to hosting sites every minute. Video-on-demand through wireless networks and smartphones have profoundly changed how people consume multimedia content. Meanwhile, the widely deployed network cameras can provide live views of many parts of the world. These cameras can provide rich sources creating multimedia content. This panel will explore the opportunities and discuss the challenges using global network cameras for creating multimedia contents and understanding the world. Every year, millions of network cameras are deployed. The data from some of these network cameras are publicly available, continuously streaming live views of national parks, city halls, streets, highways, and shopping malls. A person may see multiple tourist attractions through these cameras, without leaving home. Researchers may observe the weather in different cities. Using the data from the cameras, it is possible to observe natural disasters, such as volcano eruption or tsunami, at a safe distance. News reporters may obtain instant views of an unfolding riot without risking their lives. A spectator may watch a celebration parade from multiple locations using the street cameras. Despite the many promising applications, the opportunities of using global network cameras for creating multimedia content have not been fully exploited.
Joanna Batstone, Touradj Ebrahimi, Tiejun Huang 0001, Yung-Hsiang Lu, Yonggang Wen 0001
ACM Multimedia3
2015 Learning Complementary Saliency Priors for Foreground Object Segmentation in Complex Scenes
Yonghong Tian 0001, Jia Li 0003, Shui Yu 0001, Tiejun Huang 0001
Int. J. Comput. Vis.4
2015 Finding the Secret of Image Saliency in the Frequency Domain
abstract
There are two sides to every story of visual saliency modeling in the frequency domain. On the one hand, image saliency can be effectively estimated by applying simple operations to the frequency spectrum. On the other hand, it is still unclear which part of the frequency spectrum contributes the most to popping-out targets and suppressing distractors. Toward this end, this paper tentatively explores the secret of image saliency in the frequency domain. From the results obtained in several qualitative and quantitative experiments, we find that the secret of visual saliency may mainly hide in the phases of intermediate frequencies. To explain this finding, we reinterpret the concept of discrete Fourier transform from the perspective of template-based contrast computation and thus develop several principles for designing the saliency detector in the frequency domain. Following these principles, we propose a novel approach to design the saliency detector under the assistance of prior knowledge obtained through both unsupervised and supervised learning processes. Experimental results on a public image benchmark show that the learned saliency detector outperforms 18 state-of-the-art approaches in predicting human fixations.
Jia Li 0003, Ling-Yu Duan, Xiaowu Chen 0001, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Robust multiple cameras pedestrian detection with multi-view Bayesian network
Peixi Peng, Yonghong Tian 0001, Yaowei Wang 0001, Jia Li 0003, Tiejun Huang 0001
Pattern Recognit.5
2015 Image saliency estimation via random walk guided by informativeness and latent signal correlations
Jia Li 0003, Shu Fang, Yonghong Tian 0001, Tiejun Huang 0001, Xiaowu Chen 0001
Signal Process. Image Commun.4
2015 A Low Complexity Interest Point Detector
abstract
Interest point detection is a fundamental approach to feature extraction in computer vision tasks. To handle the scale invariance, interest points usually work on the scale-space representation of an image. In this letter, we propose a novel block-wise scale-space representation to significantly reduce the computational complexity of an interest point detector. Laplacian of Gaussian (LoG) filtering is applied to implement the block-wise scale-space representation. Extensive comparison experiments have shown the block-wise scale-space representation enables the efficient and effective implementation of an interest point detector in terms of memory and time complexity reduction, as well as promising performance in visual search.
Jie Chen 0006, Ling-Yu Duan, Feng Gao 0014, Jianfei Cai 0001, Alex Chichung Kot, Tiejun Huang 0001
IEEE Signal Process. Lett.6
2015 Depth-Preserving Warping for Stereo Image Retargeting
abstract
The popularity of stereo images and various display devices poses the need of stereo image retargeting techniques. Existing warping-based retargeting methods can well preserve the shape of salient objects in a retargeted stereo image pair. Nevertheless, these methods often incur depth distortion, since they attempt to preserve depth by maintaining the disparity of a set of sparse correspondences, rather than directly controlling the warping. In this paper, by considering how to directly control the warping functions, we propose a warping-based stereo image retargeting approach that can simultaneously preserve the shape of salient objects and the depth of 3D scenes. We first characterize the depth distortion in terms of warping functions to investigate the impact of a warping function on depth distortion. Based on the depth distortion model, we then exploit binocular visual characteristics of stereo images to derive region-based depth-preserving constraints which directly control the warping functions so as to faithfully preserve the depth of 3D scenes. Third, with the region-based depth-preserving constraints, we present a novel warping-based stereo image retargeting framework. Since the depth-preserving constraints are derived regardless of shape preservation, we relax the depth-preserving constraints to fulfill a tradeoff between shape preservation and depth preservation. Finally, we propose a quad-based implementation of the proposed framework. The results demonstrate the efficacy of our method in both depth and shape preservation for stereo image retargeting.
Bing Li 0024, Ling-Yu Duan, Chia-Wen Lin, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Image Process.4
2015 Weighted Component Hashing of Binary Aggregated Descriptors for Fast Visual Search
abstract
Towards low bit rate mobile visual search, recent works have proposed to aggregate the local features and compress the aggregated descriptor (such as Fisher vector, the vector of locally aggregated descriptors) for low latency query delivery as well as moderate search complexity. Even though Hamming distance can be computed very fast, the computational cost of exhaustive linear search over the binary descriptors grows linearly with either the length of a binary descriptor or the number of database images. In this paper, we propose a novel weighted component hashing (WeCoHash) algorithm for long binary aggregated descriptors to significantly improve search efficiency over a large scale image database. Accordingly, the proposed WeCoHash has attempted to address two essential issues in Hashing algorithms: “what to hash” and “how to search.” “What to hash” is tackled by a hybrid approach, which utilizes both image-specific component (i.e., visual word) redundancy and bit dependency within each component of a binary aggregated descriptor to produce discriminative hash values for bucketing. “How to search” is tackled by an adaptive relevance weighting based on the statistics of hash values. Extensive comparison results have shown that WeCoHash is at least 20 times faster than linear search and 10 times faster than local sensitive hash (LSH) when maintaining comparable search accuracy. In particular , the WeCoHash solution has been adopted by the emerging MPEG compact descriptor for visual search (CDVS) standard to significantly speed up the exhaustive search of the binary aggregated descriptors.
Ling-Yu Duan, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Multim.4
2015 TASC: A Transformation-Aware Soft Cascading Approach for Multimodal Video Copy Detection
abstract
How to precisely and efficiently detect near-duplicate copies with complicated audiovisual transformations from a large-scale video database is a challenging task. To cope with this challenge, this article proposes a transformation-aware soft cascading (TASC) approach for multimodal video copy detection. Basically, our approach divides query videos into some categories and then for each category designs a transformation-aware chain to organize several detectors in a cascade structure. In each chain, efficient but simple detectors are placed in the forepart, whereas effective but complex detectors are located in the rear. To judge whether two videos are near-duplicates, a Detection-on-Copy-Units mechanism is introduced in the TASC, which makes the decision of copy detection depending on the similarity between their most similar fractions, called copy units (CUs), rather than the video-level similarity. Following this, we propose a CU search algorithm to find a pair of CUs from two videos and a CU-based localization algorithm to find the precise locations of their copy segments that are with the asserted CUs as the center. Moreover, to address the problem that the copies and noncopies are possibly linearly inseparable in the feature space, the TASC also introduces a flexible strategy, called soft decision boundary , to replace the single threshold strategy for each detector. Its basic idea is to automatically learn two thresholds for each detector to examine the easy-to-judge copies and noncopies, respectively, and meanwhile to train a nonlinear classifier to further check those hard-to-judge ones. Extensive experiments on three benchmark datasets showed that the TASC can achieve excellent copy detection accuracy and localization precision with a very high processing efficiency.
Yonghong Tian 0001, Mengren Qian, Tiejun Huang 0001
ACM Trans. Inf. Syst.3
2014 Hybrid-Indexing Multi-type Features for Large-Scale Image Search
Qingjun Luo, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001
ACCV (1)3
2014 Joint optimization of JPEG quantization table and coefficient thresholding for low bitrate mobile visual search
abstract
Low latency query delivery over wireless network is a key problem for mobile visual search. Extracting compact descriptors directly on the mobile device is computational expensive, an alternate approach is to send highly compressed JPEG query images. As JPEG baseline optimizes the rate-distortion from a perceptual perspective rather than maintaining search performance, recent work proposed to learn a feature-preserving JPEG quantization table for improved search accuracy. However, this method is data-dependent and the quantization table cannot adapt to image blocks. To address these issues, we propose to jointly optimize the JPEG quantization table and coefficient thresholding. The matching score between uncompressed image and its compressed JPEG image is employed as the distortion measure to avoid time consuming image labeling, and coefficient thresholding eliminates the redundant coefficients. Extensive experiments on benchmark datasets show that our approach obtains superior performance than state-of-the-art at low bitrates, meanwhile, it consumes lower cost including processing time, memory and battery on mobile device.
Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001
ICIP4
2014 Component hashing of variable-length binary aggregated descriptors for fast image search
abstract
Compact locally aggregated binary features have shown great advantages in image search. As the exhaustive linear search in Hamming space still entails too much computational complexity for large datasets, recent works proposed to directly use binary codes as hash indices, yielding a dramatic increase in speedup. However, these methods cannot be directly applied to variable-length binary features. In this paper, we propose a Component Hashing (CoHash) algorithm to handle the variable-length binary aggregated descriptors indexing for fast image search. The main idea is to decompose the distance measure between variable-length descriptors into aligned component-to-component matching problems independently, and build multiple hash tables for the visual word components. Given a query, its candidate neighbors are found by using the query binary sub-vectors as indices into their corresponding hash tables. In particular, a bit selection based on conditional mutual information maximization is proposed to reduce the dimensionality of visual word components, which provides a light storage of indices and balances the retrieval accuracy and search cost. Extensive experiments on benchmark datasets show that our approach is 20~25 times faster than linear search, without any noticeable retrieval performance loss.
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001, Miroslaw Bober
ICIP4
2014 Multi-view gait recognition with incomplete training data
abstract
Changes in the viewing angles pose a major challenge for gait recognition because the human gait silhouettes can be different under the various viewing angles. Recently, View Transformation Model (VTM) was proposed to tackle this problem by transforming gait features from across views to a common viewing angle. However, VTM must use the data of subjects crossing all views to train the pre-constructed model, which might be unsuitable for the real applications. To address this problem, this paper proposes a View Feature Recovering Model (VFRM) to generate the VTM with incomplete training data. In our algorithm, if the gait signature of a pedestrian is missing under a view, it can be recovered from the K-nearest pedestrians whose gait features are available in the same view. Moreover, the Geodesic distance based K-Nearest Neighbor (GKNN) algorithm is adopted in our algorithm to better measure the neighborhood between two pedestrians. Experimental results on a benchmark database has demonstrated the effectiveness of our method.
Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
ICME4
2014 Background-foreground division based search for motion estimation in surveillance video coding
abstract
Basically, motion search is very time-consuming in the process of video coding. For surveillance videos, however, there exist a large amount of static background regions whose motion vectors actually are equal to zero. By utilizing the background and foreground information of coding units, this paper proposes a background-foreground division based search algorithm (BFDS) to accelerate the motion search in surveillance video coding. The basic idea of BFDS is to classify a predicting unit (PU) into a background predicting unit (BPU) or a foreground predicting unit (FPU) and then adopt different search strategies respectively for BPUs and FPUs. That is, a zero motion vector biased search strategy is applied in BPUs to reduce the search complexity on a large scale while a precise global search strategy is applied in FPUs to get higher coding performance. Compared with the current TZ search algorithm used in HEVC, the proposed BFDS algorithm can reduce the number of search points by 57.73% while remaining the coding performance almost unchanged.
Yonghong Tian 0001, Tiejun Huang 0001
ICME3
2014 Superimage: Packing Semantic-Relevant Images for Indexing and Retrieval
abstract
As an important procedure in image retrieval, off-line indexing focuses on organizing relevant images together and making them easy to access. However, most of existing indexing strategies view database images individually and only consider partial relevance, i.e., either visual or semantic relevance among them. To overcome these issues and design better indexing strategy, we propose to package semantically relevant images into superimages, and then index superimages instead of single images. Superimage effectively packages multiple images into one new unit, hence significantly decreases the number of images to be indexed. This naturally saves the memory cost and retrieval time. To make the final index file discriminative to both visual and semantic relevances, we extract local descriptors from superimages and index them with inverted file. During online retrieval, we only need to extract local descriptors from queries, but could get semantic-aware retrieval results. This is because during our off-line indexing stage, both the semantically and visually relevant images are organized together. Therefore, our approach is superior to many online retrieval fusion algorithms. Experimental results on UKbench, Holidays, and one large-scale dataset all manifest the promising performance of our approach, i.e., competitive precision, better efficiency, and only about 1/2 memory consumption compared with state-of-the-arts.
Qingjun Luo, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001
ICMR3
2014 Visual Saliency with Statistical Priors
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001
Int. J. Comput. Vis.3
2014 Mining Compact Bag-of-Patterns for Low Bit Rate Mobile Visual Search
abstract
Visual patterns, i.e., high-order combinations of visual words, contributes to a discriminative abstraction of the high-dimensional bag-of-words image representation. However, the existing visual patterns are built upon the 2D photographic concurrences of visual words, which is ill-posed comparing with their real-world 3D concurrences, since the words from different objects or different depth might be incorrectly bound into an identical pattern. On the other hand, designing compact descriptors from the mined patterns is left open. To address both issues, in this paper, we propose a novel compact bag-of-patterns (CBoPs) descriptor with an application to low bit rate mobile landmark search. First, to overcome the ill-posed 2D photographic configuration, we build up a 3D point cloud from the reference images of each landmark, therefore more accurate pattern candidates can be extracted from the 3D concurrences of visual words. A novel gravity distance metric is then proposed to mine discriminative visual patterns. Second, we come up with compact image description by introducing a CBoPs descriptor. CBoP is figured out by sparse coding over the mined visual patterns, which maximally reconstructs the original bag-of-words histogram with a minimum coding length. We developed a low bit rate mobile landmark search prototype, in which CBoP descriptor is directly extracted and sent from the mobile end to reduce the query delivery latency. The CBoP performance is quantized in several large-scale benchmarks with comparisons to the state-of-the-art compact descriptors, topic features, and hashing descriptors. We have reported comparable accuracy to the million-scale bag-of-words histogram over the million scale visual words, with high descriptor compression rate (approximately 100-bits) than the state-of-the-art bag-of-words compression scheme.
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Image Process.4
2014 Background-Modeling-Based Adaptive Prediction for Surveillance Video Coding
abstract
The exponential growth of surveillance videos presents an unprecedented challenge for high-efficiency surveillance video coding technology. Compared with the existing coding standards that were basically developed for generic videos, surveillance video coding should be designed to make the best use of the special characteristics of surveillance videos (e.g., relative static background). To do so, this paper first conducts two analyses on how to improve the background and foreground prediction efficiencies in surveillance video coding. Following the analysis results, we propose a background-modeling-based adaptive prediction (BMAP) method. In this method, all blocks to be encoded are firstly classified into three categories. Then, according to the category of each block, two novel inter predictions are selectively utilized, namely, the background reference prediction (BRP) that uses the background modeled from the original input frames as the long-term reference and the background difference prediction (BDP) that predicts the current data in the background difference domain. For background blocks, the BRP can effectively improve the prediction efficiency using the higher quality background as the reference; whereas for foreground-background-hybrid blocks, the BDP can provide a better reference after subtracting its background pixels. Experimental results show that the BMAP can achieve at least twice the compression ratio on surveillance videos as AVC (MPEG-4 Advanced Video Coding) high profile, yet with a slightly additional encoding complexity. Moreover, for the foreground coding performance, which is crucial to the subjective quality of moving objects in surveillance videos, BMAP also obtains remarkable gains over several state-of-the-art methods.
Xianguo Zhang, Tiejun Huang 0001, Yonghong Tian 0001, Wen Gao 0001
IEEE Trans. Image Process.2
2014 Optimizing the Hierarchical Prediction and Coding in HEVC for Surveillance and Conference Videos With Background Modeling
abstract
For the real-time and low-delay video surveillance and teleconferencing applications, the newly video coding standard HEVC can achieve much higher coding efficiency over H.264/AVC. However, we still argue that the hierarchical prediction structure in the HEVC low-delay encoder still does not fully utilize the special characteristics of surveillance and conference videos that are usually captured by stationary cameras. In this case, the background picture (G-picture), which is modeled from the original input frames, can be used to further improve the HEVC low-delay coding efficiency meanwhile reducing the complexity. Therefore, we propose an optimization method for the hierarchical prediction and coding in HEVC for these videos with background modeling. First, several experimental and theoretical analyses are conducted on how to utilize the G-picture to optimize the hierarchical prediction structure and hierarchical quantization. Following these results, we propose to encode the G-picture as the long-term reference frame to improve the background prediction, and then present a G-picture-based bit-allocation algorithm to increase the coding efficiency. Meanwhile, according to the proportions of background and foreground pixels in coding units (CUs), an adaptive speed-up algorithm is developed to classify each CU into different categories and then adopt different speed-up strategies to reduce the encoding complexity. To evaluate the performance, extensive experiments are performed on the HEVC test model. Results show our method can averagely save 39.09% bits and reduce the encoding complexity by 43.63% on surveillance videos, whereas those are 5.27% and 43.68% on conference videos.
Xianguo Zhang, Yonghong Tian 0001, Tiejun Huang 0001, Siwei Dong, Wen Gao 0001
IEEE Trans. Image Process.3
2014 Towards Mobile Document Image Retrieval for Digital Library
abstract
With the proliferation of mobile devices, recent years have witnessed an emerging potential to integrate mobile visual search techniques into digital library. Such a mobile application scenario in digital library has posed significant and unique challenges in document image search. The mobile photograph makes it tough to extract discriminative features from the landmark regions of documents, like line drawings, as well as text layouts. In addition, both search scalability and query delivery latency remain challenging issues in mobile document search. The former relies on an effective yet memory-light indexing structure to accomplish fast online search, while the latter puts a bit budget constraint of query images over the wireless link. In this paper, we propose a novel mobile document image retrieval framework, consisting of a robust Local Inner-distance Shape Context (LISC) descriptor of line drawings, a Hamming distance KD-Tree for scalable and memory-light document indexing, as well as a JBIG2 based query compression scheme, together with a Retinex based enhancement and an OTSU based binarization, to reduce the latency of delivering query while maintaining query quality in terms of search performance. We have extensively validated the key techniques in this framework by quantitative comparison to alternative approaches.
Ling-Yu Duan, Rongrong Ji, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Multim.4
2013 Hierarchical-and-Adaptive Bit-Allocation with Selective Background Prediction for High Efficiency Video Coding (HEVC)
abstract
Summary form only given. Recently, a low-delay and high-efficiency hierarchical prediction structure (HPS) has been proposed for the forthcoming HEVC. Actually, frames and coding units (CUs) at different HPS positions have different importance to predict following frames and CUs. This paper firstly analyzes what frames and CUs should be quantified less. Based on the analysis, we propose a Hierarchical-and-Adaptive BIT-allocation method with Selective background prediction (HABITS) to optimize the video performance of HEVC. Extensive experiments on HM8.0 show that, HABITS saves 13.3% and 35.5% of the total bit rate for eight HEVC conference videos and eight common used surveillance videos. Even for the normal videos in HEVC's Class B and C, there is still 2.2% bit-saving.
Xianguo Zhang, Tiejun Huang 0001, Yonghong Tian 0001, Wen Gao 0001
DCC2
2013 On the interoperability of local descriptors compression
abstract
There are a number of component technologies that are useful for visual search, including format of visual descriptors, descriptor extraction process, as well as indexing, and matching algorithms. As a minimum, the format of descriptors as well as parts of their extraction process should be defined to ensure interoperability. In this paper, we study the problem of interoperability among compressed local descriptors at different bit-rates; that is, allowing effective and efficient comparison of compact descriptors, which is fundamentally important to mobile visual search applications. We propose to combine feature transform and multi-stage vector quantization to implement the interoperability of compact local descriptors. First, an orthogonal transform (e.g. Principle component analysis, PCA) is employed to eliminate the correlation between local feature dimensions, which improves the performance of compressed domain descriptor matching with the well-aligned distance computing of sorted important features in transform space. Second, a multi-stage vector quantization (MSVQ) is applied to generate compact codes for local descriptors. At light quantization tables, MSVQ takes advantage of the transform domain features to properly allocate different budgets to each group of transformed feature dimensions, respectively. The interoperability between compressed descriptors at different bit rates can be achieved by the descriptors' fast matching in the orthogonal feature space. In other words, descriptor decoding into the original feature space (SIFT space) is unnecessary, as the distance can be calculated by pre-computed lookup tables. In particular, such efficient matching in transform domain is significant for large-scale visual search. Over a set of benchmark datasets, we have reported superior performance over state-of-the-arts.
Jie Chen 0006, Ling-Yu Duan, Jie Lin 0001, Rongrong Ji, Tiejun Huang 0001, Wen Gao 0001
ICASSP5
2013 Robust fisher codes for large scale image retrieval
abstract
Fisher vectors (FV) have shown great advantages in large scale visual search. However, traditional FV suffers from noisy local descriptors, which may deteriorate the FV discriminative power. In this paper, we propose a robust Fisher vectors (RFV). To fulfill fast search and light storage over a large scale image dataset, we employ a simple binarization method to compress RFV to generate compact robust Fisher codes (RFC). Extensive comparison experiments on benchmark datasets have shown that both RFV and RFC outperforms the state-of-the-art performance. The scalability of RFC has been validated on a dataset of over 1 million images as well.
Jie Lin 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
ICASSP3
2013 A system based on sequence learning for event detection in surveillance video
abstract
Event detection in crowded surveillance videos is a challenging yet important problem. In this paper, we present our eSur (Event detection system on SURveillance video) system, which is derived from TRECVid'12 surveillance tasks. Currently, eSur attempts to detect two categories of events: 1) pair-wise events (e.g., PeopleMeet, PeopleSplitUp and Embrace); 2) action-like events (e.g., ObjectPut, CellToEar, PersonRuns and Pointing). In eSur system, we first employ people detection and tracking algorithms to locate target persons in 3D space-time domain. Then the video sequences in which target persons occur are partitioned into several spatio-temporal cubes. Visual features (i.e. cubic feature and MoSIFT) are computed over these cubes. After that, a sequence learning method, (namely SVM with dynamic time alignment kernel), is employed to infer the existence of an event for the video sequence. According to the TRECVid SED formal evaluation, eSur has yielded fairly encouraging results on TRECVid'12 dataset.
Xiaoyu Fang, Ziwei Xia, Chi Su, Teng Xu 0002, Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
ICIP7
2013 A novel pair-wise image matching strategy with compact descriptors
abstract
In this paper, we address the problem of pair-wise image matching which determines whether two images depict the same objects or scenes. SIFT-like local descriptor-based matching is the most widely adopted method for this purpose and has achieved the state-of-the-art performance. However, local descriptor-based methods usually fail when an image pair contains multiple similar local regions. This problem becomes more serious when coming to limited computational and storage resources. Although global descriptors, e.g., Fisher Vectors, can solve this issue, it is difficult for global descriptors to distinguish images containing different objects of the same class. Therefore, we propose a novel strategy to integrate local and global descriptors for better matching accuracy. To further fulfill the efficiency requirement of applications, we combine dimension reduction and product quantization to obtain compact descriptors and speed up the matching process with pre-computed lookup tables. Extensive comparisons to the state-of-the-art methods demonstrate our advantages in both matching accuracy and efficiency.
Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001
ICIP4
2013 Overview of the IEEE 1857 surveillance groups
abstract
Among the multiple application-oriented video groups of IEEE 1857 video part, surveillance groups are the first specific video coding standards targeting on the exploring surveillance system. In this paper, we firstly present an overview of the technical features and characteristics of the Surveillance Baseline and Surveillance Groups. The video coding technologies are then described in greater detail on three main directions, including the background modeling based prediction techniques for high-efficiency surveillance video coding, error resilience methods for channel-noisy surveillance video transmission and the high-level syntax for surveillance video analysis. The surveillance groups can provide a good support for kinds of video analysis applications of computer vision and make the video transmission more applicable for noisy channels. Moreover, experimental results show that the background modeling based prediction techniques can well exploit the special characteristics of surveillance video and double the traditional compression performance.
Xianguo Zhang, Tiejun Huang 0001, Yonghong Tian 0001, Wen Gao 0001
ICIP2
2013 MPLBoost-based mixture model for effective human detection with Deformable Part Model
abstract
The Deformable Part Model has shown high accuracy in tackling certain occlusion or deformations of objects such as cars and bikes. However, as for human category characterized by a larger number of articulated parts and more significant appearance variations, its performance gain is not so remarkable. To address this issue, we propose an MPLBoost-based mixture model which splits data into coherent groups and trains one root classifier for each, resulting in automated selection of discriminative root models and better representation of intra-class variations through visual feature clustering. Based on this boosting framework, multiple complementary features are combined to capture shape, texture and color information. Experimental results demonstrate that the proposed model can achieve an impressive performance improvement, especially in handling larger variations of human poses and viewpoints.
Chaoran Gu, Luntian Mou, Yonghong Tian 0001, Tiejun Huang 0001
ICME4
2013 A background proportion adaptive Lagrange multiplier selection method for surveillance video on HEVC
abstract
In the recent video coding standards, the selection of Lagrange multiplier is crucial to achieve trade-off between the choices of low-distortion and low-bitrate prediction modes. For surveillance video coding, the rate-distortion analysis shows that, a larger Lagrange multiplier should be used if the background in a coding unit took a larger proportion. Therefore, a modified Lagrange multiplier might be better for rate-distortion optimization. To address this problem, we perform an in-depth analysis on the relationship between the optimal Lagrange multiplier and the background proportion, and then propose a Lagrange multiplier selection model to obtain the optimal coding performance for surveillance videos. Following this, we further develop a Lagrange multiplier optimized video coding method. Experimental results show that our coding method can averagely achieve 18.07% bitrate saving on CIF sequences and 11.88% on SD sequences against the background-irrelevant Lagrange multiplier selection method.
Xianguo Zhang, Yonghong Tian 0001, Ronggang Wang, Tiejun Huang 0001
ICME5
2013 Compact descriptors for mobile visual search and MPEG CDVS standardization
abstract
In this paper, we present the state-of-the-art compact descriptors for mobile visual search. In particular, we introduce our MPEG contributions in global descriptor aggregation and local descriptor compression, which have been adopted by the ongoing MPEG standardization of compact descriptor for visual search (CDVS). Standardization progress will be introduced. Other issues including visual object databases and MPEG CDVS impact on visual search industry will be discussed as well.
Ling-Yu Duan, Feng Gao 0014, Jie Chen 0006, Jie Lin 0001, Tiejun Huang 0001
ISCAS5
2013 Single underwater image enhancement with a new optical model
abstract
As light is attenuated when disseminating in water, the clarity of images or videos captured under water is usually degraded to varying degrees. By exploring the difference in light attenuation between in atmosphere and in water, we derive a new underwater optical model to describe the formation of an underwater image in the true physical process, and then propose an effective enhancement algorithm with the derived optical model to improve the perception of underwater images or video frames. In our algorithm, a new underwater dark channel is derived to estimate the scattering rate, and an effective method is also presented to estimate the background light in the underwater optical model. Experimental results show that our algorithm can well handle underwater images, especially for deep-sea images and those captured from turbid waters.
Haocheng Wen, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ISCAS3
2013 Surveillance video coding with quadtree partition based ROI extraction
abstract
To reduce the surveillance video coding cost, it is intuitive to encode surveillance videos by dealing with the foreground objects and the background separately. One widely used method following this strategy is Region-of-Interest (ROI) based coding. To achieve significant improvement for the coding efficiency of ROI based methods, this paper presents a surveillance video coding method with High Efficiency Video Coding (HEVC) quadtree partition based ROI extraction. With automatically generated foreground mask and modeled background frame, a ROI extraction following the block partition in HEVC's quadtree structure is firstly performed. Afterwards, surveillance videos can be compressed by coding two-layer videos. One is the ROI-layer video generated by merging ROIs and background data in each frame together. The other is the background-layer video produced by subtracting the ROIs from the original input video. Results show our method can achieve remarkable total bit-rate saving and significant bit-rate cost reduction on ROIs.
Peiyin Xing, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
PCS3
2013 Wavelet based smoke detection method with RGB Contrast-image and shape constrain
abstract
Smoke detection in video surveillance is very important for early fire detection. A general viewpoint assumes that smoke is a low frequency signal which may smoothen the background. However, some pure-color objects also have this characteristic, and smoke also produces high frequency signal because the rich edge information of its contour. In order to solve these problems, an improved smoke detection method with RGB Contrast-image and shape constrain is proposed. In this method, wavelet transformation is implemented based on the RGB Contrast-image to distinguish smoke from other low frequency signals, and the existence of smoke is determined by analyzing the combination of the shape and the energy change of the region. Experimental results show our method outperforms the conventional methods remarkably.
Jiaqiu Chen, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001
VCIP4
2013 A coding unit classification based AVC-to-HEVC transcoding with background modeling for surveillance videos
abstract
To save the storage and transmission cost, it is applicable now to develop fast and efficient methods to transcode the perennial surveillance videos to HEVC ones, since HEVC has doubled the compression ratio. Considering the long-time static background characteristic of surveillance videos, this paper presents a coding unit (CU) classification based AVC-to-HEVC transcoding method with background modeling. In our method, the background frame modeled from originally decoded frames is firstly transcoded into HEVC stream as long-term reference to enhance the prediction efficiency. Afterwards, a CU classification algorithm which employs decoded motion vectors and the modeled background frame as input is proposed to divide the decoded data into background, foreground and hybrid CUs. Following this, different transcoding strategies of CU partition termination, prediction unit candidate selection and motion estimation simplification are adopted for different CU categories to reduce the complexity. Experimental results show our method can achieve 45% bit saving and 50% complexity reduction against traditional AVC-to-HEVC transcoding.
Peiyin Xing, Yonghong Tian 0001, Xianguo Zhang, Yaowei Wang 0001, Tiejun Huang 0001
VCIP5
2013 A local shape descriptor for mobile linedrawing retrieval
abstract
Coming with the rapid spread of Intelligent terminals with camera, mobile visual search techniques have undergone a revolution, where visual information can be easily browsed and retrieved upon simply capturing a query photo. However, most existing work targets at compact description of natural scene image statistics, while dealing with line drawing images retains an open problem. This paper presents a unified framework of line drawing problems in mobile visual search. We propose a compact description of line drawing image named Local Inner-Distance Shape Context (LISC) which is robust to the distortion and occlusion and enjoys scale and rotation invariance. Together with an innovative compression scheme using JBIG2 to reduce query delivery latency, our framework works well on both a self-built dataset and MPEG- 7 CE Shape-1 dataset. Promising results on both datasets show significant improvement over state-of-the-art algorithms.
Yucong Xuan, Ling-Yu Duan, Tiejun Huang 0001
VCIP3
2013 Learning from mobile contexts to minimize the mobile location search latency
Ling-Yu Duan, Rongrong Ji, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001
Signal Process. Image Commun.5
2013 Estimating Visual Saliency Through Single Image Optimization
abstract
This letter presents a novel approach for visual saliency estimation through single image optimization. Instead of directly mapping visual features to saliency values with a unified model, we treat regional saliency values as the optimization objective on each single image. By using a quadratic programming framework, our approach can adaptively optimize the regional saliency values on each specific image to simultaneously meet multiple saliency hypotheses on visual rarity, center-bias and mutual correlation. Experimental results show that our approach can outperform 14 state-of-the-art approaches on a public image benchmark.
Jia Li 0003, Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001
IEEE Signal Process. Lett.4
2013 Selective Eigenbackground for Background Modeling and Subtraction in Crowded Scenes
abstract
Background subtraction is a fundamental preprocessing step in many surveillance video analysis tasks. In spite of significant efforts, however, background subtraction in crowded scenes remains challenging, especially, when a large number of foreground objects move slowly or just keep still. To address the problem, this paper proposes a selective eigenbackground method for background modeling and subtraction in crowded scenes. The contributions of our method are three-fold: First, instead of training eigenbackgrounds using the original video frames that may contain more or less foregrounds, a virtual frame construction algorithm is utilized to assemble clean background pixels from different original frames so as to construct some virtual frames as the training and update samples. This can significantly improve the purity of the trained eigenbackgrounds. Second, for a crowded scene with diversified environmental conditions (e.g., illuminations), it is difficult to use only one eigenbackground model to deal with all these variations, even using some online update strategies. Thus given several models trained offline, we utilize peak signal-to-noise ratio to adaptively choose the optimal one to initialize the online eigenbackground model. Third, to tackle the problem that not all pixels can obtain the optimal results when the reconstruction is performed at once for the whole frame, our method selects the best eigenbackground for each pixel to obtain an improved quality of the reconstructed background image. Extensive experiments on the TRECVID-SED dataset and the Road video dataset show that our method outperforms several state-of-the-art methods remarkably.
Yonghong Tian 0001, Yaowei Wang 0001, Zhipeng Hu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2013 Fast and Efficient Transcoding Based on Low-Complexity Background Modeling and Adaptive Block Classification
abstract
It is in urgent need to develop fast and efficient transcoding methods so as to remarkably save the storage of surveillance videos and synchronously transmit conference videos over different bandwidths. Towards this end, the special characteristics of these videos, e.g., the relatively static background, should be utilized for transcoding. Therefore, we propose a fast and efficient transcoding method (FET) based on background modeling and block classification in this paper. To improve the transcoding efficiency, FET adds the background picture, which is modeled from the originally decoded frames in low complexity, into stream in the form of an intra-coded G-picture. And then, FET utilizes the reconstructed G-picture as the long-term reference frame to transcode the following frames. This is mainly because our theoretical analyses show that G-picture can significantly improve the transcoding performance. To reduce the complexity, FET utilizes an adaptive threshold updating model for block classification and then adopts different transcoding strategies for different categories. This is due to the following statistics: after dividing blocks into categories of foreground, background and hybrid ones, different block categories have different distributions of prediction modes, motion vectors and reference frames. Extensive experiments on transcoding high-bit-rate H.264/AVC streams to low-bit-rate ones are carried out to evaluate our FET. Over the traditional full-decoding-and-full-encoding methods, FET can save more than 35% of the transcoding bit-rate with a speed-up ratio of larger than 10 on the surveillance videos. On the conference videos which should be transcoded more timely, FET achieves more than 20 times speed-up ratio with 0.2 dB gain.
Xianguo Zhang, Tiejun Huang 0001, Yonghong Tian 0001, Mingchao Geng, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.2
2013 Content-based copy detection through multimodal feature representation and temporal pyramid matching
abstract
Content-based copy detection (CBCD) is drawing increasing attention as an alternative technology to watermarking for video identification and copyright protection. In this article, we present a comprehensive method to detect copies that are subjected to complicated transformations. A multimodal feature representation scheme is designed to exploit the complementarity of audio features, global and local visual features so that optimal overall robustness to a wide range of complicated modifications can be achieved. Meanwhile, a temporal pyramid matching algorithm is proposed to assemble frame-level similarity search results into sequence-level matching results through similarity evaluation over multiple temporal granularities. Additionally, inverted indexing and locality sensitive hashing (LSH) are also adopted to speed up similarity search. Experimental results over benchmarking datasets of TRECVID 2010 and 2009 demonstrate that the proposed method outperforms other methods for most transformations in terms of copy detection accuracy. The evaluation results also suggest that our method can achieve competitive copy localization preciseness.
Luntian Mou, Tiejun Huang 0001, Yonghong Tian 0001, Menglin Jiang, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2012 Automatic Webcam-Based Human Heart Rate Measurements Using Laplacian Eigenmap
Yonghong Tian 0001, Yaowei Wang 0001, Touradj Ebrahimi, Tiejun Huang 0001
ACCV (2)5
2012 An Efficient Background Reconstruction Based Coding Method for Surveillance Videos Captured by Moving Camera
abstract
With the proliferation of moving surveillance cameras, how to effectively compress videos captured from them is becoming more and more important. One significant characteristic is that, these cameras always go and return cyclically within a limited area. Thus we propose to dynamically build up a background frame for each input frame from a generated panorama background and employ it for a background frame based motion compensation to improve the coding efficiency. For the background reconstruction procedure, we firstly extract limited number of feature point pairs between the robustly searched area in the decoded panorama and the current frame. Afterwards, the global motion transformation matrix is obtained to rectify the searched area into a projective plane of the current frame, and then the reconstructed background is produced. Experiments on six in-door and out-door surveillance videos show that, the background reconstruction based coding method achieves significant performance gain.
Shumin Han, Xianguo Zhang, Yonghong Tian 0001, Tiejun Huang 0001
AVSS4
2012 Single and Multiple View Detection, Tracking and Video Analysis in Crowded Environments
abstract
In this paper, we present our detection, tracking and event recognition methods and the results for PETS 2012. First, ROIs (Regions of Interest) based on geometric constraints are utilized in single view detection to eliminate the negative influence of clutter environment. Then, an optimized observation model is applied to address the ID switching or tracking drifting problem in single view tracking. Third, we introduce the multi-view Bayesian network (MBN) to reduce the "phantom" phenomena which frequently happen in general multi-view detection tasks. At last, a motion-based event recognition method is proposed to handle the event recognition task. Experimental results on the PETS 2012 dataset indicate that our methods are very promising.
Teng Xu 0002, Peixi Peng, Xiaoyu Fang, Chi Su, Yaowei Wang 0001, Yonghong Tian 0001, Wei Zeng 0006, Tiejun Huang 0001
AVSS8
2012 Multi-camera Pedestrian Detection with Multi-view Bayesian Network Model
Peixi Peng, Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
BMVC4
2012 A Fast and Performance-Maintained Transcoding Method Based on Background Modeling for Surveillance Video
abstract
Low-complexity and high-performance surveillance video Transcoding methods play an important role for a wide range of surveillance video transmission and storage applications. Towards this end, the special characteristics of surveillance video should be utilized for Transcoding. In this paper, we propose a fast and performance-maintained Transcoding method. This method firstly divides macro blocks (MBs) into foreground MBs, foreground border MBs and background MBs. Statistics show that the three categories have different distributions of prediction modes, motion vectors and reference frames. Following this, we adopt different Transco ding strategies in terms of removing the redundant prediction modes, narrowing motion search range and reducing reference frames. In particular, we propose an algorithm to exploit the decoded motion vector to adaptively calculate motion search range. Experimental results show that, compared with the recent background modeling based full-decoding-full-encoding, our Transcoding method saves more than 93% time with ignorable quality loss.
Mingchao Geng, Xianguo Zhang, Yonghong Tian 0001, Luhong Liang, Tiejun Huang 0001
ICME5
2012 Video Copy Detection Using a Soft Cascade of Multimodal Features
abstract
In the video copy detection task, it is widely recognized that none of any single feature can work well for all transformations. Thus more and more approaches adopt a set of complementary features to cope with complex audio-visual transformations. However, most of them utilize individual features separately and the final result is obtained by fusing results of several basic detectors. Often, this will lead to low detection efficiency. Moreover, there are some thresholds or parameters to be elaborately tuned. To address these problems, we propose a soft cascade approach to integrate multiple features for efficient copy detection. In our approach, basic detectors are organized in a cascaded framework, which processes a query video in sequence until one detector asserts it as a copy. To fully exert the complementarity of these detectors, a learning algorithm is proposed to estimate the optimal decision thresholds in the cascade architecture. Excellent performance on the benchmark dataset of TRECVid 2011 CBCD task demonstrates the effectiveness and efficiency of our approach.
Menglin Jiang, Yonghong Tian 0001, Tiejun Huang 0001
ICME3
2012 Macro-Block-Level Selective Background Difference Coding for Surveillance Video
abstract
Utilizing the special properties to improve the surveillance video coding efficiency still has much room, although there have been three typical paradigms of methods: object-oriented, background-prediction-based and background-difference-based methods. However, due to the inaccurate foreground segmentation, the low-quality or unclear background frame, and the potential "foreground pollution" phenomenon, there is still much room for improvement. To address this problem, this paper proposes a macro-block-level selective background difference coding method (MSBDC). MSBDC selects the following two ways to encode each macro-block (MB): coding the original MB, and directly coding the difference data between the MB and its corresponding background. MSBDC also features at employs the classification of MBs to facilitate the selection, through which, prediction and motion compensation turns more accurate, both on foreground and background. Results show that, MSBDC significantly decreases the total bitrate and obtains a remarkable performance gain on foreground compared with several state-of-the-art methods.
Xianguo Zhang, Yonghong Tian 0001, Luhong Liang, Tiejun Huang 0001, Wen Gao 0001
ICME4
2012 Robust and discriminative image authentication based on standard model feature
abstract
The goal of image authentication is to accept content-preserving operations and reject content-altering manipulations. So,it is increasingly approached by extracting content-based invariant features from original images and verifying their preservation in received images at later times. Since sparsity usually implies invariance, sparse feature representation has drawn significant attention from the research community. But only if discrimination is also found with a sparse feature, can it be successfully applied in image authentication. This paper proposes a sparse feature for image authentication by exploring the biologically-motivated standard model. Experimental results demonstrate both robustness and discrimination of the feature, and its effectiveness in tamper detection and location as well.
Luntian Mou, Xilin Chen 0001, Yonghong Tian 0001, Tiejun Huang 0001
ISCAS4
2012 An efficient surveillance coding method based on a timely and bit-saving background updating model
abstract
Background modeling is an important pre-processing step for object detection in surveillance video analysis systems. Recently, it has been proved to be useful for high-efficiency surveillance video coding. In existing works, the modeling background frame often needs to be high-quality encoded so as to achieve a large bit-rate saving. However, the high-quality background frame requires lots of bits in the code stream, so it is infeasible to update the background frame too frequently. Therefore, a better bit-allocation method is desirable to facilitate in-time background updating and bit-saving background coding. In this paper, we firstly build up a background updating model from a detailed analysis of results on surveillance video. Following this, we propose a bit-saving and quality-maintaining background frame coding method. In our method, the background frame can be updated more timely, consequently leading to the better coding efficiency. Experimental results show that our method can achieve more than 15% bit-rate decrease compared with three state-of-art methods.
Xianguo Zhang, Yonghong Tian 0001, Tiejun Huang 0001
VCIP4
2012 Optimizing JPEG quantization table for low bit rate mobile visual search
abstract
Smart phones is bringing about emerging potentials in mobile visual search. Extensive research efforts have been made in compact visual descriptors. However, directly extracting visual descriptors on a mobile device is computationally intensive and time consuming. Towards low bit rate visual search, we propose to deeply compress query images by learning a customized JPEG quantization table in the context of visual search. Distinct from traditional image compression, by incorporating pair-wise image matching precision into distortion measure, we optimize quantization table to seek a better trade-off between image compression rate and visual search performance. An evolutionary algorithm is employed to learn an optimal quantization table. Under MPEG CDVS evaluation framework, extensive evaluation has been done including image retrieval and pair-wise matching over 1 million database images. Experimental results have demonstrated that our optimized quantization table works much better than JPEG default one in terms of retrieval/matching performance vs. a set of different operating points. The proposed low bit rate solution may be easily deployed to smart phones without hardware support, as a useful complement to the ongoing MPEG CDVS standardization efforts.
Ling-Yu Duan, Xiangkai Liu, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001
VCIP4
2012 Low-complexity and high-efficiency background modeling for surveillance video coding
abstract
Recently, background modeling (shortly BgModeling) plays a more and more important role in high-efficiency surveillance video coding. Meanwhile, many practical video coding applications also present some specific requirements for BgModeling, such as the low memory cost and low computational complexity. However, existing BgModeling methods are mostly designed for video content analysis such as object detection. Thus they may be not directly applicable for video coding. In this paper, we firstly present an analysis for the features of BgModeling in surveillance video coding and make a comparison of the performances of existing BgModeling methods. Then we propose a segment-and-weight based running average (SWRA) method for surveillance video coding. SWRA firstly divides pixels at each position in the training frames into several temporal segments, and then calculate their corresponding mean values and weights. After that, a running and weighted average procedure is used to reduce the influence of foreground pixels and finally obtain the modeling results. Experimental results show that, the SWRA-based encoder achieves the best performance over several state-of-the-art methods, with much less cost of memory and modeling time.
Xianguo Zhang, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
VCIP3
2012 Group-Sensitive Multiple Kernel Learning for Object Recognition
abstract
In this paper, a group-sensitive multiple kernel learning (GS-MKL) method is proposed for object recognition to accommodate the intraclass diversity and the interclass correlation. By introducing the "group" between the object category and individual images as an intermediate representation, GS-MKL attempts to learn group-sensitive multikernel combinations together with the associated classifier. For each object category, the image corpus from the same category is partitioned into groups. Images with similar appearance are partitioned into the same group, which corresponds to the subcategory of the object category. Accordingly, intraclass diversity can be represented by the set of groups from the same category but with diverse appearances; interclass correlation can be represented by the correlation between groups from different categories. GS-MKL provides a tractable solution to adapt multikernel combination to local data distribution and to seek a tradeoff between capturing the diversity and keeping the invariance for each object category. Different from the simple hybrid grouping strategy that solves sample grouping and GS-MKL training independently, two sample grouping strategies are proposed to integrate sample grouping and GS-MKL training. The first one is a looping hybrid grouping method, where a global kernel clustering method and GS-MKL interact with each other by sharing group-sensitive multikernel combination. The second one is a dynamic divisive grouping method, where a hierarchical kernel-based grouping process interacts with GS-MKL. Experimental results show that performance of GS-MKL does not significantly vary with different grouping strategies, but the looping hybrid grouping method produces slightly better results. On four challenging data sets, our proposed method has achieved encouraging performance comparable to the state-of-the-art and outperformed several existing MKL methods.
Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Image Process.4
2011 Robust and discriminative image authentication based on sparse coding
abstract
Image authentication is usually approached by checking the preservation of some invariant features, which are expected to be both robust and discriminative so that content-preserving operations are accepted while content-altering manipulations are rejected. However, most of existing features have not obtained convincing performance due to insufficiency of experiments and over biasing of robustness. Motivated by the sparse coding strategy discovered in primary visual cortex, we explore the possibility of using sparse coding coefficients for image authentication. Through extensive experiments, we discover that the proposed feature bears great discrimination as well as robustness, which indicates the effectiveness of sparse coding as a new invariant feature for image authentication.
Luntian Mou, Tiejun Huang 0001, Yonghong Tian 0001, Shiguo Lian, Xilin Chen 0001
CCNC2
2011 Generating vocabulary for global feature representation towards commerce image retrieval
abstract
This paper studies the problem of retrieving images by color, texture and shape in the context of visual assisted product recommendation in E-commerce sites. Different from general CBIR applications, commerce image retrieval puts more emphasis on outlier-free ranking (top N) to gain perfect user experience. We suggest to extend the bag-of-words (BoW) model to global feature characterization rather than commonly used histogram based low-level feature representation. Although BoW is a common practice in object recognition, we argue generating feature vocabulary is useful to address the global feature characterization that could be elegantly adapted to domain specific commerce image search. The representation is compact and discriminative, which may adapt with individual websites. Quantitative as well as subjective evaluation demonstrates the functionality of the proposed method. In practice, the vocabulary based global features greatly reduce outliers in top rank images, so that desirable user experience can be obtained in E-Commerce applications.
Ling-Yu Duan, Chunyu Wang 0001, Tiejun Huang 0001, Wen Gao 0001
ICIP4
2011 Selective eigenbackgrounds method for background subtraction in crowed scenes
abstract
In this paper, a selective eigenbackgrounds method is proposed for background subtraction in crowded scenes. In order to train and update the eigenbackground model with frames containing few objects (i.e. clean frames), virtual frames are constructed based on a frame selection map. Then, the eigenbackground that best depicts background is selected for each pixel based on an eigenbackground selection map. Experimental results show the performance of the proposed method is better than those of some state-of-the-art methods in crowded scenes.
Zhipeng Hu, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001
ICIP4
2011 PKUBench: A context rich mobile visual search benchmark
abstract
While there are ever growing focuses on mobile visual search in recent years, a comprehensive benchmark database with rich context information (such as GPS) for fair evaluation among different strategies is still missing. This paper introduces a PKUBench benchmark for the quantitative evaluations of mobile visual search with the support of GPS. It contains 13,179 images organized into 198 distinct landmark locations within the Peking University campus. Each location is captured with multiple shot sizes and viewing angles, using both digital cameras and phone cameras, each photo being tagged with rich contextual information in the mobile scenario. Moreover, this benchmark studies typical quality degeneration scenarios in mobile photographing, including variable resolutions, blurring, lighting changes, occlusions, as well as various viewing angles. Together with this benchmark, we provide the bag-of-visual-words search baselines involving contextual information refinement. Finally, distractor images are further introduced to evaluate the robustness of visual search methods in this database.
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001, Hongxun Yao, Wen Gao 0001
ICIP5
2011 Learning the trip suggestion from landmark photos on the web
abstract
In this paper, we introduce a novel touristic trip suggestion system to facilitate the traveling of mobile users in a given city. Given the current user location and his touristic destination, our system can suggest a shortest trip path that visits as many popular landmarks as possible. To this end, we collect geographical tagged photos from Flickr [1] and Panoramio [2] photo sharing websites. Then a geographical graph is constructed by modeling photos as vertices and their geographical and visual closenesses as connection strengths. In this graph, we mine a dominant subgraph by quantizing nearby and visually duplicated vertices, and then trimming unpopular subgraphs. Such dominant subgraph only retains the popular landmarks from the consensus of travelers in this city. In online suggestion, we map the current user location and the target location to the nearest vertices in this subgraph, based on which an optimal trip is suggested through a shortest path search. We have quantitatively validated our system in typical areas including Beijing and New York City, with quantitative comparisons to alternative approaches.
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001
ICIP6
2011 A multimodal video copy detection approach with sequential pyramid matching
abstract
Content-based video copy detection over large corpus with complex transformations is important but challenging. It is not surprising that most existing methods fall short of either sufficient robustness to detect severely deformed copies or high accuracy to localize copy segments. In this paper, we propose a video copy detection approach which exploits complementary audio-visual features and sequential pyramid matching (SPM). Several independent detectors first match visual key frames or audio clips using individual features, and then aggregate the frame level results into video level results with SPM, which calculates video similarities by sequence matching at multiple granularities. Finally, detection results from basic detectors are fused and further filtered to generate the final result. Excellent performance evaluated on TRECVid 2010 copy detection task demonstrates the effectiveness of our approach.
Yonghong Tian 0001, Menglin Jiang, Luntian Mou, Xiaoyu Fang, Tiejun Huang 0001
ICIP5
2011 Learning Compact Visual Descriptor for Low Bit Rate Mobile Landmark Search
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001
IJCAI5
2011 Salient region detection and segmentation for general object recognition and image understanding
Tiejun Huang 0001, Yonghong Tian 0001, Jia Li 0003, Haonan Yu
Sci. China Inf. Sci.1
2011 Multi-Task Rank Learning for Visual Saliency Estimation
abstract
Visual saliency plays an important role in various video applications such as video retargeting and intelligent video advertising. However, existing visual saliency estimation approaches often construct a unified model for all scenes, thus leading to poor performance for the scenes with diversified contents. To solve this problem, we propose a multi-task rank learning approach which can be used to infer multiple saliency models that apply to different scene clusters. In our approach, the problem of visual saliency estimation is formulated in a pair-wise rank learning framework, in which the visual features can be effectively integrated to distinguish salient targets from distractors. A multi-task learning algorithm is then presented to infer multiple visual saliency models simultaneously. By an appropriate sharing of information across models, the generalization ability of each model can be greatly improved. Extensive experiments on a public eye-fixation dataset show that our multi-task rank learning approach outperforms 12 state-of-the-art methods remarkably in visual saliency estimation.
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2010 Dynamic multi-cue tracking with detection responses association
abstract
Multi-cue integration has proved successful at increasing the robustness of tracking algorithms and overcoming the failure cases of individual cue. But considering dynamic appearance of objects or clutter background, the integration based on constant weights may weaken the performance of this scheme. In this paper, we propose a dynamic weights update mechanism for multiple cues tracking with detection responses as supervision. We integrate multiple cues based on the observation hypotheses compared with detection association results and adjust the weights according to the approximation degree. The integration is adapted on-the-fly during tracking, in order to keep the tracker adaptive. The proposed method allows flexible combination of different cues and we select cues based on color and local feature for tracking. Experiments are carried out on 602 trajectories extracted from TRECVID 2008 event detection dataset which is recorded in an airport scenario. Comparison results prove the effectiveness of our method.
Guochen Jia, Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
ACM Multimedia4
2010 Saliency detection based on 2D log-gabor wavelets and center bias
abstract
Visual saliency can be a useful tool for image content analysis such as automatic image cropping and image compression. In existing methods on visual saliency detection, most of them are related to the model of receptive field. In this paper, we propose a bottom-up model which introduces 2D Log-Gabor wavelets for saliency detection. Compared with the traditional model of receptive field, the 2D Log-Gabor wavelets can better simulate the biological characteristics of the simple cortical cell in the receptive filed. Moreover, we also incorporate the influence of center bias into our model, which is a common phenomenon that directs visual attention to the center of images in natural scenes. Experimental results show that our approach outperforms three state-of-the-art approaches remarkably.
Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001, Ling-Yu Duan, Guochen Jia
ACM Multimedia3
2010 Automatic interesting object extraction from images using complementary saliency maps
abstract
Automatic interesting object extraction is widely used in many image applications. Among various extraction approaches, saliency-based ones usually have a better performance since they well accord with human visual perception. However, nearly all existing saliency-based approaches suffer the integrity problem, namely, the extracted result is either a small part of the object (referred to as sketch-like) or a large region that contains some redundant part of the background (referred to as envelope-like). In this paper, we propose a novel object extraction approach by integrating two kinds of "complementary" saliency maps (i.e., sketch-like and envelope-like maps). In our approach, the extraction process is decomposed into two sub-processes, one used to extract a high-precision result based on the sketch-like map, and the other used to extract a high-recall result based on the envelope-like map. Then a classification step is used to extract an exact object based on the two results. By transferring the complex extraction task to an easier classification problem, our approach can effectively break down the integrity problem. Experimental results show that the proposed approach outperforms six state-of-art saliency-based methods remarkably in automatic object extraction, and is even comparable to some interactive approaches.
Haonan Yu, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001
ACM Multimedia4
2010 A background model based method for transcoding surveillance videos captured by stationary camera
abstract
Real-world video surveillance applications require storing videos without neglecting any part of scenarios for weeks or months. To reduce the storage cost, the high bit-rate videos from cameras should be transcoded into a more efficient compressed format with as little quality loss as possible. In this paper, we propose a background model based method to improve the transcoding efficiency for surveillance videos captured by stationary cameras, and objectively measure it. The background model is trained by pre-decoded I frames, and then used to transcode the source stream. Following this method, an H.264/AVC based transcoder employing the background model as long-term reference frame and a difference frame coding based transcoder are implemented and evaluated. Experimental results show that both trancoders save nearly half the used bits while maintaining quality compared with the full-decoding-full-encoding method, and the latter one has slightly better performance.
Xianguo Zhang, Luhong Liang, Qian Huang 0008, Tiejun Huang 0001, Wen Gao 0001
PCS4
2010 An efficient coding scheme for surveillance videos captured by stationary cameras
abstract
In this paper, a new scheme is presented to improve the coding efficiency of sequences captured by stationary cameras (or namely, static cameras) for video surveillance applications. We introduce two novel kinds of frames (namely background frame and difference frame) for input frames to represent the foreground/background without object detection, tracking or segmentation. The background frame is built using a background modeling procedure and periodically updated while encoding. The difference frame is calculated using the input frame and the background frame. A sequence structure is proposed to generate high quality background frames and efficiently code difference frames without delay, and then surveillance videos can be easily compressed by encoding the background frames and difference frames in a traditional manner. In practice, the H.264/AVC encoder JM 16.0 is employed as a build-in coding module to encode those frames. Experimental results on eight in-door and out-door surveillance videos show that the proposed scheme achieves 0.12 dB~1.53 dB gain in PSNR over the JM 16.0 anchor specially configured for surveillance videos.
Xianguo Zhang, Luhong Liang, Qian Huang 0008, Yazhou Liu, Tiejun Huang 0001, Wen Gao 0001
VCIP5
2010 Probabilistic Multi-Task Learning for Visual Saliency Estimation in Video
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
Int. J. Comput. Vis.3
2010 A ranking SVM based fusion model for cross-media meta-search engine
abstract
Recently, we designed a new experimental system MSearch, which is a cross-media meta-search system built on the database of the WikipediaMM task of ImageCLEF 2008. For a meta-search engine, the kernel problem is how to merge the results from multiple member search engines and provide a more effective rank list. This paper deals with a novel fusion model employing supervised learning. Our fusion model employs ranking SVM in training the fusion weight for each member search engine. We assume the fusion weight of each member search engine as a feature of a result document returned by the meta-search engine. For a returned result document, we first build a feature vector to represent the document, and set the value of each feature as the document’s score returned by the corresponding member search engine. Then we construct a training set from the documents returned from the meta-search engine to learn the fusion parameter. Finally, we use the linear fusion model based on the overlap set to merge the results set. Experimental results show that our approach significantly improves the performance of the cross-media meta-search (MSearch) and outperforms many of the existing fusion methods.
Ya-li Cao, Tiejun Huang 0001, Yonghong Tian 0001
J. Zhejiang Univ. Sci. C2
2010 Salient object extraction for user-targeted video content association
abstract
The increasing amount of videos on the Internet and digital libraries highlights the necessity and importance of interactive video services such as automatically associating additional materials (e.g., advertising logos and relevant selling information) with the video content so as to enrich the viewing experience. Toward this end, this paper presents a novel approach for user-targeted video content association (VCA). In this approach, the salient objects are extracted automatically from the video stream using complementary saliency maps. According to these salient objects, the VCA system can push the related logo images to the users. Since the salient objects often correspond to important video content, the associated images can be considered as content-related. Our VCA system also allows users to associate images to the preferred video content through simple interactions by the mouse and an infrared pen. Moreover, by learning the preference of each user through collecting feedbacks on the pulled or pushed images, the VCA system can provide user-targeted services. Experimental results show that our approach can effectively and efficiently extract the salient objects. Moreover, subjective evaluations show that our system can provide content-related and user-targeted VCA services in a less intrusive way.
Jia Li 0003, Han-nan Yu, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
J. Zhejiang Univ. Sci. C4
2010 Cost-Sensitive Rank Learning From Positive and Unlabeled Data for Visual Saliency Estimation
abstract
This paper presents a cost-sensitive rank learning approach for visual saliency estimation. This approach avoids the explicit selection of positive and negative samples, which is often used by existing learning-based visual saliency estimation approaches. Instead, both the positive and unlabeled data are directly integrated into a rank learning framework in a cost-sensitive manner. Compared with existing approaches, the rank learning framework can take the influences of both the local visual attributes and the pair-wise contexts into account simultaneously. Experimental results show that our algorithm outperforms several state-of-the-art approaches remarkably in visual saliency estimation.
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
IEEE Signal Process. Lett.3
2010 Sequence Multi-Labeling: A Unified Video Annotation Scheme With Spatial and Temporal Context
abstract
Automatic video annotation is a challenging yet important problem for content-based video indexing and retrieval. In most existing works, annotation is formulated as a multi-labeling problem over individual shots. However, video is by nature informative in spatial and temporal context of semantic concepts. In this paper, we formulate video annotation as a sequence multi-labeling (SML) problem over a shot sequence. Different from many video annotation paradigms working on individual shots, SML aims to predict a multi-label sequence for consecutive shots in a global optimization manner by incorporating spatial and temporal context into a unified learning framework. A novel discriminative method, called sequence multi-label support vector machine (SVMSML), is accordingly proposed to infer the multi-label sequence for a given shot sequence. In SVMSML, a joint kernel is employed to model the feature-level and concept-level context relationships (i.e., the dependencies of concepts on the low-level features, spatial and temporal correlations of concepts). A multiple-kernel learning (MKL) algorithm is developed to optimize the kernel weights of the joint kernel as well as the SML score function. To efficiently search the desirable multi-label sequence over the large output space in both training and test phases, we adopt an approximate method to maximize the energy of a binary Markov random field (BMRF). Extensive experiments on TRECVID'05 and TRECVID'07 datasets have shown that our proposed SVMSMLgains superior performance over the state-of-the-art.
Yuanning Li, Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Multim.5
2009 Robust video fingerprinting based on visual attention regions
abstract
This paper presents a robust video fingerprinting based on visual attention regions. Video fingerprints, which are a set of short feature vectors, are unique to video clips and used for video identification. The performance of video fingerprinting is usually measured in terms of robustness and accuracy of identification. In our proposed approach, we extract video fingerprints using visual attention regions which remain the same for the perceptually same scenes with different types of distortions and different for different scenes. The experimental results show that the proposed video fingerprinting is effective for constructing video fingerprints that are robust against various content-preserving distortions and accurate in identifying different video clips.
Tiejun Huang 0001, Wen Gao 0001
ICASSP2
2009 A dataset and evaluation methodology for visual saliency in video
abstract
Recently, visual saliency has drawn great research interest in the field of computer vision and multimedia. Various approaches aiming at calculating visual saliency have been proposed. To evaluate these approaches, several datasets have been presented for visual saliency in images. However, there are few datasets to capture spatiotemporal visual saliency in video. Intuitively, visual saliency in video is strongly affected by temporal context and might vary significantly even in visually similar frames. In this paper, we present an extensive dataset with 7.5-hour videos to capture spatiotemporal visual saliency. The salient regions in frames sequentially sampled from these videos are manually labeled by 23 subjects and then averaged to generate the ground-truth saliency maps. We also present three metrics to evaluate competing approaches. Several typical algorithms were evaluated on the dataset. The experimental results show that this dataset is very suitable for evaluating visual saliency. We also discover some interesting findings that would be addressed in future research. Currently, the dataset is freely available online together with the source code for evaluation.
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ICME3
2009 A Proposed AVS Decoder Configuration in the Reconfigurable Video Coding Framework
abstract
This demonstration shows an AVS intra decoder configuration in the RVC framework. It explains how to use the dataflow mechanism offered by the RVC framework to support AVS decoder configuration. It also shows the flexibility and convenience to reconfigure decoders in the RVC framework. In this work, the AVS VTL is established containing FUs from AVS. The proposed AVS decoder configuration is implemented by connecting some FUs from AVS VTL and reusing some FUs from MPEG VTL. The demonstration shows that the decoder can decode AVS conformance bitstreams correctly in RVC simulator.
Dandan Ding, Honggang Qi, Lu Yu 0003, Tiejun Huang 0001, Wen Gao 0001
ISCAS4
2009 A SIFT-based Image Fingerprinting Approach Robust to Geometric Transformations
abstract
Among approaches in implementing digital rights management, image fingerprinting technique is considered to be one of the most attractive solutions, especially in detecting illegal use of image works. The deficiency of the existing image fingerprinting methods is that they can not deal with geometric transformations, such as aspect ratio changes, rotations, cropping, combining, etc. Aiming at this shortcoming, we propose a SIFT-based image fingerprinting algorithm which is robust to geometric transformations. Firstly, we introduce SIFT-based algorithm to extract features as a unique fingerprint. Secondly, a method based on area ratio invariance of affine transformation is utilized to verify valid matched keypoint pairs between the queried image and the pre-registered image. Finally, by counting the valid matched pairs, we estimate whether the two images are homologous or not. Experimental results demonstrate that the proposed method exhibits an excellent performance when geometric transformation occurs.
Xinghua Yu, Tiejun Huang 0001
ISCAS2
2009 Reconfigurable video coding framework and decoder reconfiguration instantiation of AVS
Dandan Ding, Honggang Qi, Lu Yu 0003, Tiejun Huang 0001, Wen Gao 0001
Signal Process. Image Commun.4
2009 A secure media streaming mechanism combining encryption, authentication, and transcoding
Luntian Mou, Tiejun Huang 0001, Longshe Huo, Weiping Li 0002, Wen Gao 0001, Xilin Chen 0001
Signal Process. Image Commun.2
2008 Streaming of Governed Content - Time for a Standard
abstract
This paper presents ISO/IEC 23000-5 -media streaming player, an ISO standard targeting the distribution of governed content in streaming mode. This standard is a major integration of MPEG technologies, ranging from MPEG-2 and MPEG-4 to MPEG-21 including a set of inter-device protocols developed by the Digital Media Project, now ISO/IEC 29116-1 -Media Streaming MAF Protocols. The paper also presents Chillout, adopted for the open source reference software of this standard. Chillout is written in Java and provides a set of libraries for developing applications conforming to the standard.
Filippo Chiariglione, Tiejun Huang 0001, Hyon-Gon Choo
CCNC2
2008 Multi-polarity text segmentation using graph theory
abstract
Text segmentation, or named text binarization, is usually an essential step for text information extraction from images and videos. However, most existing text segmentation methods have difficulties in extracting multi-polarity texts, where multi-polarity texts mean those texts with multiple colors or intensities in the same line. In this paper, we propose a novel algorithm for multi-polarity text segmentation based on graph theory. By representing a text image with an undirected weighted graph and partitioning it iteratively, multi-polarity text image can be effectively split into several single-polarity text images. As a result, these text images are then segmented by single-polarity text segmentation algorithms. Experiments on thousands of multi-polarity text images show that our algorithm can effectively segment multi-polarity texts.
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ICIP3
2007 Interoperability Issues in DRM and DMP Solutions
abstract
This paper presents an ongoing digital right management (DRM) standard: Digital Media Project (DMP), which is active as an interoperable DRM solution in recent three years. We start by reviewing the important works of DRM standards at present, mainly focusing on their interoperability issues. Then we describe the core techniques in DMP and explain how they fit into an interoperable DRM system. We also give a report on the development of the reference software of DMP -Chillout, which is an open source project. Finally, we outline several tasks that we will undertake in the next step of DMP to optimize this technical standard.
Tiejun Huang 0001
ICME2
2007 A Perception-based Scalable Encryption Model for AVS Audio
abstract
Audio Video coding Standard (AVS) is China's second-generation source coding/decoding standard with fully Intellectual Properties. As the sixth part of AVS standard, AVS-DRM aims to offer the universal and open interoperable standard for various requirements of digital rights management (DRM) in digital media industry. In its first version, AVS DRM committee document (CD) did not take into account the encryption modes on AVS audio. However, audio content protection also plays an important role in many digital media applications. Therefore, this paper proposes a perception-based scalable encryption approach for AVS audio, which specifies different audio encryption modes by utilizing the perception classification of the audio bitstream and provides multiple security levels by encrypting different audio coding layers. Moreover, the paper incorporates the proposed scalable encryption method in AVS audio fine-granularity scalable codec technique and then implements an AVS audio trusted coder/decoder. Several objective and subjective tests were performed on a set of 12 critical stereo excerpts provided by AVS audio group. The experimental results show that the proposed scalable encryption approach is able to provide a feasible and effective audio protection scheme for a wide range of applications with different DRM requirements. Currently, the scalable encryption scheme has been partly accepted by AVS-DRM final committee document (FCD).
Juan Lan, Tiejun Huang 0001, Junhua Qu
ICME2
2007 A DRM Architecture for Manageable P2P Based IPTV System
Xiaoyun Liu, Tiejun Huang 0001, Longshe Huo, Luntian Mou
ICME2
2007 Towards multi-granularity multi-facet e-book retrieval
abstract
Generally speaking, digital libraries have multiple granularities of semantic units: book, chapter, page, paragraph and word. However, there are two limitations of current eBook retrieval systems: (1) the granularity of retrievable units is either too big or too small, scales such as chapters, paragraphs are ignored; (2) the retrieval results should be grouped by facets to facilitate user's browsing and exploration. To overcome these limitations, we propose a multi-granularity multi-facet eBook retrieval approach.
Chong Huang 0006, Yonghong Tian 0001, Tiejun Huang 0001
WWW4
2006 Semantic Scoring Based on Small-World Phenomenon for Feature Selection in Text Mining
Chong Huang 0006, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ADMA3
2006 Robust Collective Classification with Contextual Dependency Network Models
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ADMA2
2006 Keyphrase Extraction Using Semantic Networks Structure Analysis
abstract
Keyphrases play a key role in text indexing, summarization and categorization. However, most of the existing keyphrase extraction approaches require human-labeled training sets. In this paper, we propose an automatic keyphrase extraction algorithm, which can be used in both supervised and unsupervised tasks. This algorithm treats each document as a semantic network. Structural dynamics of the network are used to extract keyphrases (key nodes) unsupervised. Experiments demonstrate the proposed algorithm averagely improves 50% in effectiveness and 30% in efficiency in unsupervised tasks and performs comparatively with supervised extractors. Moreover, by applying this algorithm to supervised tasks, we develop a classifier with an overall accuracy up to 80%.
Chong Huang 0006, Yonghong Tian 0001, Charles Ling 0001, Tiejun Huang 0001
ICDM5
2006 Diversifying the image retrieval results
abstract
In the area of image retrieval, post-retrieval processing is often used to refine the retrieval results to better satisfy users' requirements. Previous methods mainly focus on presenting users with relevant results. However, in most cases, users cannot clearly present their requirements by several query words. Therefore, relevant results with rich topic coverage are more likely to meet users' ambiguous needs. In this paper, a re-ranking method based on topic richness analysis is proposed to enrich topic coverage in retrieval results. Furthermore, a quantitative criterion called diversity scores (DS) is proposed to evaluate the improvement. Given a set of images, topics that are rarely included in the set are scarce topics, as oppose to rich topics that are widely distributed among the set. Scarce topics contribute more than rich topics do to the DS of images. Five researchers are invited to evaluate the re-ranked results both in topic coverage and relevance. Experimental results on over 20,000 images demonstrate that our proposed approach is effective in improving the topic coverage of retrieval results without loss of relevance.
Yonghong Tian 0001, Wen Gao 0001, Tiejun Huang 0001
ACM Multimedia4
2006 Basic Considerations on AVS DRM Architecture
Tiejun Huang 0001, Yongliang Liu
J. Comput. Sci. Technol.1
2006 Latent linkage semantic kernels for collective classification of link data
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
J. Intell. Inf. Syst.2
2006 Learning Contextual Dependency Network Models for Link-Based Classification
abstract
Links among objects contain rich semantics that can be very helpful in classifying the objects. However, many irrelevant links can be found in real-world link data such as Web pages. Often, these noisy and irrelevant links do not provide useful and predictive information for categorization. It is thus important to automatically identify which links are most relevant for categorization. In this paper, we present a contextual dependency network (CDN) model for classifying linked objects in the presence of noisy and irrelevant links. The CDN model makes use of a dependency function that characterizes the contextual dependencies among linked objects. In this way, CDNs can differentiate the impacts of the related objects on the classification and consequently reduce the effect of irrelevant links on the classification. We show how to learn the CDN model effectively and how to use the Gibbs inference framework over the learned model for collective classification of multiple linked objects. The experiments show that the CDN model demonstrates relatively high robustness on data sets containing irrelevant links.
Yonghong Tian 0001, Qiang Yang 0001, Tiejun Huang 0001, Charles Ling 0001, Wen Gao 0001
IEEE Trans. Knowl. Data Eng.3
2005 Visual Ontology Construction for Digitized Art Image Retrieval
Shuqiang Jiang, Qingming Huang, Tiejun Huang 0001, Wen Gao 0001
J. Comput. Sci. Technol.4
2004 A new method to segment playfield and its applications in match analysis in sports video
abstract
With the growing popularity of digitized sports video, automatic analysis of them need be processed to facilitate semantic summarization and retrieval. Playfield plays the fundamental role in automatically analyzing many sports programs. Many semantic clues could be inferred from the results of playfield segmentation. In this paper, a novel playfield segmentation method based on Gaussian mixture models (GMMs) is proposed. Firstly, training pixels are automatically sampled from frames. Then, by supposing that field pixels are the dominant components in most of the video frames, we build the GMMs of the field pixels and use these models to detect playfield pixels. Finally region-growing operation is employed to segment the playfield regions from the background. Experimental results show that the proposed method is robust to various sports videos even for very poor grass field conditions. Based on the results of playfield segmentation, match situation analysis is investigated, which is also desired for sports professionals and longtime fanners. The results are encouraging.
Shuqiang Jiang, Qixiang Ye, Wen Gao 0001, Tiejun Huang 0001
ACM Multimedia4
2004 An Ontology-based Approach to Retrieve Digitized Art Images
abstract
Although much progress has been made, current low-level based visual information retrieval technology does not allow users to formulate queries through high-level semantics. More and more digitized art images appear on the Internet, and techniques need to be established on how to organize and retrieve them. In this work, a framework for retrieving art images using an ontology-based method is introduced. The proposed ontology describes images in various aspects. Non-objectionable semantics are first introduced, and how to express these semantics is given. Concepts in the ontology could be automatically derived. The retrieval scheme makes users more naturally find visual information and experimental implementation demonstrates good potential on retrieving art images in a human-centered manner.
Shuqiang Jiang, Tiejun Huang 0001, Wen Gao 0001
Web Intelligence2
2004 Two-phase Web site classification based on Hidden Markov Tree models
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
Web Intell. Agent Syst.2
2003 Two-Phase Web Site Classification Based on Hidden Markov Tree Models
abstract
With the exponential growth of both the amount and diversity of the information that the Web encompasses, automatic classification of topic-specific Web sites is highly desirable. We propose a novel approach for Web site classification based on the content, structure and context information of Web sites. In our approach, the site structure is represented as a two-layered tree in which each page is modeled as a DOM (document object model) tree and a site tree is used to hierarchically link all pages within the site. Two context models are presented to capture the topic dependences in the site. Then the hidden Markov tree (HMT) model is utilized as the statistical model of the site tree and the DOM tree, and an HMT-based classifier is presented for their classification. Moreover, for reducing the download size of Web sites but still keeping high classification accuracy, an entropy-based approach is introduced to dynamically prune the site trees. On these bases, we employ the two-phase classification system for classifying Web sites through a fine-to-coarse recursion. The experiments show our approach is able to offer high accuracy and efficient process performance.
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001, PingBo Kang
Web Intelligence2