EDBT 2026 Demo / reviewers in the wild / expert
Hua Huang 0001
dblp:70/5618-1
· DBLP profile ↗
185ranked-venue papers
14as first author
98since 2021 · last 2026
0000-0003-2587-1702ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 9 first-author · 51 since 2021Artificial intelligence and machine learning · 75 · 3 first-author · 49 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMART: Evaluating LLMs' Mathematical Reasoning via a Human Cognitive Process-Inspired BenchmarkabstractLarge Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks.However, concerns remain as to whether these successes reflect genuine reasoning or superficial pattern recognition.Existing evaluation methods, which typically focus either on the final answer or on the intermediate reasoning steps, reduce mathematical reasoning to a shallow input-output mapping, overlooking its inherently multi-stage and multi-dimensional cognitive nature.Inspired by Pólya's problem-solving theory, we propose SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: Semantic Understanding, Mathematical Reasoning, Arithmetic Computation, and Reflection & Refinement, and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs.We apply SMART to 22 state-of-the-art open-and closed-source LLMs and uncover substantial discrepancies in their capabilities across dimensions.Our findings reveal genuine weaknesses in current models and motivate a new metric, the All-Pass Score, designed to better capture true problem-solving capability. Yujie Hou, Yaoyao Zhong, Ting Zhang 0002, Xuetao Ma 0001, Hua Huang 0001 |
ACL (1) | 6 |
| 2026 | MagicGeo: Training-free text-guided geometric diagram generationabstractWhile text-to-image generation has made strides in photorealistic imagery, creating accurate geometric diagrams remains a challenge due to the need for precise spatial relationships and the scarcity of geometry-specific datasets. This paper presents MagicGeo, a training-free framework for generating geometric diagrams from textual descriptions. MagicGeo formulates the diagram generation process as a coordinate optimization problem, ensuring geometric correctness through a formal language solver, and then employs coordinate-aware generation. The framework leverages the strong language translation capability of large language models, while formal mathematical solving ensures geometric correctness. We further introduce MagicGeoBench, a benchmark dataset of 220 geometric diagram descriptions, and demonstrate that MagicGeo outperforms current methods in both qualitative and quantitative evaluations. This work provides a scalable, accurate solution for automated diagram generation, with significant implications for educational and academic applications. Ting Zhang 0002, Qunyi Xie, Jingdong Wang 0001, Hua Huang 0001 |
Graph. Model. | 6 |
| 2026 | Optical-to-SAR domain adaptation with inversion regularization for unsupervised ship detection
Yuanfei Huang, Ping Wang 0009, Hua Huang 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Learning Physics-Informed Noise Models from Dark Frames for Low-Light Raw Image DenoisingabstractRecently, the mainstream practice for training low-light raw image denoising methods has shifted towards employing synthetic data. Noise modeling, which focuses on characterizing the noise distribution of real-world sensors, profoundly influences the effectiveness and practicality of synthetic data. Currently, physics-based noise modeling struggles to characterize the entire real noise distribution, while learning-based noise modeling impractically depends on paired real data. In this paper, we propose a novel strategy: learning the noise model from dark frames instead of paired real data, to break down the data dependency. Based on this strategy, we introduce an efficient physics-informed noise neural proxy (PNNP) to approximate the real-world sensor noise model. Specifically, we integrate physical priors into neural proxies and introduce three efficient techniques: physics-guided noise decoupling (PND), physics-aware proxy model (PPM), and differentiable distribution loss (DDL). PND decouples the dark frame into different components and handles different levels of noise flexibly, which reduces the complexity of noise modeling. PPM incorporates physical priors to constrain the synthetic noise, which promotes the accuracy of noise modeling. DDL provides explicit and reliable supervision for noise distribution, which promotes the precision of noise modeling. PNNP exhibits powerful potential in characterizing the real noise distribution. Extensive experiments on public datasets demonstrate superior performance in practical low-light raw image denoising. The source code will be publicly available at the https://fenghansen.github.io/publication/PNNP. Hansen Feng, Lizhi Wang 0001, Yiqi Huang, Yuzhi Wang, Lin Zhu 0012, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Tackling Ill-Posedness of Reversible Image Conversion With Well-Posed Invertible NetworkabstractReversible image conversion (RIC) suffers from ill-posedness issues due to its forward conversion process being considered an underdetermined system. Despite employing invertible neural networks (INN), existing RIC methods intrinsically remain ill-posed as inevitably introducing uncertainty by incorporating randomly sampled variables. To tackle the ill-posedness dilemma, we focus on developing a reliable approximate left inverse for the underdetermined system by constructing an overdetermined system with a non-zero Gram determinant, thus ensuring a well-posed solution. Based on this principle, we propose a well-posed invertible $1\times 1$1×1 convolution (WIC), which eliminates the reliance on random variable sampling and enables the development of well-posed invertible networks. Furthermore, we design two innovative networks, WIN-Naïve and WIN, with the latter incorporating advanced skip-connections to enhance long-term memory. Our methods are evaluated across diverse RIC tasks, including reversible image hiding, image rescaling, and image decolorization, consistently achieving state-of-the-art performance. Extensive experiments validate the effectiveness of our approach, demonstrating its ability to overcome the bottlenecks of existing RIC solutions and setting a new benchmark in the field. Yuanfei Huang, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | DSNeRF: Dynamic View Synthesis for Ultra-Fast Scenes From Continuous Spike StreamsabstractSpike cameras generate binary spikes in response to light intensity changes, enabling high-speed visual perception with unprecedented temporal resolution. However, the unique characteristics of spike stream present significant challenges for reconstructing dense 3D scene representations, particularly in dynamic environments and under non-ideal lighting conditions. In this paper, we introduce DSNeRF, the first method to derive a NeRF-based volumetric scene representation from spike camera data. Our approach leverages NeRF's multi-view consistency to establish robust self-supervision, effectively eliminating erroneous measurements and uncovering coherent structures within exceedingly noisy input amidst diverse real-world illumination scenarios. We propose a novel mapping from pixel rays to the spike domain, integrating the spike generation process directly into NeRF training. Specifically, DSNeRF introduces an integrate-and-fire neuron layer that models non-idealities to capture intrinsic camera noise, including both random and fixed-pattern spike noise, thereby enhancing scene fidelity. Additionally, we propose a motion-guided spiking neuron layer and a long-term rendering photometric loss to better align dynamic spike streams, ensuring accurate scene geometry. Our method optimizes neural radiance fields to render photorealistic novel views from continuous spike streams, demonstrating advantages over other vision sensors in certain scenes. Empirical evaluations on both real and simulated sequences validate the effectiveness of our approach. Lin Zhu 0012, Kangmin Jia, Yifan Zhao 0002, Yunshan Qi, Lizhi Wang 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Scalable portrait matte creation with layer diffusion and connectivity priors
Hao Lu 0003, Hua Huang 0001 |
Pattern Recognit. | 3 |
| 2026 | Self-Supervised Depth Completion Guided by 3D Perception and Geometry ConsistencyabstractDepth completion which aims at predicting dense depth maps from sparse depth measurements, plays a crucial role in many computer graphics and computer vision applications. Previous supervised learning based approaches have demonstrated overwhelming success in this task, while unsupervised high-precision depth completion without relying on the ground-truth data still remains challenging. One main drawback of most previous unsupervised solutions comes from the ignorance of 3D structural information, which often leads to inaccurate spatial propagation and mixed-depth problems. To alleviate the above challenges, this paper explores the utilization of 3D perceptual features and multi-view geometry consistency to devise a high-precision self-supervised depth completion method. Our key contribution is a 3D perceptual spatial propagation constructed with a point cloud representation and an attention weighting mechanism, to capture more reasonable and favorable neighbors during the depth propagation process. Based on the 3D perceptual spatial propagation, we also introduce multi-view geometric constraints between adjacent views to guide the optimization of the whole depth completion model, which achieves geometry consistent depth completion in a self-supervised manner. Extensive experiments on benchmark datasets of NYU-Depth-v2, VOID and KITTI Depth Completion demonstrate that the proposed model achieves the state-of-the-art depth completion performance compared with other unsupervised methods, and even competitive performance compared with previous supervised methods. Tianyu Shen, Shi-Sheng Huang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | GANG: Geometrically-Aligned Neural Gaussians for Efficient and Realistic RelightingabstractEfficient and realistic relighting of complex scenes with unknown illumination remains a crucial but challenging task. Recent advancements in 3D Gaussian Splatting (3DGS) have shown impressive object-level relighting. However, they still struggle with complex real-world scenes, mainly due to the challenges of accurately decoupling intricate geometry, materials, and lighting using concise 3D Gaussian primitives. In this paper, we propose a new Geometrically-Aligned Neural Gaussian Splatting (GANG) method, which performs efficient physically based rendering (PBR) directly on anchor-based relightable neural Gaussians. Our key idea is to regularize the decoded neural Gaussians geometrically aligned with the latent signed distance field (SDF) surface spawned from anchors using a differentiable implicit indicator function (IIF) solver. It brings effective geometric association to accurate decoupling of materials and lighting for efficient and realistic relighting of complex scenes. Furthermore, we propose a locally consistent geometry regularization to guide more concise neural Gaussian learning with a hybrid lighting model, which combines position-learnable spherical Gaussians (SGs) and an environment map, allowing accurate modeling of both local and global illumination. Experimental results on public datasets demonstrate that GANG consistently outperforms previous PBR methods in material decomposition and relighting quality, while representing complex scenes with concise anchors. To the best of our knowledge, GANG is a new state-of-the-art 3DGS method for realistic relighting, enabling efficient rendering and flexible editing materials and illumination, especially for complex scenes. Deqi Li, Shi-Sheng Huang, Hongbo Fu 0001, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Positive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image DenoisingabstractImage denoising enhances image quality, serving as a foundational technique across various computational photography applications. The obstacle to clean image acquisition in real scenarios necessitates the development of self-supervised image denoising methods only depending on noisy images, especially a single noisy image. Existing self-supervised image denoising paradigms (Noise2Noise and Noise2Void) rely heavily on information-lossy operations, such as downsampling and masking, culminating in low-quality denoising performance. In this paper, we propose a novel self-supervised single image denoising paradigm, Positive2Negative, to break the information-lossy barrier. Our paradigm involves two key steps: Renoised Data Construction (RDC) and Denoised Consistency Supervision (DCS). RDC renoises the predicted denoised image by the predicted noise to construct multiple noisy images, preserving all the information of the original image. DCS ensures consistency across the multiple denoised images, supervising the network to learn robust denoising. Our Positive2Negative paradigm achieves state-of-the-art performance in self-supervised single image denoising with significant speed improvements. The code is released to the public at https://github.com/Li-Tong-621/P2N. Tong Li 0016, Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 6 |
| 2025 | Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image DenoisingabstractExisting single-image denoising algorithms often struggle to restore details when dealing with complex noisy images. The introduction of near-infrared (NIR) images offers new possibilities for RGB image denoising. However, due to the inconsistency between NIR and RGB images, the existing works still struggle to balance the contributions of two fields in the process of image fusion. In response to this, in this paper, we develop a cross-field Frequency Correlation Exploiting Network (FCENet) for NIR-assisted image denoising. We first propose the frequency correlation prior based on an in-depth statistical frequency analysis of NIR-RGB image pairs. The prior reveals the complementary correlation of NIR and RGB images in the frequency domain. Leveraging frequency correlation prior, we then establish a frequency learning framework composed of Frequency Dynamic Selection Mechanism (FDSM) and Frequency Exhaustive Fusion Mechanism (FEFM). FDSM dynamically selects complementary information from NIR and RGB images in the frequency domain, and FEFM strengthens the control of common and differential features during the fusion process of NIR and RGB features. Extensive experiments on simulated and real data validate that the proposed method outperforms other state-of-the-art methods. The code will be released at https://github.com/yuchenwang815/FCENet. Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 7 |
| 2025 | Process-Supervised Reinforcement Learning for Code GenerationabstractExisting reinforcement learning (RL) strategies based on outcome supervision have proven effective in enhancing the performance of large language models (LLMs) for code generation.While reinforcement learning based on process supervision shows great potential in multi-step reasoning tasks, its effectiveness in the field of code generation still lacks sufficient exploration and verification.The primary obstacle stems from the resource-intensive nature of constructing a high-quality process-supervised reward dataset, which requires substantial human expertise and computational resources.To overcome this challenge, this paper proposes a "mutation/refactoring-execution verification" strategy.Specifically, the teacher model is used to mutate and refactor the statement lines or blocks, and the execution results of the compiler are used to automatically label them, thus generating a process-supervised reward dataset.Based on this dataset, we have carried out a series of RL experiments.The experimental results show that, compared with the method relying only on outcome supervision, reinforcement learning based on process supervision performs better in handling complex code generation tasks.In addition, this paper for the first time confirms the advantages of the Direct Preference Optimization (DPO) method in the RL task of code generation based on process supervision, providing new ideas and directions for code generation research. Yufan Ye, Ting Zhang 0002, Wenbin Jiang 0002, Hua Huang 0001 |
EMNLP | 4 |
| 2025 | GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray DiffusionabstractAccurate surface reconstruction from unposed images is crucial for efficient 3D object or scene creation. However, it remains challenging, particularly for the joint camera pose estimation. Previous approaches have achieved impressive pose-free surface reconstruction results in dense-view settings, but could easily fail for sparse-view scenarios without sufficient visual overlap. In this paper, we propose a new technique for pose-free surface reconstruction, which follows triplane-based signed distance field (SDF) learning but regularizes the learning by explicit points sampled from ray-based diffusion of camera pose estimation. Our key contribution is a novel Geometric Consistent Ray Diffusion model (GCRayDiffusion), where we represent camera poses as neural bundle rays and regress the distribution of noisy rays via a diffusion model. More importantly, we further condition the denoising process of RGRayDiffusion using the triplane-based SDF of the entire scene, which provides effective 3D consistent regularization to achieve multi-view consistent camera pose estimation. Finally, we incorporate RGRayDiffusion into the triplane-based SDF learning by introducing on-surface geometric regularization from the sampling points of the neural bundle rays, which leads to highly accurate pose-free surface reconstruction results even for sparse-view inputs. Extensive evaluations on public datasets show that our GCRayDiffusion achieves more accurate camera pose estimation than previous approaches, with geometrically more consistent surface reconstruction results, especially given sparse-view inputs. Li-Heng Chen, Zixin Zou, Tianjiao Jing, Yan-Pei Cao 0001, Shi-Sheng Huang, Hongbo Fu 0001, Hua Huang 0001 |
ICCV | 8 |
| 2025 | Noise-Modeled Diffusion Models for Low-Light Spike Image Restoration
Lin Zhu 0012, Xijie Xiang, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 5 |
| 2025 | Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
Ziming Yu, Pan Zhou 0002, Sike Wang, Jia Li 0002, Mi Tian 0008, Hua Huang 0001 |
ICCV | 6 |
| 2025 | EMatch: A Unified Framework for Event-Based Optical Flow and Stereo MatchingabstractEvent cameras have shown promise in vision applications like optical flow estimation and stereo matching, with many specialized architectures leveraging the asynchronous and sparse nature of event data. However, existing works only focus event data within the confines of task-specific domains, overlooking how tasks across the temporal and spatial domains can reinforce each other. In this paper, we reformulate event-based flow estimation and stereo matching as a unified dense correspondence matching problem, enabling us to solve both tasks within a single model by directly matching features in a shared representation space. Specifically, our method utilizes a Temporal Recurrent Network to aggregate event features across temporal or spatial domains, and a Spatial Contextual Attention to enhance knowledge transfer across event flows via temporal or spatial interactions. By utilizing a shared feature similarities module that integrates knowledge from event streams via temporal or spatial interactions, our network performs optical flow estimation from temporal event segment inputs and stereo matching from spatial event segment inputs simultaneously. We demonstrate that our unified model inherently supports multi-task fusion and cross-task transfer. Without the need for retraining for specific task, our model can effectively handle both optical flow and stereo estimation, achieving state-of-the-art performance on both tasks. Pengjie Zhang, Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 5 |
| 2025 | Pi-GPS: Enhancing Geometry Problem Solving by Unleashing the Power of Diagrammatic Information
Ting Zhang 0002, Mi Tian 0008, Hua Huang 0001 |
ICCV | 5 |
| 2025 | Spatio-Temporally Consistent Depth Estimation for Dynamic Scenes using 3D Scene FlowsabstractDynamic depth estimation continues to be crucial but challenging mainly due to the violation of multi-view consistency raised by dynamic areas. Recent approaches have made impressive progress by implicitly fusing the intra-relation features, but is still limited for heterogeneous dynamic scenes. In this paper, we propose a new intra-relation feature fusion, which can significantly improve the fusion quality using an explicit regularization from 3D scene flow cues. We first introduces a Dual Cross-Cue Fusion (D-CCF) module for depth prediction, and further build up an efficient 3D scene flow estimation as explicit 3D spatio-temporal corresponding priors to regularize the depth prediction. Finally, by jointly learning both the depth prediction and 3D scene flow estimation in a unsupervised manner, we achieve more accurate dynamic depth estimation towards spatio-temporal consistency. By extensive evaluation on challenging benchmarks (KITTI and DDAD), our approach can achieve better depth estimation results than state-of-the-art approaches in both static and dynamic areas, which especially maintains the spatio-temporal consistency for dynamic scenes. Tianjiao Jing, Zhengxuan Lian, Shi-Sheng Huang, Hua Huang 0001 |
ICME | 6 |
| 2025 | EvFocus: Learning to Reconstruct Sharp Images from Out-of-Focus Event StreamsabstractEvent cameras are innovative sensors that capture brightness changes as asynchronous events rather than traditional intensity frames. These cameras offer substantial advantages over conventional cameras, including high temporal resolution, high dynamic range, and the elimination of motion blur. However, defocus blur, a common image quality degradation resulting from out-of-focus lenses, complicates the challenge of event-based imaging. Due to the unique imaging mechanism of event cameras, existing focusing algorithms struggle to operate efficiently on sparse event data. In this work, we propose EvFocus, a novel architecture designed to reconstruct sharp images from defocus event streams for the first time. Our work includes the development of an event-based out-of-focus camera model and a simulator to generate realistic defocus event streams for robust training and testing. EvDefous integrates a temporal information encoder, a blur-aware two-branch decoder, and a reconstruction and re-defocus module to effectively learn and correct defocus blur. Extensive experiments on both simulated and real-world datasets demonstrate that EvFocus outperforms existing methods across varying lighting conditions and blur sizes, proving its robustness and practical applicability in event-based defocus imaging. Lin Zhu 0012, Xiantao Ma, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICML | 5 |
| 2025 | Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events
Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 5 |
| 2025 | DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare RemovalabstractLens flare removal remains an information confusion challenge in the underlying image background and the optical flares, due to the complex optical interactions between light sources and camera lens. While recent solutions have shown promise in decoupling the flare corruption from image, they often fail to maintain contextual consistency, leading to incomplete and inconsistent flare removal. To eliminate this limitation, we propose DeflareMamba, which leverages the efficient sequence modeling capabilities of state space models while maintains the ability to capture local-global dependencies. Particularly, we design a hierarchical framework that establishes long-range pixel correlations through varied stride sampling patterns, and utilize local-enhanced state space models that simultaneously preserves local details. To the best of our knowledge, this is the first work that introduces state space models to the flare removal task. Extensive experiments demonstrate that our method effectively removes various types of flare artifacts, including scattering and reflective flares, while maintaining the natural appearance of non-flare regions. Further downstream applications demonstrate the capacity of our method to improve visual object recognition and cross-modal semantic understanding. Code is available at https://github.com/BNU-ERC-ITEA/DeflareMamba. Yuanfei Huang, Junhui Lin, Hua Huang 0001 |
ACM Multimedia | 4 |
| 2025 | TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting PriorsabstractReconstructing transparent surfaces is essential for tasks such as robotic manipulation in labs, yet it poses a significant challenge for 3D reconstruction techniques like 3D Gaussian Splatting (3DGS). These methods often encounter a transparency-depth dilemma, where the pursuit of photorealistic rendering through standard α-blending undermines geometric precision, resulting in considerable depth estimation errors for transparent materials. To address this issue, we introduce Transparent Surface Gaussian Splatting (TSGS), a new framework that separates geometry learning from appearance refinement. In the geometry learning stage, TSGS focuses on geometry by using specular-suppressed inputs to accurately represent surfaces. In the second stage, TSGS improves visual fidelity through anisotropic specular modeling, crucially maintaining the established opacity to ensure geometric accuracy. To enhance depth inference, TSGS employs a first-surface depth extraction method. This technique uses a sliding window over α-blending weights to pinpoint the most likely surface location and calculates a robust weighted average depth. To evaluate the transparent surface reconstruction task under realistic conditions, we collect a TransLab dataset that includes complex transparent laboratory glassware. Extensive experiments on TransLab show that TSGS achieves accurate geometric reconstruction and realistic rendering of transparent objects simultaneously within the efficient 3DGS framework. Specifically, TSGS significantly surpasses current leading methods, achieving a 37.3% reduction in chamfer distance and an 8.0% improvement in F1 score compared to the top baseline. Additionally, TSGS maintains high-quality novel view synthesis, evidenced by a 0.41dB gain in PSNR, demonstrating that TSGS overcomes the transparency-depth dilemma. The code and dataset are available at https://longxiang-ai.github.io/TSGS/. Pu Pang, Hehe Fan, Hua Huang 0001, Yi Yang 0001 |
ACM Multimedia | 4 |
| 2025 | APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechabstractAutomatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in predicting mean opinion scores (MOS) to assess synthetic speech, the neglect of fundamental auditory perception mechanisms limits consistency with human judgments. To address this issue, we propose an auditory perception guided-MOS prediction model (APG-MOS) that synergistically integrates auditory modeling with semantic analysis to enhance consistency with human judgments. Specifically, we first design a perceptual module, grounded in biological auditory mechanisms, to simulate cochlear functions, which encodes acoustic signals into biologically aligned electrochemical representations. Secondly, we propose a residual vector quantization (RVQ)-based semantic distortion modeling method to quantify the degradation of speech quality at the semantic level. Finally, we design a residual cross-attention module, coupled with a progressive learning strategy, to enable multimodal fusion of encoded electrochemical signals and semantic representations. Experiments demonstrate that APG-MOS achieves superior performance on two primary benchmarks. The implementation code is available at https://github.com/BNU-ERC-ITEA/APG-MOS. Zhicheng Lian, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 3 |
| 2025 | Mixture of Group Experts for Learning Invariant RepresentationsabstractSparsely activated Mixture-of-Experts (MoE) models effectively increase the number of parameters while maintaining consistent computational costs per token. However, vanilla MoE models often suffer from limited diversity and specialization among experts, constraining their performance and scalability, especially as the number of experts increases. In this paper, we present a novel perspective on vanilla MoE with top-k routing inspired by sparse representation. This allows us to bridge established theoretical insights from sparse representation into MoE models. Building on this foundation, we propose a group sparse regularization approach for the input of top-k routing, termed Mixture of Group Experts (MoGE). MoGE indirectly regularizes experts by imposing structural constraints on the routing inputs, while preserving the original MoE architecture. Furthermore, we organize the routing input into a 2D topographic map, spatially grouping neighboring elements. This structure enables MoGE to capture representations invariant to minor transformations, thereby significantly enhancing expert diversity and specialization. Comprehensive evaluations across various Transformer models for image classification and language modeling tasks demonstrate that MoGE substantially outperforms its MoE counterpart, with minimal additional memory and computation overhead. Our approach provides a simple yet effective solution to scale the number of experts and reduce redundancy among them. Our code is available at: https://github.com/wangyuankl123/MoGE. Lei Kang 0007, Jia Li 0002, Mi Tian 0008, Hua Huang 0001 |
MMAsia | 4 |
| 2025 | Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionabstractEvent cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms. Code: https://github.com/BIT-Vision/SATE. Lin Zhu 0012, Tengyu Long, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
NeurIPS | 5 |
| 2025 | Mixture of Group Experts for Multi-task Dense Prediction
Lei Kang 0007, Jia Li 0002, Hua Huang 0001 |
PRCV (3) | 3 |
| 2025 | Beyond Image Prior: Embedding Noise Prior into Latent Space of Conditional Denoising Transformer
Yuanfei Huang, Hua Huang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Learning Refractive-Diffractive Optics with Unidirectional Transformer for Large Field-of-View Imaging
Xiangtian Ma, Lizhi Wang 0001, Qilin Sun 0001, Lin Zhu 0012, Hua Huang 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | A Survey of Recent Advances in Generative 3D Reconstruction
Shi-Sheng Huang, Shao-Kui Zhang, Sheng Yang 0007, Jian-Wei Guo, Hua Huang 0001 |
J. Comput. Sci. Technol. | 5 |
| 2025 | Building Non-Uniform Degradation Model for Position-Aware Hyperspectral Image FusionabstractThe fusion of low-spatial-resolution hyperspectral image (LR-HSI) with high-spatial-resolution multispectral image (HR-MSI) has become an effective way to obtain the high-spatial-resolution hyperspectral image (HR-HSI). Currently, learning-based methods have emerged as the mainstream solution in this field. However, these methods typically rely on predefined or simplified degradation models during fusion training, resulting in inaccurate supervision of the fusion networks. Meanwhile, most methods overlook the degradation characteristics in designing the fusion networks, leading to a mismatch between the degradation and fusion processes. These limitations ultimately result in unsatisfactory fusion performance on real data. To enhance the practicality of learning-based methods, accurate degradation modeling and effective network design have become the critical priorities. We observe that, in practical scenarios, the degree of pixel degradation varies across different positions due to the unforeseen factors such as illumination variations and imaging system fluctuations. Considering this, we propose a non-uniform degradation model (NUD), which introduces non-uniformity into the degradation processes of LR-HSI and HR-MSI. In addition, we emphasize that the essence of fusion is to reverse the degradation process. Therefore, to align with the non-uniform degradation process, the fusion process should exhibit similar positional specificity. For this purpose, we propose a position-aware fusion network (PAF), which employs positional encoding to endow the fusion process with the position-aware attribute. Experimental results show that our proposed methods provide an effective solution for HSI fusion in practical scenarios. Lizhi Wang 0001, Lin Zhu 0012, Renwei Dian, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Continuous-Time Object Segmentation Using High Temporal Resolution Event CameraabstractEvent cameras are novel bio-inspired sensors, where individual pixels operate independently and asynchronously, generating intensity changes as events. Leveraging the microsecond resolution (no motion blur) and high dynamic range (compatible with extreme light conditions) of events, there is considerable promise in directly segmenting objects from sparse and asynchronous event streams in various applications. However, different from the rich cues in video object segmentation, it is challenging to segment complete objects from the sparse event stream. In this paper, we present the first framework for continuous-time object segmentation from event stream. Given the object mask at the initial time, our task aims to segment the complete object at any subsequent time in event streams. Specifically, our framework consists of a Recurrent Temporal Embedding Extraction (RTEE) module based on a novel ResLSTM, a Cross-time Spatiotemporal Feature Modeling (CSFM) module which is a transformer architecture with long-term and short-term matching modules, and a segmentation head. The historical events and masks (reference sets) are recurrently fed into our framework along with current-time events. The temporal embedding is updated as new events are input, enabling our framework to continuously process the event stream. To train and test our model, we construct both real-world and simulated event-based object segmentation datasets, each comprising event streams, APS images, and object annotations. Extensive experiments on our datasets demonstrate the effectiveness of the proposed recurrent architecture. Lin Zhu 0012, Xianzhang Chen, Lizhi Wang 0001, Xiao Wang 0014, Yonghong Tian 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | A Destriping Framework With Arbitrary Bounded Image DenoisersabstractExplicit stripes in digital images have long posed challenges for computer vision tasks, which significantly disturbs visual perceptions and is not desirable for subsequent applications. Existing methods for destriping tasks, often constrained by image prior assumptions and lacking flexibility, demonstrate limited practical applicability. In response to these challenges, this paper proposes a flexible Destriping framework with arbitrary bounded Image Denoisers, called DID. The proposed framework decouples the destriping task into conditional expectation calculation and stripe estimation, and alternates between these two parts, which finally obtains the maximum likelihood estimation of the image. The former calculates the conditional expectation of the clean image given the estimated stripe, while the latter estimates the mean of each column in the residual image. To calculate the conditional expectation, this paper analyzes the equivalence between general image denoising and conditional expectation calculation based on Bayesian statistics. It is proven that the proposed DID framework flexibly incorporates existing denoisers to calculate the conditional expectation, without the need to explicitly define image prior assumptions. Furthermore, the fixed-point convergence of the DID framework is guaranteed postulating that the applied denoiser is bounded in an F-norm manner. Experimental results on both synthesized and real data validate the effectiveness and generalization of the proposed method, both quantitatively and qualitatively. Chenyang Qian, Lingfei Song, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Enhancing Real-Time Object Detection With Optical Flow-Guided Streaming PerceptionabstractReal-time object detection in Unmanned Aerial Vehicle (UAV) videos remains a significant challenge due to the fast motion and small scale of objects. Existing streaming perception models struggle to accurately capture fine-grained motion cues between consecutive frames, leading to suboptimal performance in dynamic UAV scenarios. To address these challenges, Stream-Flow is proposed to integrate optical flow information and enhance real-time object detection in UAV videos. StreamFlow incorporates Flow-Guided Dynamic Prediction (FGDP) to refine position predictions using local optical flow information and Optical Flow Guided Optimization (OFGO) to optimize model parameters considering both localization loss and optical flow reliability. Central to OFGO is the Adaptive Flow Weighting (AFW) module, which focuses on reliable flow samples during training. The proposed integration of optical flow and adaptive weighting scheme significantly enhances the ability of streaming perception models to handle fast-moving objects in dynamic UAV environments. Extensive experiments on four challenging UAV video datasets demonstrate the superior performance of StreamFlow compared to state-of-the-art methods in terms of accuracy. Tongbo Wang, Lin Zhu 0012, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Simultaneous Learning Intensity and Optical Flow From High-Speed Spike StreamabstractBio-inspired vision sensors, which emulate the human retina by recording light intensity as binary spikes, have gained increasing interest in recent years. Among them, the spike camera is capable of perceiving fine textures by simulating a small retinal region called the fovea and producing high temporal resolution (20,000 Hz) spatiotemporal spike streams. To bridge the gap between binary spike streams and human vision in high-speed scenes, reconstructing intensity and optical flow from high temporal resolution spikes is particularly important. In this paper, we present a hybrid SNN-ANN network designed for simultaneous intensity and optical flow learning from spike streams. To adaptively extract spatial and temporal features from continuous spike streams, we propose a spiking neuron module with dense connections that efficiently processes both short-term and long-term spike data, while maintaining low power consumption characteristics. Subsequently, we introduce two decoders for optical flow and intensity estimation that complement each other. A temporal-aware warping module, based on flow features, is specifically designed to align the temporal features of the intensity decoder, thereby reducing motion artifacts. Concurrently, improved intensity features contribute to more accurate flow feature predictions, resulting in a mutually beneficial relationship within our network. To evaluate the effectiveness of our proposed network, we conduct experiments on both simulated and real spike datasets. Our network outperforms existing state-of-the-art spike-based reconstruction and optical flow estimation methods, demonstrating its potential for advancing the field of bio-inspired vision sensors. Our code is available athttps://github.com/LinZhu111/SLIO. Lin Zhu 0012, Weiquan Yan, Yi Chang 0002, Yonghong Tian 0001, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Infrared Video Dynamic Range Compression Based on Global and Local Temporal CoherenceabstractDynamic range compression (DRC) for infrared (IR) video aims to compress high dynamic range (HDR) IR video into low dynamic range (LDR) IR video, facilitating display on common devices and enhancing visual perception. This research requires balancing the preservation of spatial and temporal information, particularly detail preservation and temporal brightness coherence. However, existing methods fail to simultaneously maintain both global and local temporal coherence due to the lack of distinction between global and local motion, which inevitably introduces flickering and reduces the visual experience. To address this issue, this paper proposes an IR video DRC (vDRC) method based on global and local temporal coherence (GLTC). Specifically, a motion mask strategy based on structural similarity is introduced to differentiate between global motion regions, local motion regions, and static regions. A motion estimation strategy using different block-matching scales is then applied to estimate HDR motion information between consecutive frames, which is used to constrain the LDR of the previous frame and construct a temporal constraint term to preserve both global and local temporal coherence. Extensive experiments conducted on two public IR video datasets demonstrate that the proposed method outperforms state-of-the-art methods both quantitatively and qualitatively, offering a more effective solution for IR vDRC and enhancing visualization. Jinyi Qiu, Yuanfei Huang, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Bias Field Correction for Infrared Images Based on Surface Normal Consistency ConstraintsabstractDue to the mechanism and structure of infrared (IR) imaging devices, IR sensors inevitably capture thermal radiation from non-imaging sources, resulting in thermal radiational bias fields superimposed on IR images. These bias fields manifest as low-frequency nonuniformities, which are difficult to disentangle from low-frequency image regions (background planes and object surfaces). Existing methods frequently fail to reliably isolate true bias fields in IR images containing extensive flat regions, resulting in reduced contrast and degraded intensity fidelity in the corrected outputs. To overcome this limitation, we investigate the geometric patterns underlying IR image degradation and make a critical observation: the unbiased images typically contain extensive flat regions, while in biased images the surface normals of those regions exhibit strong consistency with surface normals of the bias field. Leveraging this insight, we propose a novel bias field correction method that utilizes surface normal consistency as a structural prior. The proposed method begins with the local planar fitting to estimate surface normals from the observed image efficiently. Subsequently, the cosine similarity is computed between surface normals of adjacent blocks to identify regions dominated solely by bias field normals. Finally, a least squares optimization problem is formulated and solved to estimate the bias field. In addition, by exploiting the temporal consistency of the bias field, this method could be extended to IR video via a motion-frame accumulation framework. Extensive experiments on synthesized and real data demonstrate the generality and effectiveness of the proposed method. Lingfei Song, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Specificity-Guided Cross-Modal Feature Reconstruction for RGB-Infrared Object DetectionabstractRGB-Infrared object detection is an essential technology for the intelligent transportation system. Existing most works on RGB-Infrared object detection focus on how to fuse RGB and infrared features. However, these works overlook the inherent differences between RGB and infrared modalities, leading to insufficient modal feature fusion and limiting the performance of RGB-Infrared object detection. To address the above issues, a Specificity-guided Cross-modal Feature Reconstruction(SCFR) algorithm is proposed to establish modality-specific correlation for RGB-Infrared object detection. Specifically, the proposed SCFR involves the modality-specific cross-modal feature reconstruction network and two modality-specific losses. The modality-specific cross-modal feature reconstruction network performs cross-modal feature reconstruction on RGB and infrared modalities to establish modality-specific correlation. The modality-specific losses guide the direction of feature learning for reconstructing the expressive modality-specific features. These specific features can be used to achieve more efficient feature fusion, thus improving object detection performance. Comprehensive experimental results on three RGB-Infrared detection datasets demonstrate the effectiveness and the superiority of the proposed method. Our code will be available athttps://github.com/SXYSUOSUO/SCFR.git. Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | GT-HAD: Gated Transformer for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to distinguish between the background and anomalies in a scene, which has been widely adopted in various applications. Deep neural network (DNN)-based methods have emerged as the predominant solution, wherein the standard paradigm is to discern the background and anomalies based on the error of self-supervised hyperspectral image (HSI) reconstruction. However, current DNN-based methods cannot guarantee correspondence between the background, anomalies, and reconstruction error, which limits the performance of HAD. In this article, we propose a novel gated transformer network for HAD (GT-HAD). Our key observation is that the spatial-spectral similarity in HSI can effectively distinguish between the background and anomalies, which aligns with the fundamental definition of HAD. Consequently, we develop GT-HAD to exploit the spatial-spectral similarity during HSI reconstruction. GT-HAD consists of two distinct branches that model the features of the background and anomalies, respectively, with content similarity as constraints. Furthermore, we introduce an adaptive gating unit to regulate the activation states of these two branches based on a content-matching method (CMM). Extensive experimental results demonstrate the superior performance of GT-HAD. The original code is publicly available at https://github.com/jeline0110/ GT-HAD, along with a comprehensive benchmark of state-of-the-art HAD methods. Lizhi Wang 0001, He Sun 0009, Hua Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | RGAvatar: Relightable 4D Gaussian Avatar From Monocular VideosabstractRelightable 4D avatar reconstruction which enables high fidelity and real-time rendering continues to be a crucial but challenging problem, especially from monocular videos. Previous NeRF-based 4D avatars enable photo-realistic relighting but are too slow for rendering, while point-based or mesh-based 4D avatars are efficient but have limited rendering quality. The recent success of 3D Gaussian Splatting, i.e., 3DGS, has inspired a series of impressive 4D Gaussian avatars, however, most of which only focus on faithful appearance reconstruction but are not relightable. To address such issues, this article proposes a new Relightable 4D Gaussian Avatar, i.e., RGAvatar, tailored for high fidelity relightable rendering from monocular videos. Our key idea is to introduce a new relightable 4D Gaussian representation, based on which we can directly perform high fidelity Physically Based Rendering, and an effective joint learning mechanism for compact 4D Gaussian reconstruction with SDF regulation and accurate materials and lighting decomposition. By comparing with previous state-of-the-art approaches, RGAvatar can significantly outperform previous approaches in relightable rendering quality and speed. To our best knowledge, RGAvatar contributes a new state-of-the-art 4D Gaussian avatar from monocular videos, which enables high fidelity relightable rendering in a quite efficient manner. Zhe Fan, Shi-Sheng Huang, Dachao Shang, Juyong Zhang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | MPGS: Multi-Plane Gaussian Splatting for Compact Scenes RenderingabstractAccurate reconstruction of heterogeneous scenes for high-fidelity rendering in an efficient manner remains a crucial but challenging task in many Virtual Reality and Augmented Reality applications. The recent 3D Gaussian Splatting (3DGS) has shown impressive quality in scene rendering with real-time performance. However, for heterogeneous scenes with many weak-textured regions, the original 3DGS can easily produce numerously wrong floaters with unbalanced reconstruction using redundant 3D Gaussians, which often leads to unsatisfied scene rendering. This paper proposes a novel multi-plane Gaussian Splatting (MPGS), which aims to achieve high-fidelity rendering with compact reconstruction for heterogeneous scenes. The key insight of our MPGS is the introduction of a novel multi-plane Gaussian optimization strategy, which effectively adjusts the Gaussian distribution for both rich-textured and weak-textured regions in heterogeneous scenes. Moreover, we further propose a multi-scale geometric correction mechanism to effectively mitigate degradation of the 3D Gaussian distribution for compact scene reconstruction. Besides, we regularize the Gaussian distributions using normal information extracted from the compact scene learning. Experimental results on public datasets demonstrate that the proposed MPGS achieves much better rendering quality compared to previous methods, while using less storage and offering more efficient rendering. To our best knowledge, MPGS is a new state-of-the-art 3D Gaussian splatting method for compact reconstruction of heterogeneous scenes, enabling high-fidelity rendering in novel view synthesis, especially improving rendering quality for weak-textured regions. The code will be released at https://github.com/wanglids/MPGS. Deqi Li, Shi-Sheng Huang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | EGAvatar: Efficient GAN Inversion for Generalizable Head Avatar From Few-Shot ImagesabstractControllable head avatar reconstruction via the inversion of few-shot images using 3D generative models has demonstrated significant potential for efficient avatar creation. However, under limited input conditions, existing one-shot inversion methods often fail to produce high-fidelity results, frequently leading to shape distortions, expression deviations, and identity inconsistencies. To address these limitations, we propose EGAvatar, a novel and efficient 3DGAN inversion framework designed to generate high-fidelity, generalizable head avatars from few-shot images. The core principle of EGAvatar is a decoupling-by-inverting strategy, built upon an animatable 3DGAN prior. Specifically, we introduce an effective animatable 3DGAN model that synthesizes high-quality 3D avatars by integrating a coarse 3D triplane representation (derived from a latent 3DGAN) with an offset 3D triplane (learned via a triplane 3DGAN). Leveraging this architecture, we design a 3DGAN-based inversion approach to reconstruct 3D avatars efficiently. Additionally, we incorporate an expression-view disentanglement mechanism to maintain consistent appearance across varying expressions and viewpoints, thereby enhancing the generalizability of avatar reconstruction from limited input images. Extensive experiments conducted on two publicly available benchmarks and a private dataset demonstrate that EGAvatar outperforms existing state-of-the-art methods in both qualitative and quantitative evaluations. Notably, EGAvatar achieves superior performance while requiring significantly fewer input images and offering more efficient training and inference. Hao Pan Ren, Wan Yu Li, Shi-Sheng Huang, Juyong Zhang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | Finding Visual Saliency in Continuous Spike StreamabstractAs a bio-inspired vision sensor, the spike camera emulates the operational principles of the fovea, a compact retinal region, by employing spike discharges to encode the accumulation of per-pixel luminance intensity. Leveraging its high temporal resolution and bio-inspired neuromorphic design, the spike camera holds significant promise for advancing computer vision applications. Saliency detection mimic the behavior of human beings and capture the most salient region from the scenes. In this paper, we investigate the visual saliency in the continuous spike stream for the first time. To effectively process the binary spike stream, we propose a Recurrent Spiking Transformer (RST) framework, which is based on a full spiking neural network. Our framework enables the extraction of spatio-temporal features from the continuous spatio-temporal spike stream while maintaining low power consumption. To facilitate the training and validation of our proposed model, we build a comprehensive real-world spike-based visual saliency dataset, enriched with numerous light conditions. Extensive experiments demonstrate the superior performance of our Recurrent Spiking Transformer framework in comparison to other spike neural network-based methods. Our framework exhibits a substantial margin of improvement in capturing and highlighting visual saliency in the spike stream, which not only provides a new perspective for spike-based saliency segmentation but also shows a new paradigm for full SNN-based transformer models. The code and dataset are available at https://github.com/BIT-Vision/SVS. Lin Zhu 0012, Xianzhang Chen, Xiao Wang 0014, Hua Huang 0001 |
AAAI | 4 |
| 2024 | SpikeNeRF: Learning Neural Radiance Fields from Continuous Spike StreamabstractSpike cameras, leveraging spike-based integration sampling and high temporal resolution, offer distinct advan-tages over standard cameras. However, existing approaches reliant on spike cameras often assume optimal illumination, a condition frequently unmet in real-world scenarios. To address this, we introduce SpikeNeRF, the first work that derives a NeRF-based volumetric scene representation from spike camera data. Our approach leverages NeRF's multi-view consistency to establish robust self-supervision, effectively eliminating erroneous measurements and uncovering coherent structures within exceedingly noisy input amidst diverse real-world illumination scenarios. The framework comprises two core elements: a spike generation model incorporating an integrate-and-fire neuron layer and parameters accounting for non-idealities, such as threshold variation, and a spike rendering loss capable of general-izing across varying illumination conditions. We describe how to effectively optimize neural radiance fields to render photorealistic novel views from the novel continuous spike stream, demonstrating advantages over other vision sen-sors in certain scenes. Empirical evaluations conducted on both real and novel realistically simulated sequences affirm the efficacy of our methodology. The dataset and source code are released at https://github.com/BIT-Vision/SpikeNeRF. Lin Zhu 0012, Kangmin Jia, Yifan Zhao 0002, Yunshan Qi, Lizhi Wang 0001, Hua Huang 0001 |
CVPR | 6 |
| 2024 | In2SET: Intra-Inter Similarity Exploiting Transformer for Dual-Camera Compressive Hyperspectral ImagingabstractDual-camera compressive hyperspectral imaging (DC-CHI) offers the capability to reconstruct 3D hyperspectral image (HSI) by fusing compressive and panchromatic (PAN) image, which has shown great potential for snapshot hyperspectral imaging in practice. In this paper, we introduce a novel DCCHI reconstruction network, intra-inter similarity exploiting Transformer (In2SET). Our key insight is to make full use of the PAN image to assist the reconstruction. To this end, we propose to use the intra-similarity within the PAN image as a proxy for approximating the intra-similarity in the original HSI, thereby offering an enhanced content prior for more accurate HSI reconstruction. Furthermore, we propose to use the inter-similarity to align the features between HSI and PAN images, thereby maintaining semantic consistency between the two modalities during the reconstruction process. By integrating In2SET into a PAN-guided deep unrolling (PGDU)framework, our method substantially enhances the spatial-spectral fidelity and detail of the reconstructed images, providing a more comprehensive and accurate depiction of the scene. Experiments conducted on both real and simulated datasets demonstrate that our approach consistently outperforms existing state-of-the-art methods in terms of reconstruction quality and computational complexity. The code is available at https://github.com/2JONAS/In2SET. Lizhi Wang 0001, Xiangtian Ma, Maoqing Zhang, Lin Zhu 0012, Hua Huang 0001 |
CVPR | 6 |
| 2024 | Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
Lin Zhu 0012, Yunlong Zheng, Yijun Zhang 0003, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ECCV (40) | 6 |
| 2024 | NeuralIndicator: Implicit Surface Reconstruction from Neural Indicator PriorsabstractThe neural implicit surface reconstruction from unorganized points is still challenging, especially when the point clouds are incomplete and/or noisy with complex topology structure. Unlike previous approaches performing neural implicit surface learning relying on local shape priors, this paper proposes to utilize global shape priors to regularize the neural implicit function learning for more reliable surface reconstruction. To this end, we first introduce a differentiable module to generate a smooth indicator function, which globally encodes both the indicative prior and local SDFs of the entire input point cloud. Benefit from this, we propose a new framework, called NeuralIndicator, to jointly learn both the smooth indicator function and neural implicit function simultaneously, using the global shape prior encoded by smooth indicator function to effectively regularize the neural implicit function learning, towards reliable and high-fidelity surface reconstruction from unorganized points without any normal information. Extensive evaluations on synthetic and real-scan datasets show that our approach consistently outperforms previous approaches, especially when point clouds are incomplete and/or noisy with complex topology structure. Shi-Sheng Huang, Chen Li Heng, Hua Huang 0001 |
ICML | 4 |
| 2024 | AutoSFX: Automatic Sound Effect Generation for VideosabstractSound Effect (SFX) generation, primarily aims to automatically produce sound waves for sounding visual objects in images or videos. Rather than learning an automatic solution to this task, we aim to propose a much broader system, AutoSFX, designed to automate sound design for videos in a more efficient and applicable manner. AutoSFX capitalizes on this concept by aggregating multimodal representations by cross-attention and leverages a diffusion model to generate sound with visual information embedded. AutoSFX also optimizes the generated sounds to render the entire soundtrack for the input video, leading to a more immersive and engaging multimedia experience by performing seamless transitions between sound clips and harmoniously mixing sounds playing simultaneously. We have developed a user-friendly interface for AutoSFX enabling users to interactively engage in the SFX generation for their videos with particular needs. To validate the capability of our vision-to-sound generation, we conducted comprehensive experiments and analyses using the widely recognized VEGAS and VGGSound test sets, yielding promising results. We also conducted a user study to evaluate the performance of the optimized soundtrack and the usability of the interface. Overall, the results revealed that our AutoSFX provides a viable sound landscape solution for making attractive videos. Zhongxu Wang, Hua Huang 0001 |
ACM Multimedia | 3 |
| 2024 | ArtSpeech: Adaptive Text-to-Speech Synthesis with Articulatory Representations
Zhongxu Wang, Mingzhu Li, Hua Huang 0001 |
ACM Multimedia | 4 |
| 2024 | 4-bit Shampoo for Memory-Efficient Network TrainingabstractSecond-order optimizers, maintaining a matrix termed a preconditioner, are superior to first-order optimizers in both theory and practice.
The states forming the preconditioner and its inverse root restrict the maximum size of models trained by second-order optimizers. To address this, compressing 32-bit optimizer states to lower bitwidths has shown promise in reducing memory usage. However, current approaches only pertain to first-order optimizers. In this paper, we propose the first 4-bit second-order optimizers, exemplified by 4-bit Shampoo, maintaining performance similar to that of 32-bit ones. We show that quantizing the eigenvector matrix of the preconditioner in 4-bit Shampoo is remarkably better than quantizing the preconditioner itself both theoretically and experimentally. By rectifying the orthogonality of the quantized eigenvector matrix, we enhance the approximation of the preconditioner's eigenvector matrix, which also benefits the computation of its inverse 4-th root. Besides, we find that linear square quantization slightly outperforms dynamic tree quantization when quantizing second-order optimizer states. Evaluation on various networks for image classification and natural language modeling demonstrates that our 4-bit Shampoo achieves comparable performance to its 32-bit counterpart while being more memory-efficient. Sike Wang, Pan Zhou 0002, Jia Li 0002, Hua Huang 0001 |
NeurIPS | 4 |
| 2024 | Robust Image Restoration with an Adaptive Huber Function Based Fidelity
Lingfei Song, Hua Huang 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Learning a physics-based filter attachment for hyperspectral imaging with RGB cameras
Maoqing Zhang, Lizhi Wang 0001, Lin Zhu 0012, Hua Huang 0001 |
Neurocomputing | 4 |
| 2024 | Learnability Enhancement for Low-Light Raw Image Denoising: A Data PerspectiveabstractLow-light raw image denoising is an essential task in computational photography, to which the learning-based method has become the mainstream solution. The standard paradigm of the learning-based method is to learn the mapping between the paired real data, i.e., the low-light noisy image and its clean counterpart. However, the limited data volume, complicated noise model, and underdeveloped data quality have constituted the learnability bottleneck of the data mapping between paired real data, which limits the performance of the learning-based method. To break through the bottleneck, we introduce a learnability enhancement strategy for low-light raw image denoising by reforming paired real data according to noise modeling. Our learnability enhancement strategy integrates three efficient methods: shot noise augmentation (SNA), dark shading correction (DSC) and a developed image acquisition protocol. Specifically, SNA promotes the precision of data mapping by increasing the data volume of paired real data, DSC promotes the accuracy of data mapping by reducing the noise complexity, and the developed image acquisition protocol promotes the reliability of data mapping by improving the data quality of paired real data. Meanwhile, based on the developed image acquisition protocol, we build a new dataset for low-light raw image denoising. Experiments on public datasets and our dataset demonstrate the superiority of the learnability enhancement strategy. Hansen Feng, Lizhi Wang 0001, Yuzhi Wang, Haoqiang Fan, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Stimulating Diffusion Model for Image Denoising via Adaptive Embedding and EnsemblingabstractImage denoising is a fundamental problem in computational photography, where achieving high perception with low distortion is highly demanding. Current methods either struggle with perceptual quality or suffer from significant distortion. Recently, the emerging diffusion model has achieved state-of-the-art performance in various tasks and demonstrates great potential for image denoising. However, stimulating diffusion models for image denoising is not straightforward and requires solving several critical problems. For one thing, the input inconsistency hinders the connection between diffusion models and image denoising. For another, the content inconsistency between the generated image and the desired denoised image introduces distortion. To tackle these problems, we present a novel strategy called the Diffusion Model for Image Denoising (DMID) by understanding and rethinking the diffusion model from a denoising perspective. Our DMID strategy includes an adaptive embedding method that embeds the noisy image into a pre-trained unconditional diffusion model and an adaptive ensembling method that reduces distortion in the denoised image. Our DMID strategy achieves state-of-the-art performance on both distortion-based and perception-based metrics, for both Gaussian and real-world image denoising. Tong Li 0016, Hansen Feng, Lizhi Wang 0001, Lin Zhu 0012, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Non-Serial Quantization-Aware Deep Optics for Snapshot Hyperspectral ImagingabstractDeep optics has been endeavoring to capture hyperspectral images of dynamic scenes, where the optical encoder plays an essential role in deciding the imaging performance. Our key insight is that the optical encoder of a deep optics system is expected to keep fabrication-friendliness and decoder-friendliness, to be faithfully realized in the implementation phase and fully interacted with the decoder in the design phase, respectively. In this paper, we propose the non-serial quantization-aware deep optics (NSQDO), which consists of the fabrication-friendly quantization-aware model (QAM) and the decoder-friendly non-serial manner (NSM). The QAM integrates the quantization process into the optimization and adaptively adjusts the physical height of each quantization level, reducing the deviation of the physical encoder from the numerical simulation through the awareness of and adaptation to the quantization operation of the DOE physical structure. The NSM bridges the encoder and the decoder with full interaction through bidirectional hint connections and flexibilize the connections with a gating mechanism, boosting the power of joint optimization in deep optics. The proposed NSQDO improves the fabrication-friendliness and decoder-friendliness of the encoder and develops the deep optics framework to be more practical and powerful. Extensive synthetic simulation and real hardware experiments demonstrate the superior performance of the proposed method. Lizhi Wang 0001, Lingen Li, Lei Zhang 0021, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Deep Convolution Modulation for Image Super-ResolutionabstractRecently, deep-learning-based super-resolution methods have achieved excellent performances, but mainly focus on training a single generalized deep network by feeding numerous samples. Yet intuitively, each image has its specific representation, and is expected to acquire an adaptive model. For this issue, we propose a novel convolution modulation (CoMo) mechanism to build image-specific deep networks, by exploiting the principal information of the feature to generate a modulation weight, and thereby adaptively modulating the kernel weights of convolution without any additional parameters, which outperforms the vanilla convolution and several existing attention mechanisms when embedding into the state-of-the-art architectures. To optimize the modulated convolutions in mini-batch training, we introduce an image-specific optimization (IsO) algorithm, which tackles the infeasibility of the conventional optimization algorithms on this issue. Furthermore, we investigate the effect of CoMo on state-of-the-art architectures and design a new CoMoNet architecture by employing the U-style residual learning and hourglass dense block learning, which is an appropriate architecture to utmost improve the effectiveness of CoMo theoretically. Extensive experiments on benchmarks show that the proposed methods achieve superior performances and higher flexibility against the state-of-the-art SISR and blind SR methods. The code is available at github.com/YuanfeiHuang/CoMoNet. Yuanfei Huang, Jie Li 0001, Yanting Hu, Hua Huang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Future Feature-Based Supervised Contrastive Learning for Streaming PerceptionabstractStreaming perception, a critical task in computer vision, involves the real-time prediction of object locations within video sequences based on prior frames. While current methods like StreamYOLO mainly rely on coordinate information, they often fall short of delivering precise predictions due to feature misalignment between input data and supervisory labels. In this paper, a novel method, Future Feature-based Supervised Contrastive Learning (FFSCL), is introduced to address this challenge by incorporating appearance features from future frames and leveraging supervised contrastive learning techniques. FFSCL establishes a robust correspondence between the appearance of an object in current and past frames and its location in the subsequent frame. This integrated method significantly improves the accuracy of object position prediction in streaming perception tasks. In addition, the FFSCL method includes a sample pair construction module (SPC) for the efficient creation of positive and negative samples based on future frame labels and a feature consistency loss (FCL) to enhance the effectiveness of supervised contrastive learning by linking appearance features from future frames with those from past frames. The efficacy of FFSCL is demonstrated through extensive experiments on two large-scale benchmark datasets, where FFSCL consistently outperforms state-of-the-art methods in streaming perception tasks. This study represents a significant advancement in the incorporation of supervised contrastive learning techniques and future frame information into the realm of streaming perception, paving the way for more accurate and efficient object prediction within video streams. Tongbo Wang, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Infrared Image Dynamic Range Compression Based on Adaptive Contrast Adjustment and Structure PreservationabstractThe infrared (IR) image dynamic range compression (DRC) technology involves compressing high dynamic range (HDR) IR images into low dynamic range (LDR) images for display on common devices. To facilitate human observation, DRC methods should preserve the structural information as much as possible while adjusting the contrast of HDR IR images. However, existing DRC methods struggle to adapt to various highly dynamic IR scenes when using fixed parameter settings. To address this limitation, a novel gradient domain-based DRC method with adaptive contrast adjustment and structure preservation (ACASP) is proposed. Our ACASP adapts local contrast and gradients by analyzing local features of HDR IR images, effectively handling different HDR IR scenes. We introduce local contrast and variance to enhance visibility in low-contrast areas and preserve details in high-contrast areas. Specifically, we design a contrast-adaptive mapping curve and a gradient-adaptive modulation factor (GMF) to optimize both contrast and structure in the LDR image. Extensive experiments on three public HDR IR datasets demonstrate that the proposed method can outperform state-of-the-art DRC methods in both quantitative and qualitative analyses. This work contributes to the field by offering a more adaptive and robust approach to IR image DRC. Jinyi Qiu, Zhan Wang 0007, Yuanfei Huang, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Thermal Radiation Bias Correction for Infrared Images Using Huber Function-Based LossabstractLimited by the imaging mechanism, thermal radiation emitted from infrared (IR) imaging devices commonly contaminates the detector response to the scene, causing an additive thermal radiation bias at each pixel that dramatically reduces the contrast of the image. This degradation is considered a bias field over the image, impacting the visual perception and subsequent applications. Therefore, eliminating the thermal radiation bias field is an urgent issue. This paper proposes an adaptive thermal radiation bias field correction method using Huber function-based loss, which can adapt to image contents and hence preserve meaningful details while eliminating the bias field. The proposed method introduces the low-order bivariate polynomial surface model to fit the bias field from the observed image precisely. We establish a robust objective function based on the Huber function to estimate parameters of the surface model, which can adaptively switch between the ℓ1norm and the ℓ2norm-based loss functions according to the image region. Thus, our method not only effectively maintains optimality in flat regions but also improves robustness in edge & texture regions. To balance efficiency and robustness, we propose an adaptive threshold that controls the behavior of the Huber function. For stable convergence, an improved gradient descent strategy is utilized to solve the Huber loss-based objective function with the two-direction fitting and an adjustable step. Both simulated and real experiments against classical and state-of-the-art methods demonstrate the superior performance of the proposed method in improving the contrast and preserving details. Lingfei Song, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | IMU-Assisted Accurate Blur Kernel Re-Estimation in Non-Uniform Camera Shake DeblurringabstractImage deblurring for camera shake is a highly regarded problem in the field of computer vision. A promising solution is the patch-wise non-uniform image deblurring algorithms, where a linear transformation model is typically established between different blur kernels to re-estimate poorly estimated blur kernels. However, the linear model struggles to effectively describe the nonlinear transformation relationships between blur kernels. A key observation is that the inertial measurement unit (IMU) provides motion data of the camera, which is helpful in describing the landmarks of the blur kernel. This paper presents a new IMU-assisted method for the re-estimation of poorly estimated blur kernels. This method establishes a nonlinear transformation relationship model between blur kernels of different patches using IMU motion data. Subsequently, an optimization problem is applied to re-estimate poorly estimated blur kernels by incorporating this relationship model with neighboring well-estimated kernels. Experimental results demonstrate that this blur kernel re-estimation method outperforms existing methods. Jianxiang Rong, Hua Huang 0001, Jia Li 0002 |
IEEE Trans. Image Process. | 2 |
| 2024 | Balanced Federated Semisupervised Learning With Fairness-Aware Pseudo-LabelingabstractFederated semisupervised learning (FSSL) aims to train models with both labeled and unlabeled data in the federated settings, enabling performance improvement and easier deployment in realistic scenarios. However, the nonindependently identical distributed data in clients leads to imbalanced model training due to the unfair learning effects on different classes. As a result, the federated model exhibits inconsistent performance on not only different classes, but also different clients. This article presents a balanced FSSL method with the fairness-aware pseudo-labeling (FAPL) strategy to tackle the fairness issue. Specifically, this strategy globally balances the total number of unlabeled data samples which is capable to participate in model training. Then, the global numerical restrictions are further decomposed into personalized local restrictions for each client to assist the local pseudo-labeling. Consequently, this method derives a more fair federated model for all clients and gains better performance. Experiments on image classification datasets demonstrate the superiority of the proposed method over the state-of-the-art FSSL methods. Xiao-Xiang Wei, Hua Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Revisiting Unsupervised Local Descriptor LearningabstractConstructing accurate training tuples is crucial for unsupervised local descriptor learning, yet challenging due to the absence of patch labels. The state-of-the-art approach constructs tuples with heuristic rules, which struggle to precisely depict real-world patch transformations, in spite of enabling fast model convergence. A possible solution to alleviate the problem is the clustering-based approach, which can capture realistic patch variations and learn more accurate class decision boundaries, but suffers from slow model convergence. This paper presents HybridDesc, an unsupervised approach that learns powerful local descriptor models with fast convergence speed by combining the rule-based and clustering-based approaches to construct training tuples. In addition, HybridDesc also contributes two concrete enhancing mechanisms: (1) a Differentiable Hyperparameter Search (DHS) strategy to find the optimal hyperparameter setting of the rule-based approach so as to provide accurate prior for the clustering-based approach, (2) an On-Demand Clustering (ODC) method to reduce the clustering overhead of the clustering-based approach without eroding its advantage. Extensive experimental results show that HybridDesc can efficiently learn local descriptors that surpass existing unsupervised local descriptors and even rival competitive supervised ones. Wufan Wang, Lei Zhang 0021, Hua Huang 0001 |
AAAI | 3 |
| 2023 | Dynamic View Synthesis with Spatio-Temporal Feature Warping from Sparse ViewsabstractSignificant progress has been made in realizing novel view synthesis of dynamic scenes from sparse input views. However, achieving spatio-temporal consistency in dynamic view synthesis remains to be challenging for previous approaches, since the spatio-temporal correlation for view synthesis has not been fully explored. In this paper, we propose a spatio-temporal feature warping (STFW) mechanism, which can be embedded into a deep model to produce high-quality and spatio-temporally consistent view synthesis results. The two core components of STFW are: (1) a spatial feature warping (SFW) module, which enables adaptive perception of multi-view context-consistent geometric information with a compact point cloud representation, and (2) a temporal feature warping (TFW) module that implicitly models the dynamic geometry by approaching the pixel shift in image coordinate. In the optimization process of view synthesis, the SFW and TFW are integrated to exploit the spatio-temporal correlation cues across sparse input views and novel views. Leveraging the STFW, we further build an end-to-end dynamic view synthesis model with sparse input views. Qualitative and quantitative evaluation on public multi-view datasets demonstrate that our view synthesis pipeline achieves better performance compared to previous methods in terms of visual quality. Deqi Li, Shi-Sheng Huang, Tianyu Shen, Hua Huang 0001 |
ACM Multimedia | 4 |
| 2023 | Learning Spectral-wise Correlation for Spectral Super-Resolution: Where Similarity Meets ParticularityabstractHyperspectral images consist of multiple spectral channels, and the task of spectral super-resolution is to reconstruct hyperspectral images from 3-channel RGB images, where modeling spectral-wise correlation is of great importance. Based on the analysis of the physical process of this task, we distinguish the spectral-wise correlation into two aspects: similarity and particularity. The Existing Transformer model cannot accurately capture spectral-wise similarity due to the inappropriate spectral-wise fully connected linear mapping acting on input spectral feature maps, which results in spectral feature maps mixing. Moreover, the token normalization operation in the existing Transformer model also results in its inability to capture spectral-wise particularity and thus fails to extract key spectral feature maps. To address these issues, we propose a novel Hybrid Spectral-wise Attention Transformer (HySAT). The key module of HySAT is Plausible Spectral-wise self-Attention (PSA), which can simultaneously model spectral-wise similarity and particularity. Specifically, we propose a Token Independent Mapping (TIM) mechanism to reasonably model spectral-wise similarity, where a linear mapping shared by spectral feature maps is applied on input spectral feature maps. Moreover, we propose a Spectral-wise Re-Calibration (SRC) mechanism to model spectral-wise particularity and effectively capture significant spectral feature maps. Experimental results show that our method achieves state-of-the-art performance in the field of spectral super-resolution with the lowest error and computational costs. Lizhi Wang 0001, Chang Chen 0004, Fenglong Song, Hua Huang 0001 |
ACM Multimedia | 6 |
| 2023 | Recurrent Spike-based Image Restoration under General IlluminationabstractSpike camera is a new type of bio-inspired vision sensor that records light intensity in the form of a spike array with high temporal resolution (20,000 Hz). This new paradigm of vision sensor offers significant advantages for many vision tasks such as high speed image reconstruction. However, existing spike-based approaches typically assume that the scenes are with sufficient light intensity, which is usually unavailable in many real-world scenarios such as rainy days or dusk scenes. To unlock more spike-based application scenarios, we propose a Recurrent Spike-based Image Restoration (RSIR) network, which is the first work towards restoring clear images from spike arrays under general illumination. Specifically, to accurately describe the noise distribution under different illuminations, we build a physical-based spike noise model according to the sampling process of the spike camera. Based on the noise model, we design our RSIR network which consists of an adaptive spike transformation module, a recurrent temporal feature fusion module, and a frequency-based spike denoising module. Our RSIR can process the spike array in a recursive manner to ensure that the spike temporal information is well utilized. In the training process, we generate the simulated spike data based on our noise model to train our network. Extensive experiments on real-world datasets with different illuminations demonstrate the effectiveness of the proposed network. The code and dataset are released at https://github.com/BIT-Vision/RSIR. Lin Zhu 0012, Yunlong Zheng, Mengyue Geng, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 5 |
| 2023 | An efficient fine-grained vehicle recognition method based on part-level feature optimization
Yancheng Cai, Hua Huang 0001, Ping Wang 0009 |
Neurocomputing | 3 |
| 2023 | Transitional Learning: Exploring the Transition States of Degradation for Blind Super-resolutionabstractBeing extremely dependent on iterative estimation of the degradation prior or optimization of the model from scratch, the existing blind super-resolution (SR) methods are generally time-consuming and less effective, as the estimation of degradation proceeds from a blind initialization and lacks interpretable representation of degradations. To address it, this article proposes a transitional learning method for blind SR using an end-to-end network without any additional iterations in inference, and explores an effective representation for unknown degradation. To begin with, we analyze and demonstrate the transitionality of degradations as interpretable prior information to indirectly infer the unknown degradation model, including the widely used additive and convolutive degradations. We then propose a novel Transitional Learning method for blind Super-Resolution (TLSR), by adaptively inferring a transitional transformation function to solve the unknown degradations without any iterative operations in inference. Specifically, the end-to-end TLSR network consists of a degree of transitionality (DoT) estimation network, a homogeneous feature extraction network, and a transitional learning module. Quantitative and qualitative evaluations on blind SR tasks demonstrate that the proposed TLSR achieves superior performances and costs fewer complexities against the state-of-the-art blind SR methods. The code is available at github.com/YuanfeiHuang/TLSR. Yuanfei Huang, Jie Li 0001, Yanting Hu, Xinbo Gao 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Multi-Oriented Object Detection in Aerial Images With Double Horizontal RectanglesabstractMost existing methods adopt the quadrilateral or rotated rectangle representation to detect multi-oriented objects. Yet, the same oriented object may correspond to several different representations, due to different vertex ordering, or angular periodicity and edge exchangeability. To ensure the uniqueness of the representation, some engineered rules are usually added. This makes these methods suffer from discontinuity problem, resulting in degraded performance for objects around some orientation. In this article, we propose to encode the multi-oriented object with double horizontal rectangles (DHRec) to solve the discontinuity problem. Specifically, for an oriented object, we arrange the horizontal and vertical coordinates of its four vertices in left-right and top-down order, respectively. The first (resp. second) horizontal box is given by two diagonal points with smallest (resp. second) and third (resp. largest) coordinates in both horizontal and vertical dimensions. We then regress three factors given by area ratios between different regions, helping to guide the oriented object decoding from the predicted DHRec. Inherited from the uniqueness of horizontal rectangle representation, the proposed method is free of discontinuity issue, and can accurately detect objects of arbitrary orientation. Extensive experimental results show that the proposed method significantly improves the existing baseline representation, and outperforms state-of-the-art methods. The code is available at: https://github.com/lightbillow/DHRec. Guangtao Nie, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Fixed Pattern Noise Removal Based on a Semi-Calibration MethodabstractDue to the manufacturing imperfections, nonuniformities are ubiquitous in digital sensors, causing the notorious Fixed Pattern Noise (FPN). The ability of modern digital cameras to take images under low-light environments is severely limited by the FPN. This paper proposes a novel semi-calibration-based method for the FPN removal that utilizes a pre-calibrated Noise Pattern. The key observation of this work is that the FPN in each shot is actually a scaled Noise Pattern with an unknown scale parameter, since each pixel in the array generates a characteristic amount of dark current which is fundamentally determined by its physical properties. Given a noised image and the corresponding Noise Pattern, the scale parameter is automatically estimated, and then the FPN is removed by subtracting the scaled Noise Pattern from the noised image. The estimation of the scale parameter is based on an entropy minimization estimator, which is derived from the Maximum Likelihood principle and is further justified by subsequent analysis that minimizing the entropy uniquely identifies the true parameter. Convergence issues, as well as the optimality of the proposed estimator, are also theoretically discussed. Finally, some applications are given, illustrating the performance of the proposed FPN removal method in real-world tasks. Lingfei Song, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Universal Demosaicking for Interpolation-Friendly RGBW Color Filter ArraysabstractInterpolation-friendly RGBW color filter arrays (CFAs) and the popular sequential demosaicking contain the idea of computational photography, where the CFA and the demosaicking method are co-designed. Due to the advantages, interpolation-friendly RGBW CFAs have been extensively used in commercial color cameras. However, most associated demosaicking methods rely on strict assumptions or are limited to a few specific CFAs with a given camera. In this paper, we propose a universal demosaicking method for interpolation-friendly RGBW CFAs, which enables the comparison of different CFAs. Our new method belongs to sequential demosaicking, i.e., W channel is interpolated first and then RGB channels are reconstructed with guidance from the interpolated W channel. Specifically, it first interpolates the W channel using only available W pixels followed by an aliasing reduction technique to remove aliasing artifacts. Then it employs an image decomposition model to built relations between W channel and each of RGB channels with known RGB values, which can be easily generalized to the full-size demosaicked image. We apply the linearized alternating direction method (LADM) to solve it with convergence guarantee. Our demosaicking method can be applied to all interpolation-friendly RGBW CFAs with varying color cameras and lighting conditions. Extensive experiments confirm the universal property and advantage of our proposed method with both simulated and real raw images. Jia Li 0002, Chen-Yan Bai, Hua Huang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Simultaneous Destriping and Image Denoising Using a Nonparametric Model With the EM AlgorithmabstractDigital images often suffer from the common problem of stripe noise due to the inconsistent bias of each column. The existence of the stripe poses much more difficulties on image denoising since it requires another ${n}$ parameters, where ${n}$ is the width of the image, to characterize the total interference of the observed image. This paper proposes a novel EM-based framework for simultaneous stripe estimation and image denoising. The great benefit of the proposed framework is that it splits the overall destriping and denoising problem into two independent sub-problems, i.e., calculating the conditional expectation of the true image given the observation and the estimated stripe from the last round of iteration, and estimating the column means of the residual image, such that a Maximum Likelihood Estimation (MLE) is guaranteed and it does not require any explicit parametric modeling of image priors. The calculation of the conditional expectation is the key, here we choose a modified Non-Local Means algorithm to calculate the conditional expectation because it has been proven to be a consistent estimator under some conditions. Besides, if we relax the consistency requirement, the conditional expectation could be interpreted as a general image denoiser. Therefore other state-of-the-art image denoising algorithms have the potentials to be incorporated into the proposed framework. Extensive experiments have demonstrated the superior performance of the proposed algorithm and provide some promising results that motivate future research on the EM-based destriping and denoising framework. Lingfei Song, Hua Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Edge Devices Clustering for Federated Visual Classification: A Feature Norm Based FrameworkabstractFederated learning is a privacy-preserving distributed learning paradigm where multiple devices collaboratively train a model, which is applicable to edge computing environments. However, the non-IID data distributed in multiple devices degrades the performance of the federated model due to severe weight divergence. This paper presents a clustered federated learning framework named cFedFN for visual classification tasks in order to reduce the degradation. Especially, this framework introduces the computation of feature norm vectors in the local training process and divides the devices into multiple groups by the similarities of the data distributions to reduce the weight divergences for better performance. As a result, this framework gains better performance on non-IID data without leakage of the private raw data. Experiments on various visual classification datasets demonstrate the superiority of this framework over the state-of-the-art clustered federated learning frameworks. Xiao-Xiang Wei, Hua Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Multi-Modal Feature Pyramid Transformer for RGB-Infrared Object DetectionabstractRGB-Infrared multi-modal object detection utilizes diverse and complementary information, showing some advantages in intelligent transportation field. The main challenge of RGB-Infrared object detection is how to fuse the two modalities. The difficulty of fusion is reflected in two aspects: 1) large visual differences between modalities make it difficult to learn effective complementary features, 2) some misaligned RGB-Infrared images increase the difficulty of fusion. To this end, based on feature pyramid commonly used in object detection, we propose Multi-modal Feature Pyramid Transformer (MFPT) to fuse the two modalities. The proposed MFPT learns semantic and modal complementary information to enhance each modal features via intra-modal feature pyramid transformer and inter-modal feature pyramid transformer. The intra-modal feature pyramid transformer enables features to interact across space and scales, improving the semantic representations of features in each modality. The inter-modal feature pyramid transformer conducts feature interaction between modalities, enabling each modality to learn complementary features from other modalities. Meanwhile, the inter-modal feature pyramid transformer can also learn distance independent dependencies between modalities, which are not sensitive to misaligned images. Furthermore, a local attention mechanism is introduced within different windows into MFPT to achieve efficient correlation between regions of different scales or different modalities. Experimental results on two RGB-Infrared detection datasets demonstrate the proposed method is superior to state-of-the-art methods. Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Depth-Aware Multi-Person 3D Pose Estimation With Multi-Scale Waterfall RepresentationsabstractEstimating absolute 3D poses of multiple people from monocular image is challenging due to the presence of occlusions and the scale variation among different persons. Among the existing methods, the top-down paradigms are highly dependent on human detection which is prone to the influence from inter-person occlusions, while the bottom-up paradigms suffer from the difficulties in keypoint feature extraction caused by scale variation and unreliable joint grouping caused by occlusions. To address these challenges, we introduce a novel multi-person 3D pose estimation framework, aided by multi-scale feature representations and human depth perceiving. Firstly, a waterfall-based architecture is incorporated for multi-scale feature representations to achieve a more accurate estimation of occluded joints with a better detection of human shapes. Then the global and local representations are fused for handling the effects of inter-person occlusion and scale variation in depth perceiving and keypoint feature extraction. Finally, with the guidance of the fused multi-scale representations, a depth-aware model is exploited for better 2D joint grouping and 3D pose recovering. Quantitative and qualitative evaluations on benchmark datasets of MuCo-3DHP and MuPoTS-3D prove the effectiveness of our proposed method. Furthermore, we produce an occluded MuPoTS-3D dataset and the experiments on it validate the superiority of our method for overcoming the occlusions. Tianyu Shen, Deqi Li, Fei-Yue Wang 0001, Hua Huang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Image Stitching With Manifold OptimizationabstractImage stitching usually relies on spatial transformations to perform the overlap alignment and distortion mitigation. This paper presents a manifold optimization method to seek these transformations. The purpose is not to present a new formulation of image stitching, as the proposed method uses common transformations such as homography to align feature correspondences in the overlap and similarity transformations to preserve the shape. Instead, the proposed method is based on a new treatment of these transformations as elements of a prescribed matrix manifold. Its advantage lies in its more effective and efficient optimization in the manifold domain. Specifically, spatially varying homographies are computed by an efficient second-order minimization (ESM) of the geometric error of aligning feature correspondences, but with their intrinsic manifold parameterization. To mitigate the distortion, the interpolation between homography and similarity transformation is performed on a general matrix manifold. These on-manifold operations improve the stitching quality with fewer ghosting and distortion artifacts. The experiments show our manifold optimization for image stitching outperforms other methods. Lei Zhang 0021, Hua Huang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | VirtualClassroom: A Lecturer-Centered Consumer-Grade Immersive Teaching System in Cyber-Physical-Social SpaceabstractLecturers, as the guidance of the classroom, play a significant role in the teaching process. However, the lecturers’ sense of space immersion has been ignored in current virtual teaching systems. In this article, we explore the cyber–physical–social intelligence for Edu-Metaverse in cyber–physical–social space and specially design a lecturer-centered immersive teaching system, taking the social and lecturers’ factors into consideration. We call this system VirtualClassroom (V-Classroom). Specifically, we first introduce the cyber–physical–social system (CPSS) paradigm of V-Classroom so that the workflow is standardized and significantly simplified, and the systems can be constructed with off-the-shelf hardware. The key component of V-Classroom is a cyber-world representation of a physical-world classroom instrumented with sparse consumer-grade RGBD cameras for capturing the 3-D geometry and texture of the classrooms. We provide each V-Classroom lecturer with a physical device for sending 6DoF view-change messages and showing view-dependent content of the remote classroom. Following the above paradigm, we develop the V-Classroom algorithms, including V-Classroom depth algorithm (V-DA) and V-Classroom view algorithm (V-VA), to achieve the real-time rendering of remote classrooms. V-DA is dedicated to recovering accurate depth information of the classrooms while V-VA is devoted to real-time novel view synthesis. Finally, we illustrate our implemented CPSS-driven V-Classroom prototype, based on real-world classroom scenarios we collected, and discuss the main challenges and future direction. Tianyu Shen, Shi-Sheng Huang, Deqi Li, Fei-Yue Wang 0001, Hua Huang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2022 | Quantization-aware Deep Optics for Diffractive Snapshot Hyperspectral ImagingabstractDiffractive snapshot hyperspectral imaging based on the deep optics framework has been striving to capture the spectral images of dynamic scenes. However, existing deep optics frameworks all suffer from the mismatch between the optical hardware and the reconstruction algorithm due to the quantization operation in the diffractive optical element (DOE) fabrication, leading to the limited performance of hyperspectral imaging in practice. In this paper, we propose the quantization-aware deep optics for diffractive snapshot hyperspectral imaging. Our key observation is that common lithography techniques used in fabricating DOEs need to quantize the DOE height map to a few levels, and can freely set the height for each level. Therefore, we propose to integrate the quantization operation into the DOE height map optimization and design an adaptive mechanism to adjust the physical height of each quantization level. According to the optimization, we fabricate the quantized DOE directly and build a diffractive hyperspectral snapshot imaging system. Our method develops the deep optics framework to be more practical through the awareness of and adaptation to the quantization operation of the DOE physical structure, making the fabricated DOE and the reconstruction algorithm match each other systematically. Extensive synthetic simulation and real hardware experiments validate the superior performance of our method. Lingen Li, Lizhi Wang 0001, Lei Zhang 0021, Zhiwei Xiong, Hua Huang 0001 |
CVPR | 6 |
| 2022 | Learnability Enhancement for Low-light Raw Denoising: Where Paired Real Data Meets Noise ModelingabstractLow-light raw denoising is an important and valuable task in computational photography where learning-based methods trained with paired real data are mainstream. However, the limited data volume and complicated noise distribution have constituted a learnability bottleneck for paired real data, which limits the denoising performance of learning-based methods. To address this issue, we present a learnability enhancement strategy to reform paired real data according to noise modeling. Our strategy consists of two efficient techniques: shot noise augmentation (SNA) and dark shading correction (DSC). Through noise model decoupling, SNA improves the precision of data mapping by increasing the data volume and DSC reduces the complexity of data mapping by reducing the noise complexity. Extensive results on the public datasets and real imaging scenarios collectively demonstrate the state-of-the-art performance of our method. Hansen Feng, Lizhi Wang 0001, Yuzhi Wang, Hua Huang 0001 |
ACM Multimedia | 4 |
| 2022 | Progressive Unsupervised Learning of Local DescriptorsabstractTraining tuple construction is a crucial step in unsupervised local descriptor learning. Existing approaches perform this step relying on heuristics, which suffer from inaccurate supervision signals and struggle to achieve the desired performance. To address the problem, this work presents DescPro, an unsupervised approach that progressively explores both accurate and informative training tuples for model optimization without using heuristics. Specifically, DescPro consists of a Robust Cluster Assignment (RCA) method to infer pairwise relationships by clustering reliable samples with the increasingly powerful CNN model, and a Similarity-weighted Positive Sampling (SPS) strategy to select informative positive pairs for training tuple construction. Extensive experimental results show that, with the collaboration of the above two modules, DescPro can outperform state-of-the-art unsupervised local descriptors and even rival competitive supervised ones on standard benchmarks. Wufan Wang, Lei Zhang 0021, Hua Huang 0001 |
ACM Multimedia | 3 |
| 2022 | TFPnP: Tuning-free Plug-and-Play Proximal Algorithms with Applications to Inverse Imaging ProblemsabstractPlug-and-Play (PnP) is a non-convex optimization framework that combines proximal algorithms, for example, the alternating direction method of multipliers (ADMM), with advanced denoising priors. Over the past few years, great empirical success has been obtained by PnP algorithms, especially for the ones that integrate deep learning-based denoisers. However, a key problem of PnP approaches is the need for manual parameter tweaking which is essential to obtain high-quality results across the high discrepancy in imaging conditions and varying scene content. In this work, we present a class of tuning-free PnP proximal algorithms that can determine parameters such as denoising strength, termination time, and other optimization-specific parameters automatically. A core part of our approach is a policy network for automated parameter search which can be effectively learned via a mixture of model-free and model-based deep reinforcement learning strategies. We demonstrate, through rigorous numerical and visual experiments, that the learned policy can customize parameters to different settings, and is often more efficient and effective than existing handcrafted criteria. Moreover, we discuss several practical considerations of PnP denoisers, which together with our learned policy yield state-of-the-art results. This advanced performance is prevalent on both linear and nonlinear exemplar inverse imaging problems, and in particular shows promising results on compressed sensing MRI, sparse-view CT, single-photon imaging, and phase retrieval. Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Hua Huang 0001, Carola-Bibiane Schönlieb |
J. Mach. Learn. Res. | 5 |
| 2022 | Sensitivity-Aware Spatial Quality Adaptation for Live Video AnalyticsabstractTo address the conflict between the limited network bandwidth and high DNN inference accuracy, live video analytics desires a bandwidth-efficient streaming approach. To this end, more and more works study spatially variable quality streaming where high quality is only used for important regions. The key challenges are to accurately identify the important regions and select the right qualities for them to maximize accuracy. Existing approaches use either cheap analytics models or low-quality videos to locate important regions, and employ heuristic rules to make quality decisions, which struggle to address the above challenges. Our key insight is that the region’s accuracy “sensitivity” obtained by running the expensive DNN model on the high-quality video provides a reliable indication of the region’s importance and allows to allocate the available bandwidth optimally over regions by explicitly maximizing the frame accuracy. This work presents a sensitivity-aware algorithm Orchestra, which incorporates sensitivity into the design of spatial quality adaptation, including video zoning and quality selection. The design of Orchestra entails three main contributions: a feasible way of sensitivity estimation, sensitivity-aware zoning, and deduction-based accuracy estimation. Extensive experiments over realistic videos and network traces show that Orchestra improves accuracy by upto 14.1% with comparable bandwidth usage or reduces bandwidth usage by upto 44.2% while maintaining higher accuracy compared to baselines. Wufan Wang, Lei Zhang 0021, Hua Huang 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | Coded Hyperspectral Image Reconstruction Using Deep External and Internal LearningabstractTo solve the low spatial and/or temporal resolution problem which the conventional hyperspectral cameras often suffer from, coded hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Specifically, we first develop a CNN-based channel attention reconstruction network to effectively exploit the spatial-spectral correlation of the HSI. Then, the reconstruction network is learned by leveraging an arbitrary external hyperspectral dataset to exploit the general spatial-spectral correlation under adversarial loss. Finally, we customize the network by internal learning with spatial-spectral constraint and total variation regularization for each coded image, which can make use of the internal imaging model to learn specific prior for current desirable image and effectively avoids overfitting. Experimental results using both synthetic data and real images show that our method outperforms the state-of-the-art methods on several popular coded hyperspectral imaging systems under both comprehensive quantitative metrics and perceptive quality. Ying Fu 0001, Tao Zhang 0042, Lizhi Wang 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Joint Camera Spectral Response Selection and Hyperspectral Image RecoveryabstractHyperspectral image (HSI) recovery from a single RGB image has attracted much attention, whose performance has recently been shown to be sensitive to the camera spectral response (CSR). In this paper, we present an efficient convolutional neural network (CNN) based method, which can jointly select the optimal CSR from a candidate dataset and learn a mapping to recover HSI from a single RGB image captured with this algorithmically selected camera under multi-chip or single-chip setups. Given a specific CSR, we first present a HSI recovery network, which accounts for the underlying characteristics of the HSI, including spectral nonlinear mapping and spatial similarity. Later, we append a CSR selection layer onto the recovery network, and the optimal CSR under both multi-chip and single-chip setups can thus be automatically determined from the network weights under the nonnegative sparse constraint. Experimental results on three hyperspectral datasets and two camera spectral response datasets demonstrate that our HSI recovery network outperforms state-of-the-art methods in terms of both quantitative metrics and perceptive quality, and the selection layer always returns a CSR consistent to the best one determined by exhaustive search. Finally, we show that our method can also perform well in the real capture system, and collect a hyperspectral flower dataset to evaluate the effect from HSI recovery on classification problem. Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Novel hyperbolic clustering-based band hierarchy (HCBH) for effective unsupervised band selection of hyperspectral images
He Sun 0009, Lei Zhang 0054, Jinchang Ren, Hua Huang 0001 |
Pattern Recognit. | 4 |
| 2022 | Stochastic gate-based autoencoder for unsupervised hyperspectral band selection
He Sun 0009, Lei Zhang 0021, Lizhi Wang 0001, Hua Huang 0001 |
Pattern Recognit. | 4 |
| 2022 | Robust Extraction and Super-Resolution of Low-Resolution Flying Airplane From Satellite VideoabstractExtracting the flying airplane from the satellite video and enhancing its resolution are significant and demanding tasks in the remote sensing community. The challenge mainly lies in that the flying airplane target in the satellite video often suffers from detail loss due to complex background and limited spaceborne imaging device. In this article, a novel constructive model is proposed to model the airplane of low resolution for more complete extraction, and a new reflective symmetry shape prior is integrated into the super-resolution process to obtain the higher resolution result. Concretely, each frame can be decomposed as a linear combination of foreground and background with specific mixture ratios. With the assumption of uniform linear motion and the rigidity of the airplane, a periodic change of mixture ratios through frames is induced, which can construct the airplane as complete as possible by adopting the proposed iterative matting optimization. To further enhance the resolution of the extracted airplane, an improved alternating direction method of multipliers (ADMM) is utilized to solve the super-resolution problem with the reflective symmetry of the shape as prior. The effectiveness of our method with respect to extraction and super-resolution is borne out by the experiments on both synthetic and real data. De-Lei Chen, Lei Zhang 0021, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Extracting Small Flying Airplane With Spatially Accurate and Temporally Consistent Foreground Modeling
De-Lei Chen, Lei Zhang 0021, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | IMU-Assisted Online Video Background IdentificationabstractDistinguishing between dynamic foreground objects and a mostly static background is a fundamental problem in many computer vision and computer graphics tasks. This paper presents a novel online video background identification method with the assistance of inertial measurement unit (IMU). Based on the fact that the background motion of a video essentially reflects the 3D camera motion, we leverage IMU data to realize a robust camera motion estimation for identifying background feature points by only investigating a few historical frames. We observe that the displacement of the 2D projection of a scene point caused by camera rotation is depth-invariant, and the rotation estimation by using IMU data can be quite accurate. We thus propose to analyze 2D feature points by decomposing the 2D motion into two components: rotation projection and translation projection. In our method, after establishing the 3D camera rotations, we generate the depth-relevant 2D feature point movement induced by the camera 3D translation. Then, by examining the disparity between inter-frame offset and the projection of estimated 3D camera motion, we can identify the background feature points. In the experiments, our online method is able to run at 30FPS with only 1 frame latency and outperforms state-of-the-art background identification and other relevant methods. Our method directly leads to a better camera motion estimation, which is beneficial to many applications like online video stabilization, SLAM, image stitching, etc. Jianxiang Rong, Lei Zhang 0021, Hua Huang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Few-Shot Class-Incremental Learning via Compact and Separable Features for Fine-Grained Vehicle RecognitionabstractMost of the existing deep learning-based fine-grained vehicle recognition methods collect a large-scale training set in advance and train a model based on the closed-world assumption. However, in the real world, new classes of vehicles are released over time but it is difficult to collect sufficient labeled data for new classes, which results in a typical few-shot class-incremental learning problem (FSCIL). To solve this problem, this work proposes a compact and separable feature learning method (CSFL) which exploits a decoupled learning scheme to prevent the feature extractor from updating during class-incremental learning. CSFL trains an initial model to learn discriminative features of fine-grained vehicles using deep metric learning. Then an incremental linear discriminant analysis algorithm is applied to the learned features to further discriminate potential confused classes. Specifically, the decoupled components share the same objective of enhancing intra-class compactness and inter-class separability, which is beneficial for classification. Extensive experiments on three fine-grained vehicle datasets demonstrate that the proposed CSFL achieves better results than state-of-the-art incremental learning methods, validating the importance of compact and separable features in the problem of FSCIL. De-Wang Li, Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | A Large-Scale Frontal Vehicle Image Dataset for Fine-Grained Vehicle CategorizationabstractFine-grained vehicle categorization has evolved into a significant subject of study due to its importance in the Intelligent Transportation System. A highly accurate and real-time vehicle categorization system will help to support many applications not only in the security aspect but also many walks of life. In this paper, facing the growing importance of this study, we present an image dataset named Frontal-103 to promote the development of the vision-based research on the vehicle, and particularly for the task of fine-grained vehicle categorization. This paper provides a detailed analysis of Frontal-103 in its current state: 1,759 fine-grained vehicle models in 103 vehicle makes and 65,433 web-nature images in total. Apart from the specific viewpoint and vehicle hierarchy, Frontal-103 is superior to the other state-of-the-art vehicle image datasets not only in the scale and diversity but also the accuracy and fine-grained level. We further discuss the peculiar challenges and issues lies in the task of fine-grained vehicle categorization and illustrate the usefulness of our dataset in addressing those problems. We hope Frontal-103 will be beneficial to the vision-based vehicle analysis and contribute to the computer vision community. Ping Wang 0009, Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Towards Universal Physical Attacks on Single Object TrackingabstractRecent studies show that small perturbations in video frames could misguide single object trackers. However, such attacks have been mainly designed for digital-domain videos (i.e., perturbation on full images), which makes them practically infeasible to evaluate the adversarial vulnerability of trackers in real-world scenarios. Here we made the first step towards physically feasible adversarial attacks against visual tracking in real scenes with a universal patch to camouflage single object trackers. Fundamentally different from physical object detection, the essence of single object tracking lies in the feature matching between the search image and templates, and we therefore specially design the maximum textural discrepancy (MTD), a resolution-invariant and target location-independent feature de-matching loss. The MTD distills global textural information of the template and search images at hierarchical feature scales prior to performing feature attacks. Moreover, we evaluate two shape attacks, the regression dilation and shrinking, to generate stronger and more controllable attacks. Further, we employ a set of transformations to simulate diverse visual tracking scenes in the wild. Experimental results show the effectiveness of the physically feasible attacks on SiamMask and SiamRPN++ visual trackers both in digital and physical scenes. Kaiwen Yuan, Minyang Jiang, Ping Wang 0009, Hua Huang 0001, Z. Jane Wang 0001 |
AAAI | 6 |
| 2021 | Learning Tensor Low-Rank Prior for Hyperspectral Image ReconstructionabstractSnapshot hyperspectral imaging has been developed to capture the spectral information of dynamic scenes. In this paper, we propose a deep neural network by learning the tensor low-rank prior of hyperspectral images (HSI) in the feature domain to promote the reconstruction quality. Our method is inspired by the canonical-polyadic (CP) decomposition theory, where a low-rank tensor can be expressed as a weight summation of several rank-1 component tensors. Specifically, we first learn the tensor low-rank prior of the image features with two steps: (a) we generate rank-1 tensors with discriminative components to collect the contextual information from both spatial and channel dimensions of the image features; (b) we aggregate those rank-1 tensors into a low-rank tensor as a 3D attention map to exploit the global correlation and refine the image features. Then, we integrate the learned tensor low-rank prior into an iterative optimization algorithm to obtain an end-to-end HSI reconstruction. Experiments on both synthetic and real data demonstrate the superiority of our method. Lizhi Wang 0001, Lei Zhang 0021, Hua Huang 0001 |
CVPR | 4 |
| 2021 | Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense FrameworkabstractDeep learning based image classification models are shown vulnerable to adversarial attacks by injecting deliberately crafted noises to clean images. To defend against adversarial attacks in a training-free and attack-agnostic manner, this work proposes a novel and effective reconstruction-based defense framework by delving into deep image prior (DIP). Fundamentally different from existing reconstruction-based defenses, the proposed method analyzes and explicitly incorporates the model decision process into our defense. Given an adversarial image, firstly we map its reconstructed images during DIP optimization to the model decision space, where cross-boundary images can be detected and on-boundary images can be further localized. Then, adversarial noise is purified by perturbing on-boundary images along the reverse direction to the adversarial image. Finally, on-manifold images are stitched to construct an image that can be correctly predicted by the victim classifier. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art reconstruction-based methods both in defending white-box attacks and defense-aware attacks. Moreover, the proposed method can maintain a high visual quality during adversarial image reconstruction. Xin Ding 0004, Kaiwen Yuan, Ping Wang 0009, Hua Huang 0001, Z. Jane Wang 0001 |
ACM Multimedia | 6 |
| 2021 | Adaptive Dimension-Discriminative Low-Rank Tensor Recovery for Computational Hyperspectral Imaging
Lizhi Wang 0001, Hua Huang 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | Sparse additive discriminant canonical correlation analysis for multiple features fusion
Zhan Wang 0007, Lizhi Wang 0001, Hua Huang 0001 |
Neurocomputing | 3 |
| 2021 | Unified quality assessment of natural and screen content images via adaptive weighting on double scales
Ping Wang 0009, Hua Huang 0001 |
Signal Process. Image Commun. | 3 |
| 2021 | Loop Closure Detection by Using Global and Local Features With Photometric and Viewpoint InvarianceabstractLoop closure detection plays an important role in many Simultaneous Localization and Mapping (SLAM) systems, while the main challenge lies in the photometric and viewpoint variance. This paper presents a novel loop closure detection algorithm that is more robust to the variance by using both global and local features. Specifically, the global feature with the consolidation of photometric and viewpoint invariance is learned by a Siamese Network from the intensity, depth, gradient and normal vectors distribution. The local feature with rotation invariance is based on the histogram of relative pixel intensity and geometric information like curvature and coplanarity. Then, these two types of features are jointly leveraged for the robust detection of loop closures. The extensive experiments have been conducted on the publicly available RGB-D benchmark datasets like TUM and KITTI. The results demonstrate that our algorithm can effectively address challenging scenarios with large photometric and viewpoint variance, which outperforms other state-of-the-art methods. Mingfei Yu, Lei Zhang 0021, Wufan Wang, Hua Huang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Shared Low-Rank Correlation Embedding for Multiple Feature FusionabstractThe diversity of multimedia data in the real world usually forms heterogeneous types of feature sets. How to explore the structure information and the relationships among multiple features is still an open problem. In this paper, we propose an unsupervised subspace learning method, named the shared low-rank correlation embedding (SLRCE) for multiple feature fusion. First, in the learned subspace, we implement the low-rank representation on each feature set and enforce a shared low-rank constraint to uncover the common structure information of multiple features. Second, we develop an enhanced correlation analysis in the learned subspace for simultaneously removing the redundancy of each feature set and exploring the correlation of multiple features. Finally, we incorporate the shared low-rank representation and the correlation analysis into a unified framework. The shared low-rank constraint not only depicts the data distribution consistency among multiple features, but also assists robust subspace learning. Our method is robust to noise in practice and can be extended to the kernel case to handle the nonlinear feature fusion. Experimental results on several typical datasets demonstrate the superior performance of the proposed methods. Zhan Wang 0007, Lizhi Wang 0001, Jun Wan 0001, Hua Huang 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | 3-D Quasi-Recurrent Neural Network for Hyperspectral Image DenoisingabstractIn this article, we propose an alternating directional 3-D quasi-recurrent neural network for hyperspectral image (HSI) denoising, which can effectively embed the domain knowledge-structural spatiospectral correlation and global correlation along spectrum (GCS). Specifically, 3-D convolution is utilized to extract structural spatiospectral correlation in an HSI, while a quasi-recurrent pooling function is employed to capture the GCS. Moreover, the alternating directional structure is introduced to eliminate the causal dependence with no additional computation cost. The proposed model is capable of modeling spatiospectral dependence while preserving the flexibility toward HSIs with an arbitrary number of bands. Extensive experiments on HSI denoising demonstrate significant improvement over the state-of-the-art under various noise settings, in terms of both restoration accuracy and computation time. Our code is available at https://github.com/Vandermode/QRNN3D. Kaixuan Wei, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | DNU: Deep Non-Local Unrolling for Computational Spectral ImagingabstractComputational spectral imaging has been striving to capture the spectral information of the dynamic world in the last few decades. In this paper, we propose an interpretable neural network for computational spectral imaging. First, we introduce a novel data-driven prior that can adaptively exploit both the local and non-local correlations among the spectral image. Our data-driven prior is integrated as a regularizer into the reconstruction problem. Then, we propose to unroll the reconstruction problem into an optimization-inspired deep neural network. The architecture of the network has high interpretability by explicitly characterizing the image correlation and the system imaging model. Finally, we learn the complete parameters in the network through end-to-end training, enabling robust performance with high spatial-spectral fidelity. Extensive simulation and hardware experiments validate the superior performance of our method over state-of-the-art methods. Lizhi Wang 0001, Maoqing Zhang, Ying Fu 0001, Hua Huang 0001 |
CVPR | 5 |
| 2020 | A Physics-Based Noise Formation Model for Extreme Low-Light Raw DenoisingabstractLacking rich and realistic data, learned single image denoising algorithms generalize poorly in real raw images that not resemble the data used for training. Although the problem can be alleviated by the heteroscedastic Gaussian noise model, the noise sources caused by digital camera electronics are still largely overlooked, despite their significant effect on raw measurement, especially under extremely low-light condition. To address this issue, we present a highly accurate noise formation model based on the characteristics of CMOS photosensors, thereby enabling us to synthesize realistic samples that better match the physics of image formation process. Given the proposed noise model, we additionally propose a method to calibrate the noise parameters for available modern digital cameras, which is simple and reproducible for any new device. We systematically study the generalizability of a neural network trained with existing schemes, by introducing a new low-light denoising dataset that covers many modern digital cameras from diverse brands. Extensive empirical results collectively show that by utilizing our proposed noise formation model, a network can reach the capability as if it had been trained with rich real data, which demonstrates the effectiveness of our noise formation model. Kaixuan Wei, Ying Fu 0001, Jiaolong Yang, Hua Huang 0001 |
CVPR | 4 |
| 2020 | Structure Preserving Multi-View Dimensionality ReductionabstractThe multi-view features from multimedia data in the real-world are usually high-dimensional. How to simultaneously reduce their dimensions and explore the complementary information among multi-view features is of vital importance but challenging. In this paper, we propose a novel unsupervised method named structure preserving multi-view dimensionality reduction (SPMDR). We first propose a bilinear low-rank representation with an orthogonal constraint in the learning subspace. Then, we construct a 3-order rotated tensor among the low-rank coefficient matrices and utilize tensor nuclear norm to capture complementary information among multi-view representations. Finally, we develop a numerical algorithm for solving the proposed model. Our method is robust to noisy data and can capture the complex correlations among multi-view features. Experimental results on recognition tasks demonstrate the superior performance of SPMDR. Zhan Wang 0007, Lizhi Wang 0001, Hua Huang 0001 |
ICME | 3 |
| 2020 | Tuning-free Plug-and-Play Proximal Algorithm for Inverse Imaging ProblemsabstractPlug-and-play (PnP) is a non-convex framework that combines ADMM or other proximal algorithms with advanced denoiser priors. Recently, PnP has achieved great empirical success, especially with the integration of deep learning-based denoisers. However, a key problem of PnP based approaches is that they require manual parameter tweaking. It is necessary to obtain high-quality results across the high discrepancy in terms of imaging conditions and varying scene content. In this work, we present a tuning-free PnP proximal algorithm, which can automatically determine the internal parameters including the penalty parameter, the denoising strength and the terminal time. A key part of our approach is to develop a policy network for automatic search of parameters, which can be effectively learned via mixed model-free and model-based deep reinforcement learning. We demonstrate, through numerical and visual experiments, that the learned policy can customize different parameters for different states, and often more efficient and effective than existing handcrafted criteria. Moreover, we discuss the practical considerations of the plugged denoisers, which together with our learned policy yield state-of-the-art results. This is prevalent on both linear and nonlinear exemplary inverse imaging problems, and in particular, we show promising results on Compressed Sensing MRI and phase retrieval. Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Carola-Bibiane Schönlieb, Hua Huang 0001 |
ICML | 6 |
| 2020 | Snapshot Hyperspectral Imaging Based on Weighted High-order Singular Value RegularizationabstractSnapshot hyperspectral imaging can capture the 3D hyperspectral image (HSI) with a single 2D measurement and has attracted increasing attention recently. Recovering the underlying HSI from the compressive measurement is an ill-posed problem and exploiting the image prior is essential for solving this ill-posed problem. However, existing reconstruction methods always start from modeling image prior with the 1D vector or 2D matrix and cannot fully exploit the structurally spectral-spatial nature in 3D HSI, thus leading to a poor fidelity. In this paper, we propose an effective high-order tensor optimization based method to boost the reconstruction fidelity for snapshot hyperspectral imaging. We first build high-order tensors by exploiting the spatial-spectral correlation in HSI. Then, we propose a weight high-order singular value regularization (WHOSVR) based low-rank tensor recovery model to characterize the structure prior of HSI. By integrating the structure prior in WHOSVR with the system imaging process, we develop an optimization framework for HSI reconstruction, which is finally solved via the alternating minimization algorithm. Extensive experiments implemented on two representative systems demonstrate that our method outperforms state-of-the-art methods. Niankai Cheng, Hua Huang 0001, Lei Zhang 0021, Lizhi Wang 0001 |
ICPR | 2 |
| 2020 | Embedding shared low-rank and feature correlation for multi-view data analysisabstractThe diversity of multimedia data in the real-world usually forms multi-view features. How to explore the structure information and correlations among multi-view features is still a challenging problem. In this paper, we propose a novel multi-view subspace learning method, named embedding shared low-rank and feature correlation (ESLRFC), for multi-view data analysis. First, in the embedding subspace, we propose a robust low-rank model on each feature set and enforce a shared low-rank constraint to characterize the common structure information of multiple feature data. Second, we develop an enhanced correlation analysis in the embedding subspace for simultaneously removing the redundancy of each feature set and exploring the correlations of multiple feature data. Finally, we incorporate the low-rank model and the correlation analysis into a unified framework. The shared low-rank constraint not only depicts the data distribution consistency among multiple feature data, but also assists robust subspace learning. Experimental results on recognition tasks demonstrate the superior performance and noise robustness of the proposed method. Zhan Wang 0007, Lizhi Wang 0001, Lei Zhang 0021, Hua Huang 0001 |
ICPR | 4 |
| 2020 | Simultaneous hyperspectral image super-resolution and geometric alignment with a hybrid camera system
Ying Fu 0001, Yongrong Zheng, Yinqiang Zheng, Hua Huang 0001 |
Neurocomputing | 5 |
| 2020 | Component-based feature extraction and representation schemes for vehicle make and model recognition
Hua Huang 0001 |
Neurocomputing | 2 |
| 2020 | Unified non-uniform scale adaptive sampling model for quality assessment of natural scene and screen content images
Tianli Tao, Hua Huang 0001 |
Neurocomputing | 3 |
| 2020 | Joint low rank embedded multiple features learning for audio-visual emotion recognition
Zhan Wang 0007, Lizhi Wang 0001, Hua Huang 0001 |
Neurocomputing | 3 |
| 2020 | Blind image blur metric based on orientation-aware local patterns
Lixiong Liu, Jiachao Gong, Hua Huang 0001, Qingbing Sang |
Signal Process. Image Commun. | 3 |
| 2020 | Video quality assessment using space-time slice mappings
Lixiong Liu, Tianshu Wang 0003, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2020 | Blind S3D image quality prediction using classical and non-classical receptive field models
Lixiong Liu, Jiufa Zhang, Michele A. Saad, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2020 | Small Target Detection in Infrared Videos Based on Spatio-Temporal Tensor ModelabstractExisting methods of the small target detection from infrared videos are not effective with the complex background. It is mainly caused by: 1) the interference of strong edges and the similarity with other nontarget objects and 2) the lack of the context information of both the background and the target in a spatio-temporal domain. By considering these two points, we propose to slide a window in a single frame and form a spatio-temporal cube with the current frame patch and other frame patches in the spatio-temporal domain. Then, we establish a spatio-temporal tensor model based on these patches. According to the sparse prior of the target and the local correlation of the background, the separation of the target and the background can be cast as a low rank and sparse tensor decomposition problem. The target is obtained from the sparse tensor by the tensor decomposition. The experiments show that our method gains better detection performance in infrared videos with the complex background by making full use of the spatio-temporal context information. Hong-Kang Liu, Lei Zhang 0021, Hua Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Global Topology Constraint Network for Fine-Grained Vehicle RecognitionabstractFine-grained vehicle recognition is challenging due to the large intra-class variation in vehicle pose and viewpoint. Many existing methods, especially convolution neural network (CNN)-based methods, solve this problem via detecting and aligning parts individually, and do not consider the interaction between parts, which is very important for effective part detection and vehicle recognition. In this paper, we propose a global topology constraint network for fine-grained vehicle recognition, which adopts the constraint of global topology relationship to depict the interaction between parts and integrates it into CNN in an efficient way. Different CNNs for image classification truncated at intermediate layer can be used for part detection. The global topology relationship between parts is encoded into kernel of depthwise convolution layer, and can be learned from training images. The response under global topology constraint reflects the probability of topology relationship between parts existing in input image. All responses build the description for classification with translation invariance. Through training the whole network, the back-propagation of gradient information of kernel for global topology relationship will guide former layers to better detect useful parts, and thereby improve vehicle recognition. Our proposed method does not require additional annotation, such as bounding box or part annotation, and can be trained in an end-to-end way. We conduct comparison experiments on public Stanford Cars and CompCars datasets, which both show that our method achieves the state-of-the-art performance. Ye Xiang, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Hyperspectral Image Super-Resolution With Optimized RGB GuidanceabstractTo overcome the limitations of existing hyperspectral cameras on spatial/temporal resolution, fusing a low resolution hyperspectral image (HSI) with a high resolution RGB (or multispectral) image into a high resolution HSI has been prevalent. Previous methods for this fusion task usually employ hand-crafted priors to model the underlying structure of the latent high resolution HSI, and the effect of the camera spectral response (CSR) of the RGB camera on super-resolution accuracy has rarely been investigated. In this paper, we first present a simple and efficient convolutional neural network (CNN) based method for HSI super-resolution in an unsupervised way, without any prior training. Later, we append a CSR optimization layer onto the HSI super-resolution network, either to automatically select the best CSR in a given CSR dataset, or to design the optimal CSR under some physical restrictions. Experimental results show our method outperforms the state-of-the-arts, and the CSR optimization can further boost the accuracy of HSI super-resolution. Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
CVPR | 5 |
| 2019 | Hyperspectral Image Reconstruction Using a Deep Spatial-Spectral PriorabstractRegularization is a fundamental technique to solve an ill-posed optimization problem robustly and is essential to reconstruct compressive hyperspectral images. Various hand-crafted priors have been employed as a regularizer but are often insufficient to handle the wide variety of spectra of natural hyperspectral images, resulting in poor reconstruction quality. Moreover, the prior-regularized optimization requires manual tweaking of its weight parameters to achieve a balance between the spatial and spectral fidelity of result images. In this paper, we present a novel hyperspectral image reconstruction algorithm that substitutes the traditional hand-crafted prior with a data-driven prior, based on an optimization-inspired network. Our method consists of two main parts: First, we learn a novel data-driven prior that regularizes the optimization problem with a goal to boost the spatial-spectral fidelity. Our data-driven prior learns both local coherence and dynamic characteristics of natural hyperspectral images. Second, we combine our regularizer with an optimization-inspired network to overcome the heavy computation problem in the traditional iterative optimization methods. We learn the complete parameters in the network through end-to-end training, enabling robust performance with high accuracy. Extensive simulation and hardware experiments validate the superior performance of our method over the state-of-the-art methods. Lizhi Wang 0001, Ying Fu 0001, Min H. Kim 0001, Hua Huang 0001 |
CVPR | 5 |
| 2019 | Single Image Reflection Removal Exploiting Misaligned Training Data and Network EnhancementsabstractRemoving undesirable reflections from a single image captured through a glass window is of practical importance to visual computing systems. Although state-of-the-art methods can obtain decent results in certain situations, performance declines significantly when tackling more general real-world cases. These failures stem from the intrinsic difficulty of single image reflection removal -- the fundamental ill-posedness of the problem, and the insufficiency of densely-labeled training data needed for resolving this ambiguity within learning-based neural network pipelines. In this paper, we address these issues by exploiting targeted network enhancements and the novel use of misaligned data. For the former, we augment a baseline network architecture by embedding context encoding modules that are capable of leveraging high-level contextual clues to reduce indeterminacy within areas containing strong reflections. For the latter, we introduce an alignment-invariant loss function that facilitates exploiting misaligned real-world training data that is much easier to collect. Experimental results collectively show that our method outperforms the state-of-the-art with aligned data, and that significant improvements are possible when using additional misaligned data. Kaixuan Wei, Jiaolong Yang, Ying Fu 0001, David P. Wipf, Hua Huang 0001 |
CVPR | 5 |
| 2019 | Incremental Learning Using Conditional Adversarial NetworksabstractIncremental learning using Deep Neural Networks (DNNs) suffers from catastrophic forgetting. Existing methods mitigate it by either storing old image examples or only updating a few fully connected layers of DNNs, which, however, requires large memory footprints or hurts the plasticity of models. In this paper, we propose a new incremental learning strategy based on conditional adversarial networks. Our new strategy allows us to use memory-efficient statistical information to store old knowledge, and fine-tune both convolutional layers and fully connected layers to consolidate new knowledge. Specifically, we propose a model consisting of three parts, i.e., a base sub-net, a generator, and a discriminator. The base sub-net works as a feature extractor which can be pre-trained on large scale datasets and shared across multiple image recognition tasks. The generator conditioned on labeled embeddings aims to construct pseudo-examples with the same distribution as the old data. The discriminator combines real-examples from new data and pseudo-examples generated from the old data distribution to learn representation for both old and new classes. Through adversarial training of the discriminator and generator, we accomplish the multiple continuous incremental learning. Comparison with the state-of-the-arts on public CIFAR-100 and CUB-200 datasets shows that our method achieves the best accuracies on both old and new classes while requiring relatively less memory storage. Ye Xiang, Ying Fu 0001, Pan Ji, Hua Huang 0001 |
ICCV | 4 |
| 2019 | Hyperspectral Image Reconstruction Using Deep External and Internal LearningabstractTo solve the low spatial and/or temporal resolution problem which the conventional hypelrspectral cameras often suffer from, coded snapshot hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Our method can effectively exploit spatial-spectral correlation and sufficiently represent the variety nature of HSIs. Experimental results show our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality. Tao Zhang 0042, Ying Fu 0001, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 4 |
| 2019 | Computational Hyperspectral Imaging Based on Dimension-Discriminative Low-Rank Tensor RecoveryabstractExploiting the prior information is fundamental for the image reconstruction in computational hyperspectral imaging. Existing methods usually unfold the 3D signal as a 1D vector and treat the prior information within different dimensions in an indiscriminative manner, which ignores the high-dimensionality nature of hyperspectral image (HSI) and thus results in poor quality reconstruction. In this paper, we propose to make full use of the high-dimensionality structure of the desired HSI to boost the reconstruction quality. We first build a high-order tensor by exploiting the nonlocal similarity in HSI. Then, we propose a dimension-discriminative low-rank tensor recovery (DLTR) model to characterize the structure prior adaptively in each dimension. By integrating the structure prior in DLTR with the system imaging process, we develop an optimization framework for HSI reconstruction, which is finally solved via the alternating minimization algorithm. Extensive experiments implemented with both synthetic and real data demonstrate that our method outperforms state-of-the-art methods. Lizhi Wang 0001, Ying Fu 0001, Xiaoming Zhong, Hua Huang 0001 |
ICCV | 5 |
| 2019 | No-Reference Stereoscopic Video Quality Assessment Based on Spatial-Temporal Statistics
Jiufa Zhang, Lixiong Liu, Jiachao Gong, Hua Huang 0001 |
ICIG (3) | 4 |
| 2019 | An Effective Network with ConvLSTM for Low-Light Image Enhancement
Yixi Xiang, Ying Fu 0001, Lei Zhang 0021, Hua Huang 0001 |
PRCV (2) | 4 |
| 2019 | Fast HSI super resolution using linear regressionabstractHyperspectral imaging has great achievements in agriculture, astronomy, surveillance, and so on. However, the inherent low spatial resolution of hyperspectral imaging, unfortunately, limits its more widespread applications. Recently, hyperspectral image (HSI) super resolution addresses this problem by fusing a low spatial resolution HSI (LR‐HSI) with a high spatial resolution multispectral image (HR‐MSI), but most of these methods did not consider real‐time restoration of high spatial resolution HSI. In this study, the authors propose a fast HSI super‐resolution method which fills this blank. Specifically, they model the hyperspectral super resolution as a linear regression problem according to the fact that the imaging process is a linear transform and the inverse of this transform can be approximately estimated, as the spectra of a typical scene lie in a very low‐dimensional space. To further exploit the low‐dimensional nature of the spectra, they divide the HR‐MSI and LR‐HSI into several patches and learn the inverse transform patch‐by‐patch. Experiments on several public datasets show that their method approximates state‐of‐the‐art methods in accuracy, but is several orders of magnitude faster than all of them. Furthermore, they provide an efficient C language implementation of their methods, which can meet the real‐time request. Lingfei Song, Ying Fu 0001, Hua Huang 0001 |
IET Image Process. | 3 |
| 2019 | Image restoration from patch-based compressed sensing measurement
Hua Huang 0001, Guangtao Nie, Yinqiang Zheng, Ying Fu 0001 |
Neurocomputing | 1 |
| 2019 | Global relative position space based pooling for fine-grained vehicle recognition
Ye Xiang, Ying Fu 0001, Hua Huang 0001 |
Neurocomputing | 3 |
| 2019 | Defocus Hyperspectral Image Deblurring with Adaptive Reference Image and Scale Map
De-Wang Li, Linjing Lai, Hua Huang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2019 | High-Speed Hyperspectral Video Acquisition By Combining Nyquist and Compressive SamplingabstractWe propose a novel hybrid imaging system to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. The proposed system consists of two branches: one branch performs Nyquist sampling in the temporal dimension while integrating the whole spectrum, resulting in a high-frame-rate panchromatic video; the other branch performs compressive sampling in the spectral dimension with longer exposures, resulting in a low-frame-rate hyperspectral video. Owing to the high light throughput and complementary sampling, these two branches jointly provide reliable measurements for recovering the underlying HSHS video. Moreover, the panchromatic video can be used to learn an over-complete 3D dictionary to represent each band-wise video sparsely, thanks to the inherent structural similarity in the spectral dimension. Based on the joint measurements and the self-adaptive dictionary, we further propose a simultaneous spectral sparse (3S) model to reinforce the structural similarity across different bands and develop an efficient computational reconstruction algorithm to recover the HSHS video. Both simulation and hardware experiments validate the effectiveness of the proposed approach. To the best of our knowledge, this is the first time that hyperspectral videos can be acquired at a frame rate up to 100fps with commodity optical elements and under ordinary indoor illumination. Lizhi Wang 0001, Zhiwei Xiong, Hua Huang 0001, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Encoding Shaky Videos by Integrating Efficient Video StabilizationabstractThis paper presents a novel video coding method by integrating video stabilization for shaky videos. By reusing the stabilized motion of feature points and geometric transformations, a better predictor can be established to replace the original motion vectors in the motion estimation stage of video coding. Then, these motion vectors are optimized based on statistics of the residuals between the stabilization predictor and the standard one to improve the efficiency of motion search. As a result, our method brings much less computational cost to motion estimation than encoding the stabilized frames separately (e.g., 24% for the enhanced predictive zonal search algorithm, 17% for the unsymmetrical-cross multi-hexagon-grid search algorithm), while the Bjontegaard distortion (BD) bit rate and BDpeak signal-to-noise ratio still have the comparable performance with multiple quantization parameters. Specially, the implementation of our integrative system based on ×264 is very fast and of low latency, by which full HD videos can be simultaneously stabilized and encoded with more than 30 frames per second in the fastest mode, even on the mobile platform. The experiments on a variety of shaky videos demonstrate the potential of our method in terms of effectiveness and efficiency. Hua Huang 0001, Xiao-Xiang Wei, Lei Zhang 0021 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Fast Parallel Implementation of Dual-Camera Compressive Hyperspectral Imaging SystemabstractCoded aperture snapshot spectral imager (CASSI) provides a potential solution to recover the 3D hyperspectral image (HSI) from a single 2D measurement. The latest proposed design of the dual-camera compressive hyperspectral imager (DCCHI) can collect more information simultaneously with the CASSI to improve the reconstruction quality. The main bottleneck now lies in the high computation complexity of the reconstruction methods, which hinders the practical application. In this paper, we propose a fast parallel implementation based on DCCHI to reach a stable and efficient HSI reconstruction. Specifically, we develop a new optimization method for the reconstruction problem, which integrates the alternative direction multiplier method with the total variation-based regularization to boost the convergence rate. Then, to improve the time efficiency, a novel parallel implementation based on GPU is proposed. The performance of the proposed method is validated on both synthetic and real data. The experimental results demonstrate that our method has a significant advantage in time efficiency, while maintaining a comparable reconstruction fidelity. Hua Huang 0001, Ying Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | HyperReconNet: Joint Coded Aperture Optimization and Image Reconstruction for Compressive Hyperspectral ImagingabstractCoded aperture snapshot spectral imaging (CASSI) system encodes the 3D hyperspectral image (HSI) within a single 2D compressive image and then reconstructs the underlying HSI by employing an inverse optimization algorithm, which equips with the distinct advantage of snapshot but usually results in low reconstruction accuracy. To improve the accuracy, existing methods attempt to design either alternative coded apertures or advanced reconstruction methods, but cannot connect these two aspects via a unified framework, which limits the accuracy improvement. In this paper, we propose a convolution neural network (CNN) based endto- end method to boost the accuracy by jointly optimizing the coded aperture and the reconstruction method. On the one hand, based on the nature of CASSI forward model, we design a repeated pattern for the coded aperture, whose entities are learned by acting as the network weights. On the other hand, we conduct the reconstruction through simultaneously exploiting intrinsic properties within HSI - the extensive correlations across the spatial and the spectral dimensions. By leveraging the power of deep learning, the coded aperture design and the image reconstruction are connected and optimized via a unified framework. Experimental results show that our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality. Lizhi Wang 0001, Tao Zhang 0042, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | A Hierarchical Scheme for Vehicle Make and Model Recognition From Frontal Images of VehiclesabstractThis paper presents a novel recognition scheme for vehicle make and model recognition (VMMR) from frontal images of vehicles. In general, we introduce some domain knowledge to cope with this task. The structural components contained in the frontal appearance of vehicles present different visual characteristics and their discriminating ability varies when vehicle models belonging to the same brand or different brands are compared. In light of the particularities, we take advantage of the varying discriminating ability of these structural components to perform the recognition task sequentially in two stages. At the first stage, the logo sub-region (which is one of the component-related sub-regions in the region of interest) is applied to classify the vehicle models at the brand level. Different from the traditional brand-level classification that the models of the same brand are considered as a single class, in this paper, multiple sub-classes in one brand class are allowed, since the intra-brand models also exhibit a certain degree of diversity. In this way, the problem of inter-class similarity is remitted. At the second stage, several customized classifiers are trained for each sub-class in the light of the discriminant ability of the remaining sub-regions. The proposed approach has been tested on a large-scale vehicle image database collected in this paper and has achieved the state-of-the-art results. Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Pre-Attention and Spatial Dependency Driven No-Reference Image Quality AssessmentabstractThe excessive emulation of the human visual system and the lack of connection between chromatic data and distortion have been the major bottlenecks in developing image quality assessment. To address this issue, we develop a new no-reference (NR) image quality assessment (IQA) metric that accounts for the impact of pre-attention and spatial dependency on the perceived quality of distorted images. The resulting model, dubbed the Pre-attention and Spatial-dependency driven Quality Assessment (PSQA) predictor, introduces the pre-attention theory to emulate early phase visual perception by refining luminance-channel data. Chromatic data are also processed concurrently by transforming images from RGB to the perceptually optimized SCIELAB color space. Considering that the gray-tone spatial dependency matrix conveys important texture properties that are closely related to visual quality, this matrix, as a mathematical solution for subsequent visual process emulation, is calculated along with its statistical features on both gray and color channels. To clarify the influence of different regression procedures on model output, support vector regression and AdaBoosting Back Propagation (BP) neural networks are adopted separately to train the prediction models. We thoroughly evaluated PSQA on four public image quality databases: LIVE, TID2013, CSIQ, and VCL. The experimental results show that PSQA delivers highly competitive performance compared with top-rank NR and full-reference IQA metrics. Lixiong Liu, Tianshu Wang 0003, Hua Huang 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Intrinsic Motion Stability Assessment for Video StabilizationabstractThis paper presents a novel algorithm for assessing the motion stability of a video after stabilization. The assessment works in a non-reference manner that directly measures the intrinsic smoothness of the video motion path. Specifically, the motion path is cast as a curve embedded in the Lie group of homographies, and its smoothness is mathematically characterized by the intrinsic geodesic curvature. A bundle of paths are adopted to handle spatially variant motions through the frames. Then, we compute the weighted curvature for a holistic assessment on the motion stability. Other factors related to video stabilization, e.g., distortion and cropping, are also investigated as supplement. We collect 160 shaky video clips and their stabilized results for verification, and the experimental evidence shows the effectiveness of our algorithm in good correlation with human subjective judgements. Lei Zhang 0021, Qing-Zhuo Zheng, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | A Mobile Augmented Reality System For Illumination Consistency With Fisheye Camera Based On Client-Server ModelabstractThis paper proposes a client-server based mobile augmented reality system for illumination consistency between the virtual objects and the real world. The server system adopts a fisheye camera to capture real world illumination and transmits the light source information to the client system, while the client system uses the received information to render the virtual objects in the augmented reality application. The proposed method possesses the virtue of reducing the calculation cost of the clients, which is especially important for mobile devices. An evaluation method of AR image's illumination consistency has also been proposed based on analyzing the distance of the histogram of virtual and real parts of AR image with different channels. Wankui Liu, Yue Liu 0005, Hua Huang 0001 |
CASA | 3 |
| 2018 | Joint Camera Spectral Sensitivity Selection and Hyperspectral Image Recovery
Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
ECCV (3) | 5 |
| 2018 | Evaluation of Hand-Based Interaction for Near-Field Mixed Reality with Optical See-Through Head-Mounted DisplaysabstractHand-based interaction is one of the most widely-used interaction modes in the applications based on optical see-through head-mounted displays (OST-HMDs). In this paper, such interaction modes as gesture-based interaction (GBI) and physics-based interaction (PBI) are developed to construct a mixed reality system to evaluate the advantages and disadvantages of different interaction modes for near-field mixed reality. The experimental results show that PBI leads to a better performance of users regarding their work efficiency in the proposed tasks. The statistical analysis of T-test has been adopted to prove that the difference of efficiency between different interaction modes is significant. Zhenliang Zhang 0002, Benyang Cao, Dongdong Weng, Yue Liu 0005, Yongtian Wang, Hua Huang 0001 |
VR | 6 |
| 2018 | Hyperspectral image super-resolution under misaligned hybrid camera systemabstractHyperspectral imaging has been widely used for agriculture, astronomy, surveillance, and so on. However, hyperspectral imaging usually suffers from low‐spatial resolution, due to the limited photons in individual bands. Recently, more hyperspectral image super‐resolution methods have been developed by fusing the low‐resolution hyperspectral image and high‐resolution RGB image, but most of them did not consider the misalignment between two input images. In this study, the authors present an effective method to restore a high‐resolution hyperspectral image from the misaligned low‐resolution hyperspectral image and high‐resolution RGB image, which exploits spectral and spatial correlation in hyperspectral and RGB images. Specifically, they employ the spectral sparsity to restore the high‐resolution hyperspectral image on the misaligned part, and then simultaneously employ spectral and spatial structure correlation to restore the high‐resolution hyperspectral image on the aligned area, which can be fused to obtain the high‐quality hyperspectral image restoration under a misaligned hybrid camera system. Experimental results show that the proposed method outperforms the state‐of‐the‐art hyperspectral image super‐resolution methods under a misaligned hybrid camera system in terms of both objective metric and subjective visual quality. Yonggang Lin, Yongrong Zheng, Ying Fu 0001, Hua Huang 0001 |
IET Image Process. | 4 |
| 2018 | No-reference stereopair quality assessment based on singular value decomposition
Lixiong Liu, Hua Huang 0001 |
Neurocomputing | 3 |
| 2018 | Expanding Training Data for Facial Image Super-ResolutionabstractThe quality of training data is very important for learning-based facial image super-resolution (SR). The more similarity between training data and testing input is, the better SR results we can have. To generate a better training set of low/high resolution training facial images for a particular testing input, this paper is the first work that proposes expanding the training data for improving facial image SR. To this end, observing that facial images are highly structured, we propose three constraints, i.e., the local structure constraint, the correspondence constraint and the similarity constraint, to generate new training data, where local patches are expanded with different expansion parameters. The expanded training data can be used for both patch-based facial SR methods and global facial SR methods. Extensive testings on benchmark databases and real world images validate the effectiveness of training data expansion on improving the SR quality. Hua Huang 0001, Chun Qi |
IEEE Trans. Cybern. | 2 |
| 2018 | Hyperspectral Image Super-Resolution With a Mosaic RGB ImageabstractRecently, many hyperspectral (HS) image superresolution methods that merge a low spatial resolution HS image and a high spatial resolution three-channel RGB image have been proposed in spectral imaging. A largely ignored fact is that most existing commercial RGB cameras capture high resolution images by a single CCD/CMOS sensor equipped with a color filter array (CFA). In this paper, we account for the common imaging mechanism of commercial RGB cameras, and propose to use a mosaic RGB image for HS image super-resolution, which prevents demosaicing error and thus its propagation into the HS image super-resolution results. We design a proper nonlocal low-rank regularization to exploit the intrinsic properties - rich self-repeating patterns and high correlation across spectra - within HS images of natural scenes, and formulate the HS image super-resolution task into a variational optimization problem, which can be efficiently solved via the alternating direction method of multipliers (ADMM). The effectiveness of the proposed method has been evaluated on two benchmark datasets, demonstrating that the proposed method can provide substantial improvement over the current state-of-the-art HS image superresolution methods without considering the mosaicing effect. Finally, we show that our method can also perform well in the real capture system. Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001, Imari Sato, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Image Haze Removal via Reference Retrieval and Scene PriorabstractPhotography of hazy scene typically suffers from low-contrast which degrades the visibility of the scene. The performance of single-image dehazing methods is limited by the priors or constraints. In this paper, we present an effective method for haze removal, which utilizes its retrieved correlated haze-free images as external information. The correlated haze-free images are with scene prior offering scene structure and local high frequency information for dehazing, although variations in viewpoints, scales, and illumination conditions exist. To utilize those reference more effectively, global geometric registration and local block matching toward the hazy input are performed to reinforce the spatial correlations. Based on the registration, different kinds of external information are estimated. In addition, we combine that additional external information with internal constraint and regularization for estimating scene transmission map. Experiments demonstrate that our approach can produce dehazing results with better visual quality compared with other state-of-the-art methods. Hua Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Full-Reference Stability Assessment of Digital Video Stabilization Based on Riemannian MetricabstractAssessing the quality of the motion stability is important to evaluating the performance of video stabilization algorithms. This paper presents a novel quality assessment scheme for the video motion stability in a full-reference (FR) manner. Given ideally stable videos and their corresponding shaky videos, our method measures the geodesic distance between motion paths of the stable and the stabilized videos. Due to the use of the Riemannian metric defined on the manifold of spatial transformations, our method enables the intrinsic and faithful measurement on pairwise motion disparities. To facilitate the FR assessment, a data set of stable and shaky videos is constructed by directly capturing realistic stable/shaky videos with a customized device. Then, digital video stabilization algorithms can be run on shaky videos to obtain the stabilized sequence of frames, whereupon their performances are evaluated by using our stability assessment. The experiments demonstrate that our stability assessment gains good concordance with the subjective assessment. Lei Zhang 0021, Qing-Zhuo Zheng, Hong-Kang Liu, Hua Huang 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Camera spectral sensitivity, illumination and spectral reflectance estimation for a hybrid hyperspectral image capture systemabstractA variety of methods have been proposed to restore high resolution hyperspectral image (HSI) from a hybrid camera system, which captures high spatial resolution RGB images and low spatial resolution HSI. They focused unanimously on HSI super-resolution via fusion, yet did not explore the potential of this kind of system for camera spectral sensitivity (CSS), illumination spectrum, and high spatial resolution spectral reflectance recovery. In this paper, we present a sparse representation based method to estimate the CSS of the RGB camera under unknown illumination for the hybrid camera system. Furthermore, the illumination and high spatial resolution spectral reflectance are simultaneously recovered. Experimental results show the effectiveness of the proposed methods on camera spectral sensitivity, illumination spectrum and spectral reflectance recovery. Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001 |
ICIP | 4 |
| 2017 | Binocular spatial activity and reverse saliency driven no-reference stereopair quality assessment
Lixiong Liu, Che-Chun Su, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2017 | A Global Approach to Fast Video StabilizationabstractThis paper presents a novel formulation of video stabilization by directly solving for optimal image warps toward stabilized sequence. With the estimated shaky motion via long or short feature trajectories, our approach encodes another two steps, motion compensation and image warping, into a single global optimization process, rather than operating as two individual steps. This process is done only with positions of embedded mesh vertices as common variables. Spatial and temporal coherence is therein reformulated with similarity-invariant representation of motion trajectories and intra- (and inter-) frame consistency of similar transformations with respect to mesh vertices. Such a one-shot formulation converts video stabilization into a quadratic energy minimization problem defined for image warps, and thus can be efficiently resolved by using a robust solver for sparse linear systems. Experimental results demonstrate the flexibility and efficiency of our approach in producing visually plausible stabilization effects on a variety of videos. Lei Zhang 0021, Qian-Kun Xu, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Bundled Kernels for Nonuniform Blind Video DeblurringabstractWe present a novel blind video deblurring approach by estimating a bundle of kernels and applying the residual deconvolution. Our approach adopts multiple kernels to represent spatially varying motion blur, and thus can cope with nonuniform video deblurring. For each blurred frame, we build a warping-based, space-variant motion blur model based on a bundle of homographies in between its adjacent frames. Then, the nearest sharp frame is employed to form an unblurred-blurred pair for solving the motion model and obtain a bundle of kernels at the blurred frame. Finally, we apply the deconvolution on the residual between the warped unblurred frame and blurred frame with the kernels. The blur kernel estimation and residual deconvolution are iteratively performed toward the deblurred frame, as well as significantly reducing artifacts such as ringings. Experiments show that our approach can efficiently remove the nonuniform video blurring, and achieves better deblurring results than some state-of-the-art methods. Lei Zhang 0021, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Image Quality Assessment Using Directional Anisotropy Structure MeasurementabstractImage quality assessment models prefer an effective visual feature to perceive image quality. Structure-based image quality metrics have verified that a measure of structural information change can provide a good approximation to perceived image distortion. Furthermore, psychological studies have suggested that human beings awareness on image structures is perception-driven and the human visual system (HVS) is more sensitive to the distortion on dominant structures rather than on minor textures. Accordingly, the image distortion can be perceived well by measuring the information loss of the dominant structures. Considering two conclusive psychovisual observations-anisotropy and local directionality, this paper takes a more comprehensive analysis on the behavior of structures and textures, and introduces a directional anisotropic structure measurement (DASM) to represent the dominant structures that are visually important. The proposed DASM can well identify dominant structures, to which the HVS is highly sensitive, from minor textures. Using the DASM as a visual feature, we assess image quality by measuring its degradations. The proposed method was tested on the six benchmark databases and the experimental results demonstrate that our method obtains good performance and correlates well with the human perception. Hua Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Geodesic Video Stabilization in Transformation SpaceabstractWe present a novel formulation of video stabilization in the space of geometric transformations. With the setting of the Riemannian metric, the optimized smooth path is cast as the geodesics on the Lie group embedded in transformation space. While solving the geodesics has a closed-form expression in a certain space, path smoothing can be easily implemented by using geometric interpolation, rather than optimizing any space-time energy function. Specially, by using the geodesic solution in the space of rigid transformations, our approach even gains speedup 10× faster than state-of-the-art methods for path smoothing and motion compensation, and guarantees no extra distortion drawn into the stabilized frames. The experiments demonstrate the efficiency and effectiveness of our algorithm on stabilizing a variety of shaky videos. Lei Zhang 0021, Xiao-Quan Chen, Xin-Yi Kong, Hua Huang 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Pixel-wise video stabilization
Zhongqiang Wang, Hua Huang 0001 |
Multim. Tools Appl. | 2 |
| 2016 | Blind image quality assessment by relative gradient statistics and adaboosting neural network
Lixiong Liu, Qingjie Zhao, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2015 | Efficient Variational Light Field View Synthesis For Making Stereoscopic 3D ImagesabstractWe present a novel approach for making stereoscopic images by variational view synthesis on the multi-perspective light field. With the intended disparities as constraints, we specialize the generative variational model by incorporating per-pixel viewpoint assignment to synthesize the stereo pair. Also, we improve the variational solution by use of explicit weighted average on the light field. Our algorithm is able to handle arbitrary disparity remapping, thus enabling more flexible disparity control for the desired stereoscopic effect. The experiments demonstrate the effectiveness and efficiency for making the stereoscopic 3D images based on the light field. Lei Zhang 0021, Yuhang Zhang 0005, Hua Huang 0001 |
Comput. Graph. Forum | 3 |
| 2015 | Guided Adaptive Image Smoothing via Directional Anisotropic Structure MeasurementabstractImage smoothing prefers a good metric to identify dominant structures from textures adaptive of intensity contrast. In this paper, we drop on a novel directional anisotropic structure measurement (DASM) toward adaptive image smoothing. With observations on psychological perception regarding anisotropy, non-periodicity and local directionality, DASM can well characterize structures and textures independent on their contrast scales. By using such measurement as constraint, we design a guided adaptive image smoothing scheme by improving extrema localization and envelopes construction in a structure-aware manner. Our approach can well suppresses the staircase-like artifacts and blur of structures that appear in previous methods, which better suits structure-preserving image smoothing task. The algorithm is performed on a space-filling curve as the reduced domain, so it is very fast and much easy to implement in practice. We make comprehensive comparisons with previous state-of-the-art methods for a variety of applications. Experimental results demonstrate the merit using our DASM as metric to identify structures, and the effectiveness and efficiency of our adaptive image smoothing approach to produce commendable results. Hua Huang 0001, Lei Zhang 0021 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | No-reference image quality assessment in curvelet domain
Lixiong Liu, Hongping Dong, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2014 | No-reference image quality assessment based on spatial and spectral entropies
Lixiong Liu, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2014 | Efficient Structure-Aware Image Smoothingby Local Extrema on Space-Filling CurveabstractThis paper presents a novel image smoothing approach using a space-filling curve as the reduced domain to perform separation of edges and details. This structure-aware smoothing effect is achieved by modulating local extrema after empirical mode decomposition; it is highly effective and efficient since it is implemented on a one-dimensional curve instead of a two-dimensional image grid. To overcome edge staircase-like artifacts caused by a neighborhood deficiency in domain reduction, we next use a joint contrast-based filter to consolidate edge structures in image smoothing. The adoption of dimensional reduction makes our smoothing approach distinct for two reasons. First, overall structure-awareness is improved as more extrema are exploited to locate the salient edges and details. Second, envelope computation for local extrema is made much fast by using explicit interpolants on the curve. Moreover, our approach is simple and very easy to implement in practice. Experimental results demonstrate the merit of our approach, which outperforms previous state-of-the-art methods, for a variety of image processing tasks. Hua Huang 0001, Lei Zhang 0021 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | VideoGraph: a non-linear video representation for efficient exploration
Lei Zhang 0021, Qian-Kun Xu, Lei-Zheng Nie, Hua Huang 0001 |
Vis. Comput. | 4 |
| 2014 | Artistic preprocessing for painterly rendering and image stylization
Hua Huang 0001, Chenfeng Li |
Vis. Comput. | 2 |
| 2013 | Multiplane Video StabilizationabstractAbstract This paper presents a novel video stabilization approach by leveraging the multiple planes structure of video scene to stabilize inter‐frame motion. As opposed to previous stabilization procedure operating in a single plane, our approach primarily deals with multiplane videos and builds their multiple planes structure for performing stabilization in respective planes. Hence, a robust plane detection scheme is devised to detect multiple planes by classifying feature trajectories according to reprojection errors generated by plane induced homographies. Then, an improved planar stabilization technique is applied by conforming to the compensated homography in each plane. Finally, multiple stabilized planes are coherently fused by content‐preserving image warps to obtain the output stabilized frames. Our approach does not need any stereo reconstruction, yet is able to produce commendable results due to awareness of multiple planes structure in the stabilization. Experimental results demonstrate the effectiveness and efficiency of our approach to robust stabilization on multiplane videos. Zhongqiang Wang, Lei Zhang 0021, Hua Huang 0001 |
Comput. Graph. Forum | 3 |
| 2013 | Stroke Style Analysis for Painterly Rendering
Hua Huang 0001, Chenfeng Li |
J. Comput. Sci. Technol. | 2 |
| 2012 | Hierarchical Narrative Collage For Digital Photo AlbumabstractAbstract Collage can provide a summary form on the collection of photos in an album. In this paper, we introduce a novel approach to constructing photo collage in the hierarchical narrative manner. As opposed to previous methods focusing on spatial coherence in the collage layout, our narrative collage arranges the photos according to the basic narrative elements from literary writings, i.e., character, setting and plot. Face, time and place attributes are exploited to embody those narrative elements in the collage. Then, photos are organized into the hierarchical structure for the multi‐level details in the events recorded by the album. Such hierarchical narrative collage can present a visual overview in the chronological order on what happened in the album. Experimental results show that our approach offers a better summarization to browse on the photo album content than previous ones. Lei Zhang 0021, Hua Huang 0001 |
Comput. Graph. Forum | 2 |
| 2012 | A fast two-step block type decision algorithm for intra prediction in H.264/AVC high profile
Ping Wang 0009, Hua Huang 0001 |
Multim. Tools Appl. | 2 |
| 2012 | Super-Resolution Method for Multiview Face Recognition From a Single Image Per Person Using Nonlinear Mappings on Coherent FeaturesabstractIn video surveillance, the face recognition usually aims at recognizing a nonfrontal low resolution face image from the gallery in which each person has only one high resolution frontal face image. Traditional face recognition approaches have several challenges, such as the difference of image resolution, pose variation and only one gallery image per person. This letter presents a regression based method that can successfully recognize the identity given all these difficulties. The nonlinear regression models from the specific nonfrontal low resolution image to frontal high resolution features are learnt by radial basis function in subspace built by canonical correlation analysis. Extensive experiments on benchmark database show the superiority of our method. Hua Huang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2012 | EXCOL: An EXtract-and-COmplete Layering Approach to Cartoon Animation ReusingabstractWe introduce the EXtract-and-COmplete Layering method (EXCOL)--a novel cartoon animation processing technique to convert a traditional animated cartoon video into multiple semantically meaningful layers. Our technique is inspired by vision-based layering techniques but focuses on shape cues in both the extraction and completion steps to reflect the unique characteristics of cartoon animation. For layer extraction, we define a novel similarity measure incorporating both shape and color of automatically segmented regions within individual frames and propagate a small set of user-specified layer labels among similar regions across frames. By clustering regions with the same labels, each frame is appropriately partitioned into different layers, with each layer containing semantically meaningful content. Then, a warping-based approach is used to fill missing parts caused by occlusion within the extracted layers to achieve a complete representation. EXCOL provides a flexible way to effectively reuse traditional cartoon animations with only a small amount of user interaction. It is demonstrated that our EXCOL method is effective and robust, and the layered representation benefits a variety of applications in cartoon animation processing. Lei Zhang 0021, Hua Huang 0001, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Web-image driven best views of 3D shapes
Lei Zhang 0021, Hua Huang 0001 |
Vis. Comput. | 3 |
| 2011 | RepSnapping: Efficient Image Cutout for Repeated Scene ElementsabstractAbstract Repeated scene elements are copious and ubiquitous in natural images. Cutout of those repeated elements usually involves tedious and laborious user interaction by previous image segmentation methods. In this paper, we present RepSnapping, a novel method oriented to cutout of repeated scene elements with much less user interaction. By exploring inherent similarity between repeated elements, a new optimization model is introduced to thread correlated elements in the segmentation procedure. The model proposed here enables efficient solution using max‐flow/min cut on an extended graph. Experiments indicate thatRepSnappingfacilitates cutout of repeated elements better than the state‐of‐the‐art interactive image segmentation and repetition detection methods. Hua Huang 0001, Lei Zhang 0021 |
Comput. Graph. Forum | 1 |
| 2011 | Fast feature-based mode decision for 4×4 intra prediction in H.264/AVC
Ping Wang 0009, Hua Huang 0001 |
Sci. China Inf. Sci. | 2 |
| 2011 | Fast Facial Image Super-Resolution via Local Linear Transformations for Resource-Limited ApplicationsabstractMost popular learning-based super-resolution (SR) approaches suffer from complicated learning structures and highly intensive computation, especially in resource-limited applications. We propose a novel frontal facial image SR approach by using multiple local linear transformations to approximate the nonlinear mapping between low-resolution (LR) and high-resolution (HR) images in the pixel domain. We adopt Procrustes analysis to obtain orthogonal matrices representing the learned linear transformations, which cannot only well capture appearance variations in facial patches but also greatly simplify the transformation computation to matrices manipulation. An HR image can be directly reconstructed from a single LR image without need of the large training data, thus avoiding the use of a large redundant LR and HR patch database. Experimental results show that our approach is computationally fast as well the SR quality compares favorably with the state-of-the-art approaches from both subjective and objective evaluations. Besides, our approach is insensitive to the size of training data and robust to a wide range of facial variations like occlusions. More importantly, the proposed method is also much more effective than other comparative methods to reconstruct real-world images captured from the Internet and webcams. Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Super-Resolution Method for Face Recognition Using Nonlinear Mappings on Coherent FeaturesabstractLow-resolution (LR) of face images significantly decreases the performance of face recognition. To address this problem, we present a super-resolution method that uses nonlinear mappings to infer coherent features that favor higher recognition of the nearest neighbor (NN) classifiers for recognition of single LR face image. Canonical correlation analysis is applied to establish the coherent subspaces between the principal component analysis (PCA) based features of high-resolution (HR) and LR face images. Then, a nonlinear mapping between HR/LR features can be built by radial basis functions (RBFs) with lower regression errors in the coherent feature space than in the PCA feature space. Thus, we can compute super-resolved coherent features corresponding to an input LR image according to the trained RBF model efficiently and accurately. And, face identity can be obtained by feeding these super-resolved features to a simple NN classifier. Extensive experiments on the Facial Recognition Technology, University of Manchester Institute of Science and Technology, and Olivetti Research Laboratory databases show that the proposed method outperforms the state-of-the-art face recognition algorithms for single LR image in terms of both recognition rate and robustness to facial variations of pose and expression. Hua Huang 0001, Huiting He |
IEEE Trans. Neural Networks | 1 |
| 2011 | Arcimboldo-like collage using internet imagesabstractCollage is a composite artwork made from assemblage of different material forms. In this work, we present a novel approach for creating a fantastic collage artform, namely Arcimboldo-like collage, which represents an input image with multiple thematically-related cutouts from the filtered Internet images. Due to the massive data of Internet images, competent image cutouts can almost always be discovered to match the segmented components of the input image. The selected cutouts are purposefully arranged such that as a whole assembly, they can represent the input image with disguise in both shape and color; but separately, individual cutout is still recognizable as its own being. Experimental results and user study show that our algorithm can effectively produce the entertaining Arcimboldo-like collages. Hua Huang 0001, Lei Zhang 0021 |
ACM Trans. Graph. | 1 |
| 2011 | Painterly rendering with content-dependent natural paint strokes
Hua Huang 0001, TianNan Fu, Chenfeng Li |
Vis. Comput. | 1 |
| 2010 | An improved motion-search method based on pattern classificationabstractAn improved motion-search method based on pattern classification is proposed in this paper. A new feature representing the maximum motion around the current block is introduced, allowing more precise description of the motion characteristics of the block. Ada-Boost classifiers can correctly classify the harder-to-classify samples by increasing the weights of misclassified samples, so are adopted to replace the original linear classifiers. Experiments show that the improved method has lower temporal cost in motion estimation search while providing the same image quality. Hua Huang 0001 |
ICASSP | 4 |
| 2010 | Face image super resolution by linear transformationabstractA novel two-step super-resolution (SR) method for face images is proposed in this paper. The critical issue of global face reconstruction in the two-step SR framework is to construct the relationship between high resolution (HR) and low resolution (LR) features. We choose the Principal Component Analysis (PCA) coefficients of LR/HR face images as the features for global faces. These features are considered as inputs and outputs of an unknown linear system. The mapping between the inputs and outputs is estimated from training sets as the system response. The HR features corresponding to a test LR image can be obtained by applying the learnt mapping to the LR features, and hence we can reconstruct the global face. Ultimately, an HR face image is generated by using the patch-based neighbor reconstruction that imposes facial details into the global face. Experiments indicate that our method produces HR faces of higher quality and is easier to implement than traditional methods based on two-step framework. Hua Huang 0001, Chun Qi |
ICIP | 1 |
| 2010 | Orthogonal 4-tap integer multiwavelet transforms using matrix factorizationabstractAn algorithm for orthogonal 4-tap integer multiwavelet transforms is proposed. Some remarkable properties of orthogonal matrix are presented. Furthermore, the transform matrix is rewritten in a product of two block diagonal matrices and a permutation matrix by the singular value decomposition (SVD) of block recursive matrices. Each block of block diagonal matrices is factorized into triangular elementary reversible matrices (TERMs), which can map integers to integers by rounding arithmetic. Experiment results show that the proposed algorithm is an executable algorithm and outperforms the existing orthogonal 4-tap integer multiwavelet transform algorithm. Mingli Jing, Hua Huang 0001, WuLing Liu, Chun Qi |
ICIP | 2 |
| 2010 | Mesh reconstruction by meshless denoising and parameterization
Lei Zhang 0021, Ligang Liu 0001, Craig Gotsman, Hua Huang 0001 |
Comput. Graph. | 4 |
| 2010 | Video Painting via Motion Layer ManipulationabstractAbstract Temporal coherence is an important problem in Non‐Photorealistic Rendering for videos. In this paper, we present a novel approach to enhance temporal coherence in video painting. Instead of painting on video frame, our approach first partitions the video into multiple motion layers, and then places the brush strokes on the layers to generate the painted imagery. The extracted motion layers consist of one background layer and several object layers in each frame. Then, background layers from all the frames are aligned into a panoramic image, on which brush strokes are placed to paint the background in one‐shot. The strokes used to paint object layers are propagated frame by frame using smooth transformations defined by thin plate splines. Once the background and object layers are painted, they are projected back to each frame and blent to form the final painting results. Thanks to painting a single image, our approach can completely eliminate the flickering in background, and temporal coherence on object layers is also significantly enhanced due to the smooth transformation over frames. Additionally, by controlling the painting strokes on different layers, our approach is easy to generate painted video with multi‐style. Experimental results show that our approach is both robust and efficient to generate plausible video painting. Hua Huang 0001, Lei Zhang 0021, TianNan Fu |
Comput. Graph. Forum | 1 |
| 2010 | RBF network-based temporal color morphingabstractAbstract A method of RBF network‐based temporal color morphing is proposed to simulate the natural phenomena characterized by temporal color alteration, e.g., turning green of foliage, resurgence of leaves. Such phenomena usually span a long time and it is very difficult to capture their whole process. Our system accepts a source image sequence and a reference image as input. The source sequence contains the desired scene except for color alteration, and the reference image has the color style which the source sequence is expected to advance into. First, an RBF network is employed to model the mapping between the colors of the source sequence and the reference image. Then, a simple interpolation algorithm is applied to render the resulting sequence. The effectiveness of the new method is verified by experiments. Copyright © 2010 John Wiley & Sons, Ltd. XueZhong Xiao, Hua Huang 0001, Lizhuang Ma |
Comput. Animat. Virtual Worlds | 2 |
| 2010 | Super-resolution of human face image using canonical correlation analysis
Hua Huang 0001, Huiting He, Junping Zhang |
Pattern Recognit. | 1 |
| 2010 | A Simple Approach to Multiview Face HallucinationabstractMost face hallucination methods are usually limited to frontal face with small pose variations. This letter presents a simple and efficient multiview face hallucination (MFH) method to generate high-resolution (HR) multiview faces from a single given low-resolution (LR) one. The problem is addressed in two steps. A simple face transformation method is proposed by defining a constrained least square problem for LR multiview face transformation and a position-patch based face hallucination method is extended to incorporate HR multiview face details. Experimental results show that our approach has some advantages over existing MFH methods. Hua Huang 0001, Shaopeng Wang, Chun Qi |
IEEE Signal Process. Lett. | 2 |
| 2010 | Example-based contrast enhancement by gradient mapping
Hua Huang 0001, XueZhong Xiao |
Vis. Comput. | 1 |
| 2010 | Example-based painting guided by color features
Hua Huang 0001, Chenfeng Li |
Vis. Comput. | 1 |
| 2009 | Real-time content-aware image resizing
Hua Huang 0001, TianNan Fu, Paul L. Rosin, Chun Qi |
Sci. China Ser. F Inf. Sci. | 1 |
| 2009 | Edge-Aware Level Set Diffusion and Bilateral Filtering Reconstruction for Image Magnification
Hua Huang 0001, Paul L. Rosin, Chun Qi |
J. Comput. Sci. Technol. | 1 |
| 2009 | Neighbor embedding based super-resolution algorithm through edge detection and feature selection
Tak-Ming Chan, Junping Zhang, Jian Pu, Hua Huang 0001 |
Pattern Recognit. Lett. | 4 |
| 2008 | Balanced Multiwavelets Based Digital Image Watermarking
Hua Huang 0001, Chun Qi |
IWDW | 2 |
| 2006 | A hybrid parallel projection approach to object-based image restoration
Hua Huang 0001, Dequn Liang, Chun Qi |
Pattern Recognit. Lett. | 2 |
| 2005 | Probabilistic Contour Extraction Using Hierarchical Shape RepresentationabstractIn this paper, we address the issue of extracting contour of the object with a specific shape. A hierarchical graphical model is proposed to represent shape variations. A complex shape is decomposed into several components which are described as principal component analysis (PCA) based models in various levels. The hierarchical representation allows for chain-like conditional dependency within a single level and bidirectional communication between different levels. Additionally, a sequential Monte-Carlo (SMC) based inference algorithm that can explore the graphical structure is proposed to estimate the contour. The experiments performed on real-world hand and face images show that the proposed method is effective in combating occlusion and cluttered background. Moreover, it is possible to isolate the localization error to an individual component of a shape attributed to the hierarchical representation. Chun Qi, Dequn Liang, Hua Huang 0001 |
ICCV | 4 |