Lei Zhang 0021

dblp:64/5666-21 · DBLP profile ↗
← Back
51ranked-venue papers
14as first author
24since 2021 · last 2026
0000-0002-2286-0314ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 13 first-author · 18 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 WonderTex: Consistent-and-Seamless Texture Generation With Text-Guided Multi-View Image Diffusion Models
abstract
Text-guided texture generation has been rapidly developed with the proliferation of generative artificial intelligence for creating three-dimensional textured objects. However, existing text-guided texture generation methods often suffer from artifacts such as inconsistent visual appearance across different views, the Janus problems and seams in texture maps. To address these issues, we propose a novel text-guided texture generation method, named WonderTex. It achieves the generation of high-quality, view-consistent, and seamless texture maps through a two-stage pipeline. Specifically, we fine-tune a Stable Diffusion model using a large dataset to obtain a multi-view image diffusion model capable of generating a 4-view grid. This model serves as the foundation for producing four consistent views and establishing the base texture in the first stage. Subsequently, an automatic view selection and inpainting strategy is employed to effectively fill and refine the texture maps in the second stage. Extensive experiments have shown that our method is effective and robust, capable of generating high-quality textures with various meshes and prompts, outperforming baseline methods in terms of texture details, view consistency, and other metrics.
Xiaoguang Han 0001, Lei Zhang 0021
IEEE Trans. Vis. Comput. Graph.3
2026 ScribbleSense: Generative Scribble-Based Texture Editing With Intent Prediction
abstract
Interactive 3D model texture editing presents enhanced opportunities for creating 3D assets, with freehand drawing style offering the most intuitive experience. However, existing methods primarily support sketch-based interactions for outlining, while the utilization of coarse-grained scribble-based interaction remains limited. Furthermore, current methodologies often encounter challenges due to the abstract nature of scribble instructions, which can result in ambiguous editing intentions and unclear target semantic locations. To address these issues, we propose ScribbleSense, an editing method that combines multimodal large language models (MLLMs) and image generation models to effectively resolve these challenges. We leverage the visual capabilities of MLLMs to predict the editing intent behind the scribbles. Once the semantic intent of the scribble is discerned, we employ globally generated images to extract local texture details, thereby anchoring local semantics and alleviating ambiguities concerning the target semantic locations. Experimental results indicate that our method effectively leverages the strengths of MLLMs, achieving state-of-the-art interactive editing performance for scribble-based texture editing.
Yeming Geng, Lei Zhang 0021
IEEE Trans. Vis. Comput. Graph.3
2025 Talking Head Generation via Viewpoint and Lighting Simulation Based on Global Representation
abstract
NeRF-based talking head generation has made great progress, but existing methods still lack in achieving high-quality detail fidelity, mainly manifested in detail loss and intermittent blur. We attribute this to the limitations of the training video data in terms of viewpoint and lighting, which leads to the inability to fully model the global depth and brightness information of spatial points. Specifically, a fixed viewpoint may fail to provide sufficient depth information for high-frequency details, leading to inaccurate volume density estimation and the loss of details such as hair. Furthermore, constant lighting often fails to adapt to the drastic brightness changes of continuous video frames, resulting in color accumulation errors and blurring artifacts. To address these issues, we propose a novel talking head generation method that combines layered viewpoint simulation (LVS) and continuous lighting simulation (CLS). LVS simulates multiple viewpoints through the multi-scale features of the video frame to construct the global depth representation, which can improve the accuracy of volume density estimation and enhance detail description. CLS simulates multiple lighting through brightness changes of continuous video frames to construct the global brightness representation, thereby alleviating color accumulation errors and eliminating blur. Extensive experiments demonstrate that our method significantly improves the detail quality compared to the state-of-the-art methods.
Biao Dong, Lei Zhang 0021
ACM Multimedia2
2025 RTR-GS: 3D Gaussian Splatting for Inverse Rendering with Radiance Transfer and Reflection
abstract
3D Gaussian Splatting (3DGS) has demonstrated impressive capabilities in novel view synthesis. However, rendering reflective objects remains a significant challenge, particularly in inverse rendering and relighting. We introduce RTR-GS, a novel inverse rendering framework capable of robustly rendering objects with arbitrary reflectance properties, decomposing BRDF and lighting, and delivering credible relighting results. Given a collection of multi-view images, our method effectively recovers geometric structure through a hybrid rendering model that combines forward rendering for radiance transfer with deferred rendering for reflections. This approach successfully separates high-frequency and low-frequency appearances, mitigating floating artifacts caused by spherical harmonic overfitting when handling high-frequency details. We further refine BRDF and lighting decomposition using an additional physically-based deferred rendering branch. Experimental results show that our method enhances novel view synthesis, normal estimation, decomposition, and relighting while maintaining efficient training inference process.
Yongyang Zhou, Lei Zhang 0021
ACM Multimedia4
2025 Novel view synthesis with wide-baseline stereo pairs based on local-global information
Lei Zhang 0021
Comput. Graph.2
2025 Consistent Image Layout Editing With Diffusion Models
abstract
Despite the great success of large-scale text-to-image diffusion models in image generation and image editing, existing methods still struggle with editing the layout of real-world images. Although a few works have been developed to address this issue, they either fail to adjust the image layout effectively or encounter challenges in preserving the visual appearance of objects after layout adjustment. To bridge this gap, this paper proposes a novel image layout editing method that not only re-arranges a real-world image to a specified layout, but also ensures that the visual appearance of the objects remains consistent with their original state prior to editing. Concretely, a Multi-Concept Learning scheme is developed to learn the concepts of different objects from a single image, which can be seen as a novel inversion scheme tailored for image layout editing. Then, we leverage the semantic consistency within intermediate features of diffusion models to project the appearance information of objects to the target regions to improve the fidelity of objects after editing. Additionally, a novel initialization noise design is adopted to facilitate the convergence and success rate of re-arranging the layout. The phenomenon of concept entanglement is also analyzed, and resolved by a novel asynchronous editing strategy. Extensive experimental results demonstrate that the proposed method outperforms existing methods in both layout alignment and visual consistency for the task of image layout editing.
Ting Liu 0018, Lei Zhang 0021
IEEE Trans. Image Process.4
2025 IMU-Assisted Gray Pixel Shift for Video White Balance Stabilization
abstract
Video white balance is to correct the scene color of video frames to the color under the standard white illumination. Due to the camera movement, video white balance usually suffers temporal instability with unnatural color change between frames. This paper presents a video white balance stabilization method for spatially correct and temporally stable color correction. It exploits the color invariance at the position of the same object to obtain the consistent illumination color estimation through frames. Specifically, it detects gray pixels that inherit the potential illumination color, and their inter-frame motion calculated with the assistance of inertial measurement unit (IMU) is used to carry gray pixels for establishing their correspondence and color fusion between adjacent frames. Because the IMU has more robust and accurate motion cues against large camera movement and texture-less regions in the scene, our method can generate better gray pixel correspondences and illumination color estimation for the white balance stabilization. Besides, our method is computationally efficient to be deployed on mobile phones. Experimental results show that our method can significantly improve the temporal stability as well as maintain the spatial correctness of white balance for videos recorded by cameras equipped with IMU sensors.
Lei Zhang 0021
IEEE Trans. Multim.1
2025 Voxel-Mesh Hybrid Representation for Real-Time View Synthesis by Meshing Density Field
abstract
The neural radiance fields (NeRF) have emerged as a prominent methodology for synthesizing realistic images of novel views. While neural radiance representations based on voxels or mesh individually offer distinct advantages, excelling in either rendering quality or speed, each has limitations in the other aspect. In response, we propose a hybrid representation named Vosh, seamlessly combining both voxel and mesh components in hybrid rendering for view synthesis. Vosh is meticulously crafted by optimizing the voxel grid based on neural rendering, strategically meshing a portion of the volumetric density field to surface. Therefore, it excels in fast rendering scenes with simple geometry and textures through its mesh component, while simultaneously enabling high-quality rendering in intricate regions by leveraging voxel component. The flexibility of Vosh is showcased through the ability to adjust hybrid ratios, providing users the ability to control the balance between rendering quality and speed based on flexible usage. Experimental results demonstrate that our method achieves commendable trade-off between rendering quality and speed, and notably has real-time performance on mobile devices.
Chenhao Zhang 0001, Yongyang Zhou, Lei Zhang 0021
IEEE Trans. Vis. Comput. Graph.3
2024 Non-Serial Quantization-Aware Deep Optics for Snapshot Hyperspectral Imaging
abstract
Deep optics has been endeavoring to capture hyperspectral images of dynamic scenes, where the optical encoder plays an essential role in deciding the imaging performance. Our key insight is that the optical encoder of a deep optics system is expected to keep fabrication-friendliness and decoder-friendliness, to be faithfully realized in the implementation phase and fully interacted with the decoder in the design phase, respectively. In this paper, we propose the non-serial quantization-aware deep optics (NSQDO), which consists of the fabrication-friendly quantization-aware model (QAM) and the decoder-friendly non-serial manner (NSM). The QAM integrates the quantization process into the optimization and adaptively adjusts the physical height of each quantization level, reducing the deviation of the physical encoder from the numerical simulation through the awareness of and adaptation to the quantization operation of the DOE physical structure. The NSM bridges the encoder and the decoder with full interaction through bidirectional hint connections and flexibilize the connections with a gating mechanism, boosting the power of joint optimization in deep optics. The proposed NSQDO improves the fabrication-friendliness and decoder-friendliness of the encoder and develops the deep optics framework to be more practical and powerful. Extensive synthetic simulation and real hardware experiments demonstrate the superior performance of the proposed method.
Lizhi Wang 0001, Lingen Li, Lei Zhang 0021, Zhiwei Xiong, Hua Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Hyperbolic Space-Based Autoencoder for Hyperspectral Anomaly Detection
abstract
Deep-learning (DL)-based methods have been shown to be effective on the hyperspectral image (HSI) anomaly detection task because of their feature extraction ability. However, current DL-based methods lack an effective means of regularizing the background information. In this article, the hyperbolic space-based autoencoder (HSAE) is proposed for the hyperspectral anomaly detection task. We assume that an effective hierarchical structural representation can better model the HSI in the spatial domain, and this enables the background information to be effectively regularized. Motivated by this idea, the HSAE embeds the HSI into hyperbolic space, which is a non-Euclidean geometry with a constant negative curvature and an exponential growth distance between points. Using a wrapped normal prior distribution, the training of the hidden representation is supervised to preserve more hierarchical features. After the training process, a hyperbolic distance-based anomaly detector (HDB) is introduced to discover anomalies in a more robust way. Experimental results on several popular HSI benchmarks fully demonstrate the superiority of our HSAE.
He Sun 0009, Lizhi Wang 0001, Lei Zhang 0021, Lianru Gao
IEEE Trans. Geosci. Remote. Sens.3
2024 Intrinsic Omnidirectional Image Decomposition With Illumination Pre-Extraction
abstract
Capturing an omnidirectional image with a 360-degree field of view entails capturing intricate spatial and lighting details of the scene. Consequently, existing intrinsic image decomposition methods face significant challenges when attempting to separate reflectance and shading components from a low dynamic range (LDR) omnidirectional images. To address this, our article introduces a novel method specifically designed for the intrinsic decomposition of omnidirectional images. Leveraging the unique characteristics of the 360-degree scene representation, we employ a pre-extraction technique to isolate specific illumination information. Subsequently, we establish new constraints based on these extracted details and the inherent characteristics of omnidirectional images. These constraints limit the illumination intensity range and incorporate spherical-based illumination variation. By formulating and solving an objective function that accounts for these constraints, our method achieves a more accurate separation of reflectance and shading components. Comprehensive qualitative and quantitative evaluations demonstrate the superiority of our proposed method over state-of-the-art intrinsic decomposition methods.
Rong-Kai Xu, Lei Zhang 0021
IEEE Trans. Vis. Comput. Graph.2
2023 Revisiting Unsupervised Local Descriptor Learning
abstract
Constructing accurate training tuples is crucial for unsupervised local descriptor learning, yet challenging due to the absence of patch labels. The state-of-the-art approach constructs tuples with heuristic rules, which struggle to precisely depict real-world patch transformations, in spite of enabling fast model convergence. A possible solution to alleviate the problem is the clustering-based approach, which can capture realistic patch variations and learn more accurate class decision boundaries, but suffers from slow model convergence. This paper presents HybridDesc, an unsupervised approach that learns powerful local descriptor models with fast convergence speed by combining the rule-based and clustering-based approaches to construct training tuples. In addition, HybridDesc also contributes two concrete enhancing mechanisms: (1) a Differentiable Hyperparameter Search (DHS) strategy to find the optimal hyperparameter setting of the rule-based approach so as to provide accurate prior for the clustering-based approach, (2) an On-Demand Clustering (ODC) method to reduce the clustering overhead of the clustering-based approach without eroding its advantage. Extensive experimental results show that HybridDesc can efficiently learn local descriptors that surpass existing unsupervised local descriptors and even rival competitive supervised ones.
Wufan Wang, Lei Zhang 0021, Hua Huang 0001
AAAI2
2023 Semi-direct Sparse Odometry with Robust and Accurate Pose Estimation for Dynamic Scenes
Wufan Wang, Lei Zhang 0021
CAD/Graphics2
2023 SCP-SLAM: Accelerating DynaSLAM With Static Confidence Propagation
abstract
DynaSLAM is the state-of-the-art visual simultaneous localization and mapping (SLAM) in dynamic environments. It adopts a convolutional neural network (CNN) for moving object detection, but usually incurs a very high computational cost because it performs semantic segmentation using the CNN model on every frame. This paper proposes SCP-SLAM, which accelerates DynaSLAM by running the CNN only on keyframes and propagating static confidence through other frames in parallel. The proposed static confidence characterizes the moving object features by the residual defined by inter-frame geometry transformation, which can be computed quickly. Our method combines the effectiveness of a CNN with the efficiency of static confidence in a tightly coupled manner. Extensive experiments on the publicly available TUM and Bonn RGB-D dynamic benchmark datasets demonstrate the efficacy of the method. Compared with DynaSLAM, it enables acceleration by a factor of ten on average, but retains comparable localization accuracy.
Mingfei Yu, Lei Zhang 0021, Wu-Fan Wang
VR2
2023 Image Stitching With Manifold Optimization
abstract
Image stitching usually relies on spatial transformations to perform the overlap alignment and distortion mitigation. This paper presents a manifold optimization method to seek these transformations. The purpose is not to present a new formulation of image stitching, as the proposed method uses common transformations such as homography to align feature correspondences in the overlap and similarity transformations to preserve the shape. Instead, the proposed method is based on a new treatment of these transformations as elements of a prescribed matrix manifold. Its advantage lies in its more effective and efficient optimization in the manifold domain. Specifically, spatially varying homographies are computed by an efficient second-order minimization (ESM) of the geometric error of aligning feature correspondences, but with their intrinsic manifold parameterization. To mitigate the distortion, the interpolation between homography and similarity transformation is performed on a general matrix manifold. These on-manifold operations improve the stitching quality with fewer ghosting and distortion artifacts. The experiments show our manifold optimization for image stitching outperforms other methods.
Lei Zhang 0021, Hua Huang 0001
IEEE Trans. Multim.1
2022 Quantization-aware Deep Optics for Diffractive Snapshot Hyperspectral Imaging
abstract
Diffractive snapshot hyperspectral imaging based on the deep optics framework has been striving to capture the spectral images of dynamic scenes. However, existing deep optics frameworks all suffer from the mismatch between the optical hardware and the reconstruction algorithm due to the quantization operation in the diffractive optical element (DOE) fabrication, leading to the limited performance of hyperspectral imaging in practice. In this paper, we propose the quantization-aware deep optics for diffractive snapshot hyperspectral imaging. Our key observation is that common lithography techniques used in fabricating DOEs need to quantize the DOE height map to a few levels, and can freely set the height for each level. Therefore, we propose to integrate the quantization operation into the DOE height map optimization and design an adaptive mechanism to adjust the physical height of each quantization level. According to the optimization, we fabricate the quantized DOE directly and build a diffractive hyperspectral snapshot imaging system. Our method develops the deep optics framework to be more practical through the awareness of and adaptation to the quantization operation of the DOE physical structure, making the fabricated DOE and the reconstruction algorithm match each other systematically. Extensive synthetic simulation and real hardware experiments validate the superior performance of our method.
Lingen Li, Lizhi Wang 0001, Lei Zhang 0021, Zhiwei Xiong, Hua Huang 0001
CVPR4
2022 Progressive Unsupervised Learning of Local Descriptors
abstract
Training tuple construction is a crucial step in unsupervised local descriptor learning. Existing approaches perform this step relying on heuristics, which suffer from inaccurate supervision signals and struggle to achieve the desired performance. To address the problem, this work presents DescPro, an unsupervised approach that progressively explores both accurate and informative training tuples for model optimization without using heuristics. Specifically, DescPro consists of a Robust Cluster Assignment (RCA) method to infer pairwise relationships by clustering reliable samples with the increasingly powerful CNN model, and a Similarity-weighted Positive Sampling (SPS) strategy to select informative positive pairs for training tuple construction. Extensive experimental results show that, with the collaboration of the above two modules, DescPro can outperform state-of-the-art unsupervised local descriptors and even rival competitive supervised ones on standard benchmarks.
Wufan Wang, Lei Zhang 0021, Hua Huang 0001
ACM Multimedia2
2022 Sensitivity-Aware Spatial Quality Adaptation for Live Video Analytics
abstract
To address the conflict between the limited network bandwidth and high DNN inference accuracy, live video analytics desires a bandwidth-efficient streaming approach. To this end, more and more works study spatially variable quality streaming where high quality is only used for important regions. The key challenges are to accurately identify the important regions and select the right qualities for them to maximize accuracy. Existing approaches use either cheap analytics models or low-quality videos to locate important regions, and employ heuristic rules to make quality decisions, which struggle to address the above challenges. Our key insight is that the region’s accuracy “sensitivity” obtained by running the expensive DNN model on the high-quality video provides a reliable indication of the region’s importance and allows to allocate the available bandwidth optimally over regions by explicitly maximizing the frame accuracy. This work presents a sensitivity-aware algorithm Orchestra, which incorporates sensitivity into the design of spatial quality adaptation, including video zoning and quality selection. The design of Orchestra entails three main contributions: a feasible way of sensitivity estimation, sensitivity-aware zoning, and deduction-based accuracy estimation. Extensive experiments over realistic videos and network traces show that Orchestra improves accuracy by upto 14.1% with comparable bandwidth usage or reduces bandwidth usage by upto 44.2% while maintaining higher accuracy compared to baselines.
Wufan Wang, Lei Zhang 0021, Hua Huang 0001
IEEE J. Sel. Areas Commun.3
2022 Stochastic gate-based autoencoder for unsupervised hyperspectral band selection
He Sun 0009, Lei Zhang 0021, Lizhi Wang 0001, Hua Huang 0001
Pattern Recognit.2
2022 Robust Extraction and Super-Resolution of Low-Resolution Flying Airplane From Satellite Video
abstract
Extracting the flying airplane from the satellite video and enhancing its resolution are significant and demanding tasks in the remote sensing community. The challenge mainly lies in that the flying airplane target in the satellite video often suffers from detail loss due to complex background and limited spaceborne imaging device. In this article, a novel constructive model is proposed to model the airplane of low resolution for more complete extraction, and a new reflective symmetry shape prior is integrated into the super-resolution process to obtain the higher resolution result. Concretely, each frame can be decomposed as a linear combination of foreground and background with specific mixture ratios. With the assumption of uniform linear motion and the rigidity of the airplane, a periodic change of mixture ratios through frames is induced, which can construct the airplane as complete as possible by adopting the proposed iterative matting optimization. To further enhance the resolution of the extracted airplane, an improved alternating direction method of multipliers (ADMM) is utilized to solve the super-resolution problem with the reflective symmetry of the shape as prior. The effectiveness of our method with respect to extraction and super-resolution is borne out by the experiments on both synthetic and real data.
De-Lei Chen, Lei Zhang 0021, Hua Huang 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Extracting Small Flying Airplane With Spatially Accurate and Temporally Consistent Foreground Modeling
De-Lei Chen, Lei Zhang 0021, Hua Huang 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 IMU-Assisted Online Video Background Identification
abstract
Distinguishing between dynamic foreground objects and a mostly static background is a fundamental problem in many computer vision and computer graphics tasks. This paper presents a novel online video background identification method with the assistance of inertial measurement unit (IMU). Based on the fact that the background motion of a video essentially reflects the 3D camera motion, we leverage IMU data to realize a robust camera motion estimation for identifying background feature points by only investigating a few historical frames. We observe that the displacement of the 2D projection of a scene point caused by camera rotation is depth-invariant, and the rotation estimation by using IMU data can be quite accurate. We thus propose to analyze 2D feature points by decomposing the 2D motion into two components: rotation projection and translation projection. In our method, after establishing the 3D camera rotations, we generate the depth-relevant 2D feature point movement induced by the camera 3D translation. Then, by examining the disparity between inter-frame offset and the projection of estimated 3D camera motion, we can identify the background feature points. In the experiments, our online method is able to run at 30FPS with only 1 frame latency and outperforms state-of-the-art background identification and other relevant methods. Our method directly leads to a better camera motion estimation, which is beneficial to many applications like online video stabilization, SLAM, image stitching, etc.
Jianxiang Rong, Lei Zhang 0021, Hua Huang 0001
IEEE Trans. Image Process.2
2021 Learning Tensor Low-Rank Prior for Hyperspectral Image Reconstruction
abstract
Snapshot hyperspectral imaging has been developed to capture the spectral information of dynamic scenes. In this paper, we propose a deep neural network by learning the tensor low-rank prior of hyperspectral images (HSI) in the feature domain to promote the reconstruction quality. Our method is inspired by the canonical-polyadic (CP) decomposition theory, where a low-rank tensor can be expressed as a weight summation of several rank-1 component tensors. Specifically, we first learn the tensor low-rank prior of the image features with two steps: (a) we generate rank-1 tensors with discriminative components to collect the contextual information from both spatial and channel dimensions of the image features; (b) we aggregate those rank-1 tensors into a low-rank tensor as a 3D attention map to exploit the global correlation and refine the image features. Then, we integrate the learned tensor low-rank prior into an iterative optimization algorithm to obtain an end-to-end HSI reconstruction. Experiments on both synthetic and real data demonstrate the superiority of our method.
Lizhi Wang 0001, Lei Zhang 0021, Hua Huang 0001
CVPR3
2021 Loop Closure Detection by Using Global and Local Features With Photometric and Viewpoint Invariance
abstract
Loop closure detection plays an important role in many Simultaneous Localization and Mapping (SLAM) systems, while the main challenge lies in the photometric and viewpoint variance. This paper presents a novel loop closure detection algorithm that is more robust to the variance by using both global and local features. Specifically, the global feature with the consolidation of photometric and viewpoint invariance is learned by a Siamese Network from the intensity, depth, gradient and normal vectors distribution. The local feature with rotation invariance is based on the histogram of relative pixel intensity and geometric information like curvature and coplanarity. Then, these two types of features are jointly leveraged for the robust detection of loop closures. The extensive experiments have been conducted on the publicly available RGB-D benchmark datasets like TUM and KITTI. The results demonstrate that our algorithm can effectively address challenging scenarios with large photometric and viewpoint variance, which outperforms other state-of-the-art methods.
Mingfei Yu, Lei Zhang 0021, Wufan Wang, Hua Huang 0001
IEEE Trans. Image Process.2
2020 Snapshot Hyperspectral Imaging Based on Weighted High-order Singular Value Regularization
abstract
Snapshot hyperspectral imaging can capture the 3D hyperspectral image (HSI) with a single 2D measurement and has attracted increasing attention recently. Recovering the underlying HSI from the compressive measurement is an ill-posed problem and exploiting the image prior is essential for solving this ill-posed problem. However, existing reconstruction methods always start from modeling image prior with the 1D vector or 2D matrix and cannot fully exploit the structurally spectral-spatial nature in 3D HSI, thus leading to a poor fidelity. In this paper, we propose an effective high-order tensor optimization based method to boost the reconstruction fidelity for snapshot hyperspectral imaging. We first build high-order tensors by exploiting the spatial-spectral correlation in HSI. Then, we propose a weight high-order singular value regularization (WHOSVR) based low-rank tensor recovery model to characterize the structure prior of HSI. By integrating the structure prior in WHOSVR with the system imaging process, we develop an optimization framework for HSI reconstruction, which is finally solved via the alternating minimization algorithm. Extensive experiments implemented on two representative systems demonstrate that our method outperforms state-of-the-art methods.
Niankai Cheng, Hua Huang 0001, Lei Zhang 0021, Lizhi Wang 0001
ICPR3
2020 Embedding shared low-rank and feature correlation for multi-view data analysis
abstract
The diversity of multimedia data in the real-world usually forms multi-view features. How to explore the structure information and correlations among multi-view features is still a challenging problem. In this paper, we propose a novel multi-view subspace learning method, named embedding shared low-rank and feature correlation (ESLRFC), for multi-view data analysis. First, in the embedding subspace, we propose a robust low-rank model on each feature set and enforce a shared low-rank constraint to characterize the common structure information of multiple feature data. Second, we develop an enhanced correlation analysis in the embedding subspace for simultaneously removing the redundancy of each feature set and exploring the correlations of multiple feature data. Finally, we incorporate the low-rank model and the correlation analysis into a unified framework. The shared low-rank constraint not only depicts the data distribution consistency among multiple feature data, but also assists robust subspace learning. Experimental results on recognition tasks demonstrate the superior performance and noise robustness of the proposed method.
Zhan Wang 0007, Lizhi Wang 0001, Lei Zhang 0021, Hua Huang 0001
ICPR3
2020 Small Target Detection in Infrared Videos Based on Spatio-Temporal Tensor Model
abstract
Existing methods of the small target detection from infrared videos are not effective with the complex background. It is mainly caused by: 1) the interference of strong edges and the similarity with other nontarget objects and 2) the lack of the context information of both the background and the target in a spatio-temporal domain. By considering these two points, we propose to slide a window in a single frame and form a spatio-temporal cube with the current frame patch and other frame patches in the spatio-temporal domain. Then, we establish a spatio-temporal tensor model based on these patches. According to the sparse prior of the target and the local correlation of the background, the separation of the target and the background can be cast as a low rank and sparse tensor decomposition problem. The target is obtained from the sparse tensor by the tensor decomposition. The experiments show that our method gains better detection performance in infrared videos with the complex background by making full use of the spatio-temporal context information.
Hong-Kang Liu, Lei Zhang 0021, Hua Huang 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 An Effective Network with ConvLSTM for Low-Light Image Enhancement
Yixi Xiang, Ying Fu 0001, Lei Zhang 0021, Hua Huang 0001
PRCV (2)3
2019 Encoding Shaky Videos by Integrating Efficient Video Stabilization
abstract
This paper presents a novel video coding method by integrating video stabilization for shaky videos. By reusing the stabilized motion of feature points and geometric transformations, a better predictor can be established to replace the original motion vectors in the motion estimation stage of video coding. Then, these motion vectors are optimized based on statistics of the residuals between the stabilization predictor and the standard one to improve the efficiency of motion search. As a result, our method brings much less computational cost to motion estimation than encoding the stabilized frames separately (e.g., 24% for the enhanced predictive zonal search algorithm, 17% for the unsymmetrical-cross multi-hexagon-grid search algorithm), while the Bjontegaard distortion (BD) bit rate and BDpeak signal-to-noise ratio still have the comparable performance with multiple quantization parameters. Specially, the implementation of our integrative system based on ×264 is very fast and of low latency, by which full HD videos can be simultaneously stabilized and encoded with more than 30 frames per second in the fastest mode, even on the mobile platform. The experiments on a variety of shaky videos demonstrate the potential of our method in terms of effectiveness and efficiency.
Hua Huang 0001, Xiao-Xiang Wei, Lei Zhang 0021
IEEE Trans. Circuits Syst. Video Technol.3
2019 Intrinsic Motion Stability Assessment for Video Stabilization
abstract
This paper presents a novel algorithm for assessing the motion stability of a video after stabilization. The assessment works in a non-reference manner that directly measures the intrinsic smoothness of the video motion path. Specifically, the motion path is cast as a curve embedded in the Lie group of homographies, and its smoothness is mathematically characterized by the intrinsic geodesic curvature. A bundle of paths are adopted to handle spatially variant motions through the frames. Then, we compute the weighted curvature for a holistic assessment on the motion stability. Other factors related to video stabilization, e.g., distortion and cropping, are also investigated as supplement. We collect 160 shaky video clips and their stabilized results for verification, and the experimental evidence shows the effectiveness of our algorithm in good correlation with human subjective judgements.
Lei Zhang 0021, Qing-Zhuo Zheng, Hua Huang 0001
IEEE Trans. Vis. Comput. Graph.1
2018 Full-Reference Stability Assessment of Digital Video Stabilization Based on Riemannian Metric
abstract
Assessing the quality of the motion stability is important to evaluating the performance of video stabilization algorithms. This paper presents a novel quality assessment scheme for the video motion stability in a full-reference (FR) manner. Given ideally stable videos and their corresponding shaky videos, our method measures the geodesic distance between motion paths of the stable and the stabilized videos. Due to the use of the Riemannian metric defined on the manifold of spatial transformations, our method enables the intrinsic and faithful measurement on pairwise motion disparities. To facilitate the FR assessment, a data set of stable and shaky videos is constructed by directly capturing realistic stable/shaky videos with a customized device. Then, digital video stabilization algorithms can be run on shaky videos to obtain the stabilized sequence of frames, whereupon their performances are evaluated by using our stability assessment. The experiments demonstrate that our stability assessment gains good concordance with the subjective assessment.
Lei Zhang 0021, Qing-Zhuo Zheng, Hong-Kang Liu, Hua Huang 0001
IEEE Trans. Image Process.1
2017 A Global Approach to Fast Video Stabilization
abstract
This paper presents a novel formulation of video stabilization by directly solving for optimal image warps toward stabilized sequence. With the estimated shaky motion via long or short feature trajectories, our approach encodes another two steps, motion compensation and image warping, into a single global optimization process, rather than operating as two individual steps. This process is done only with positions of embedded mesh vertices as common variables. Spatial and temporal coherence is therein reformulated with similarity-invariant representation of motion trajectories and intra- (and inter-) frame consistency of similar transformations with respect to mesh vertices. Such a one-shot formulation converts video stabilization into a quadratic energy minimization problem defined for image warps, and thus can be efficiently resolved by using a robust solver for sparse linear systems. Experimental results demonstrate the flexibility and efficiency of our approach in producing visually plausible stabilization effects on a variety of videos.
Lei Zhang 0021, Qian-Kun Xu, Hua Huang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Bundled Kernels for Nonuniform Blind Video Deblurring
abstract
We present a novel blind video deblurring approach by estimating a bundle of kernels and applying the residual deconvolution. Our approach adopts multiple kernels to represent spatially varying motion blur, and thus can cope with nonuniform video deblurring. For each blurred frame, we build a warping-based, space-variant motion blur model based on a bundle of homographies in between its adjacent frames. Then, the nearest sharp frame is employed to form an unblurred-blurred pair for solving the motion model and obtain a bundle of kernels at the blurred frame. Finally, we apply the deconvolution on the residual between the warped unblurred frame and blurred frame with the kernels. The blur kernel estimation and residual deconvolution are iteratively performed toward the deblurred frame, as well as significantly reducing artifacts such as ringings. Experiments show that our approach can efficiently remove the nonuniform video blurring, and achieves better deblurring results than some state-of-the-art methods.
Lei Zhang 0021, Hua Huang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Geodesic Video Stabilization in Transformation Space
abstract
We present a novel formulation of video stabilization in the space of geometric transformations. With the setting of the Riemannian metric, the optimized smooth path is cast as the geodesics on the Lie group embedded in transformation space. While solving the geodesics has a closed-form expression in a certain space, path smoothing can be easily implemented by using geometric interpolation, rather than optimizing any space-time energy function. Specially, by using the geodesic solution in the space of rigid transformations, our approach even gains speedup 10× faster than state-of-the-art methods for path smoothing and motion compensation, and guarantees no extra distortion drawn into the stabilized frames. The experiments demonstrate the efficiency and effectiveness of our algorithm on stabilizing a variety of shaky videos.
Lei Zhang 0021, Xiao-Quan Chen, Xin-Yi Kong, Hua Huang 0001
IEEE Trans. Image Process.1
2015 Efficient Variational Light Field View Synthesis For Making Stereoscopic 3D Images
abstract
We present a novel approach for making stereoscopic images by variational view synthesis on the multi-perspective light field. With the intended disparities as constraints, we specialize the generative variational model by incorporating per-pixel viewpoint assignment to synthesize the stereo pair. Also, we improve the variational solution by use of explicit weighted average on the light field. Our algorithm is able to handle arbitrary disparity remapping, thus enabling more flexible disparity control for the desired stereoscopic effect. The experiments demonstrate the effectiveness and efficiency for making the stereoscopic 3D images based on the light field.
Lei Zhang 0021, Yuhang Zhang 0005, Hua Huang 0001
Comput. Graph. Forum1
2015 Guided Adaptive Image Smoothing via Directional Anisotropic Structure Measurement
abstract
Image smoothing prefers a good metric to identify dominant structures from textures adaptive of intensity contrast. In this paper, we drop on a novel directional anisotropic structure measurement (DASM) toward adaptive image smoothing. With observations on psychological perception regarding anisotropy, non-periodicity and local directionality, DASM can well characterize structures and textures independent on their contrast scales. By using such measurement as constraint, we design a guided adaptive image smoothing scheme by improving extrema localization and envelopes construction in a structure-aware manner. Our approach can well suppresses the staircase-like artifacts and blur of structures that appear in previous methods, which better suits structure-preserving image smoothing task. The algorithm is performed on a space-filling curve as the reduced domain, so it is very fast and much easy to implement in practice. We make comprehensive comparisons with previous state-of-the-art methods for a variety of applications. Experimental results demonstrate the merit using our DASM as metric to identify structures, and the effectiveness and efficiency of our adaptive image smoothing approach to produce commendable results.
Hua Huang 0001, Lei Zhang 0021
IEEE Trans. Vis. Comput. Graph.3
2014 Efficient Structure-Aware Image Smoothingby Local Extrema on Space-Filling Curve
abstract
This paper presents a novel image smoothing approach using a space-filling curve as the reduced domain to perform separation of edges and details. This structure-aware smoothing effect is achieved by modulating local extrema after empirical mode decomposition; it is highly effective and efficient since it is implemented on a one-dimensional curve instead of a two-dimensional image grid. To overcome edge staircase-like artifacts caused by a neighborhood deficiency in domain reduction, we next use a joint contrast-based filter to consolidate edge structures in image smoothing. The adoption of dimensional reduction makes our smoothing approach distinct for two reasons. First, overall structure-awareness is improved as more extrema are exploited to locate the salient edges and details. Second, envelope computation for local extrema is made much fast by using explicit interpolants on the curve. Moreover, our approach is simple and very easy to implement in practice. Experimental results demonstrate the merit of our approach, which outperforms previous state-of-the-art methods, for a variety of image processing tasks.
Hua Huang 0001, Lei Zhang 0021
IEEE Trans. Vis. Comput. Graph.3
2014 VideoGraph: a non-linear video representation for efficient exploration
Lei Zhang 0021, Qian-Kun Xu, Lei-Zheng Nie, Hua Huang 0001
Vis. Comput.1
2013 Multiplane Video Stabilization
abstract
Abstract This paper presents a novel video stabilization approach by leveraging the multiple planes structure of video scene to stabilize inter‐frame motion. As opposed to previous stabilization procedure operating in a single plane, our approach primarily deals with multiplane videos and builds their multiple planes structure for performing stabilization in respective planes. Hence, a robust plane detection scheme is devised to detect multiple planes by classifying feature trajectories according to reprojection errors generated by plane induced homographies. Then, an improved planar stabilization technique is applied by conforming to the compensated homography in each plane. Finally, multiple stabilized planes are coherently fused by content‐preserving image warps to obtain the output stabilized frames. Our approach does not need any stereo reconstruction, yet is able to produce commendable results due to awareness of multiple planes structure in the stabilization. Experimental results demonstrate the effectiveness and efficiency of our approach to robust stabilization on multiplane videos.
Zhongqiang Wang, Lei Zhang 0021, Hua Huang 0001
Comput. Graph. Forum2
2012 Hierarchical Narrative Collage For Digital Photo Album
abstract
Abstract Collage can provide a summary form on the collection of photos in an album. In this paper, we introduce a novel approach to constructing photo collage in the hierarchical narrative manner. As opposed to previous methods focusing on spatial coherence in the collage layout, our narrative collage arranges the photos according to the basic narrative elements from literary writings, i.e., character, setting and plot. Face, time and place attributes are exploited to embody those narrative elements in the collage. Then, photos are organized into the hierarchical structure for the multi‐level details in the events recorded by the album. Such hierarchical narrative collage can present a visual overview in the chronological order on what happened in the album. Experimental results show that our approach offers a better summarization to browse on the photo album content than previous ones.
Lei Zhang 0021, Hua Huang 0001
Comput. Graph. Forum1
2012 EXCOL: An EXtract-and-COmplete Layering Approach to Cartoon Animation Reusing
abstract
We introduce the EXtract-and-COmplete Layering method (EXCOL)--a novel cartoon animation processing technique to convert a traditional animated cartoon video into multiple semantically meaningful layers. Our technique is inspired by vision-based layering techniques but focuses on shape cues in both the extraction and completion steps to reflect the unique characteristics of cartoon animation. For layer extraction, we define a novel similarity measure incorporating both shape and color of automatically segmented regions within individual frames and propagate a small set of user-specified layer labels among similar regions across frames. By clustering regions with the same labels, each frame is appropriately partitioned into different layers, with each layer containing semantically meaningful content. Then, a warping-based approach is used to fill missing parts caused by occlusion within the extracted layers to achieve a complete representation. EXCOL provides a flexible way to effectively reuse traditional cartoon animations with only a small amount of user interaction. It is demonstrated that our EXCOL method is effective and robust, and the layered representation benefits a variety of applications in cartoon animation processing.
Lei Zhang 0021, Hua Huang 0001, Hongbo Fu 0001
IEEE Trans. Vis. Comput. Graph.1
2012 Web-image driven best views of 3D shapes
Lei Zhang 0021, Hua Huang 0001
Vis. Comput.2
2011 RepSnapping: Efficient Image Cutout for Repeated Scene Elements
abstract
Abstract Repeated scene elements are copious and ubiquitous in natural images. Cutout of those repeated elements usually involves tedious and laborious user interaction by previous image segmentation methods. In this paper, we present RepSnapping, a novel method oriented to cutout of repeated scene elements with much less user interaction. By exploring inherent similarity between repeated elements, a new optimization model is introduced to thread correlated elements in the segmentation procedure. The model proposed here enables efficient solution using max‐flow/min cut on an extended graph. Experiments indicate thatRepSnappingfacilitates cutout of repeated elements better than the state‐of‐the‐art interactive image segmentation and repetition detection methods.
Hua Huang 0001, Lei Zhang 0021
Comput. Graph. Forum2
2011 Arcimboldo-like collage using internet images
abstract
Collage is a composite artwork made from assemblage of different material forms. In this work, we present a novel approach for creating a fantastic collage artform, namely Arcimboldo-like collage, which represents an input image with multiple thematically-related cutouts from the filtered Internet images. Due to the massive data of Internet images, competent image cutouts can almost always be discovered to match the segmented components of the input image. The selected cutouts are purposefully arranged such that as a whole assembly, they can represent the input image with disguise in both shape and color; but separately, individual cutout is still recognizable as its own being. Experimental results and user study show that our algorithm can effectively produce the entertaining Arcimboldo-like collages.
Hua Huang 0001, Lei Zhang 0021
ACM Trans. Graph.2
2010 Mesh reconstruction by meshless denoising and parameterization
Lei Zhang 0021, Ligang Liu 0001, Craig Gotsman, Hua Huang 0001
Comput. Graph.1
2010 Video Painting via Motion Layer Manipulation
abstract
Abstract Temporal coherence is an important problem in Non‐Photorealistic Rendering for videos. In this paper, we present a novel approach to enhance temporal coherence in video painting. Instead of painting on video frame, our approach first partitions the video into multiple motion layers, and then places the brush strokes on the layers to generate the painted imagery. The extracted motion layers consist of one background layer and several object layers in each frame. Then, background layers from all the frames are aligned into a panoramic image, on which brush strokes are placed to paint the background in one‐shot. The strokes used to paint object layers are propagated frame by frame using smooth transformations defined by thin plate splines. Once the background and object layers are painted, they are projected back to each frame and blent to form the final painting results. Thanks to painting a single image, our approach can completely eliminate the flickering in background, and temporal coherence on object layers is also significantly enhanced due to the smooth transformation over frames. Additionally, by controlling the painting strokes on different layers, our approach is easy to generate painted video with multi‐style. Experimental results show that our approach is both robust and efficient to generate plausible video painting.
Hua Huang 0001, Lei Zhang 0021, TianNan Fu
Comput. Graph. Forum2
2010 An as-rigid-as-possible approach to sensor network localization
abstract
We present a novel approach to localization of sensors in a network given a subset of noisy inter-sensor distances. The algorithm is based on “stitching” together local structures by solving an optimization problem requiring the structures to fit together in an “As-Rigid-As-Possible” manner, hence the name ARAP. The local structures consist of reference “patches” and reference triangles, both obtained from inter-sensor distances. We elaborate on the relationship between the ARAP algorithm and other state-of-the-art algorithms, and provide experimental results demonstrating that ARAP is significantly less sensitive to sparse connectivity and measurement noise. We also show how ARAP may be distributed.
Lei Zhang 0021, Ligang Liu 0001, Craig Gotsman, Steven J. Gortler
ACM Trans. Sens. Networks1
2009 Fast approach for computing roots of polynomials using cubic clipping
Ligang Liu 0001, Lei Zhang 0021, Guojin Wang
Comput. Aided Geom. Des.2
2008 A Local/Global Approach to Mesh Parameterization
abstract
Abstract We present a novel approach to parameterize a mesh with disk topology to the plane in a shape‐preserving manner. Our key contribution is a local/global algorithm, which combines a local mapping of each 3D triangle to the plane, using transformations taken from a restricted set, with a global “stitch” operation of all triangles, involving a sparse linear system. The local transformations can be taken from a variety of families, e.g. similarities or rotations, generating different types of parameterizations. In the first case, the parameterization tries to force each 2D triangle to be an as‐similar‐as‐possible version of its 3D counterpart. This is shown to yield results identical to those of the LSCM algorithm. In the second case, the parameterization tries to force each 2D triangle to be an as‐rigid‐as‐possible version of its 3D counterpart. This approach preserves shape as much as possible. It is simple, effective, and fast, due to pre‐factoring of the linear system involved in the global phase. Experimental results show that our approach provides almost isometric parameterizations and obtains more shape‐preserving results than other state‐of‐the‐art approaches. We present also a more general “hybrid” parameterization model which provides a continuous spectrum of possibilities, controlled by a single parameter. The two cases described above lie at the two ends of the spectrum. We generalize our local/global algorithm to compute these parameterizations. The local phase may also be accelerated by parallelizing the independent computations per triangle.
Ligang Liu 0001, Lei Zhang 0021, Craig Gotsman, Steven J. Gortler
Comput. Graph. Forum2
2006 Manifold Parameterization
Lei Zhang 0021, Ligang Liu 0001, Zhongping Ji, Guojin Wang
Computer Graphics International1
2005 Study on the performance assessment of green supply chain
abstract
Supply chain management (SCM) is an effective way to enhance enterprises' adaptive ability and viability in the market, by combining the upstream enterprises and downstream enterprises to participate in the market competition together. With the increasing environmental consciousness of people, how to reduce the environmental impacts of the enterprises' actions has become the common concern of both customers and enterprises. Green supply chain management (GSCM) is an effective way to reduce the environmental impacts of products throughout their life cycles. The performance assessment of green supply chain (GSC) is an important part of GSCM, and the greenness and closed-loop of GSC make its performance assessment more complex. The performance assessment of GSC is discussed and the performance assessment index system is introduced. Furthermore, the performance assessment model of GSC is established using the fuzzy assessment method, and it is explained with a case study.
Shuwang Wang, Lei Zhang 0021, Guangfu Liu
SMC2