Haoqian Wang

dblp:57/7040 · DBLP profile ↗
← Back
132ranked-venue papers
12as first author
90since 2021 · last 2026
0000-0003-2792-8469ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 100 · 11 first-author · 62 since 2021Artificial intelligence and machine learning · 69 · 3 first-author · 59 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Theory of computation · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation
abstract
Egocentric human pose estimation (HPE) plays a crucial role in immersive applications such as virtual and augmented reality. However, existing methods relying on either visual or sparse inertial data alone often suffer from occlusion or ill-posed problems. In this work, we propose SAME, a novel spatial-aware multimodal fusion framework combining the complementary signals from the stereo images and sparse IMUs for accurate and robust egocentric HPE. It adopts a two-stage network based on a dual coordinate frame to mitigate the coordinate inconsistencies among the stereo cameras and the IMUs. In the first stage, the IMU signals are transformed into the local frame and iteratively fused with the stereo images for estimating 3D poses in the local frame. In the second stage, the local poses are transformed into the global frame with the 6DOF head poses provided by the head-mounted display's (HMD) SLAM algorithm and then temporally aggregated via a temporal Transformer network. Meanwhile, to achieve geometric and semantic alignment among multi-modal features, we present a depth-guided spatial-aware deformable stereo attention network and a modality-aware Transformer decoder for cross-view and cross-modal feature fusion. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on the public EMHI multi-modal egocentric pose estimation benchmark.
Yurong Fu, Yiqiang Feng, Haoqian Wang
AAAI6
2026 Language-Guided and Motion-Aware Gait Representation for Generalizable Recognition
abstract
Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, existing methods typically rely on complex architectures to directly extract features from images and apply pooling operations to obtain sequence-level representations. Such designs often lead to overfitting on static noise (e.g., clothing), while failing to effectively capture dynamic motion regions, such as the arms and legs. This bottleneck is particularly challenging in the presence of intra-class variation, where gait features of the same individual under different environmental conditions are significantly distant in the feature space. To address the above challenges, we present a Language-guided and Motion-aware gait recognition framework, named LMGait. To the best of our knowledge, LMGait is the first method to introduce natural language descriptions as explicit semantic priors into the gait recognition task. In particular, we utilize designed gait-related language cues to capture key motion features in gait sequences. To improve cross-modal alignment, we propose the Motion Awareness Module (MAM), which refines the language features by adaptively adjusting various levels of semantic information to ensure better alignment with the visual representations. Furthermore, we introduce the Motion Temporal Capture Module (MTCM) to enhance the discriminative capability of gait features and improve the model’s motion tracking ability. We conducted extensive experiments across multiple datasets, and the results demonstrate the significant advantages of our proposed network. Specifically, our model achieved accuracies of 88.5%, 97.1%, and 97.5% on the CCPG, SUSTech1K, and CASIAB* datasets, respectively, achieving state-of-the-art performance.
Zhengxian Wu, Chuanrui Zhang, Shenao Jiang, Hangrui Xu, Zirui Liao, Luyuan Zhang, Huaqiu Li, Peng Jiao, Haoqian Wang
AAAI9
2026 Wonder3D++: Cross-Domain Diffusion for High-Fidelity 3D Generation From a Single Image
abstract
In this work, we introduce Wonder3D++, a novel method for efficiently generating high-fidelity textured meshes from single-view images. Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works directly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of single-view reconstruction tasks, we propose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure the consistency of generation, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a cascaded 3D mesh extraction algorithm that drives high-quality surfaces from the multi-view 2D representations in only about 3 minute in a coarse-to-fine manner. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and good efficiency compared to prior works.
Xiaoxiao Long, Zhiyang Dou, Cheng Lin 0001, Yuan Liu 0025, Qingsong Yan, Yuexin Ma, Haoqian Wang, Zhiqiang Wu 0001, Wei Yin 0006
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Spatiotemporal Blind-Spot Network with Calibrated Flow Alignment for Self-Supervised Video Denoising
abstract
Self-supervised video denoising aims to remove noise from videos without relying on ground truth data, leveraging the video itself to recover clean frames. Existing methods often rely on simplistic feature stacking or apply optical flow without thorough analysis. This results in suboptimal utilization of both inter-frame and intra-frame information, and it also neglects the potential of optical flow alignment under self-supervised conditions, leading to biased and insufficient denoising outcomes. To this end, we first explore the practicality of optical flow in the self-supervised setting and introduce a SpatioTemporal Blind-spot Network (STBN) for global frame feature utilization. In the temporal domain, we utilize bidirectional blind-spot feature propagation through the proposed blind-spot alignment block to ensure accurate temporal alignment and effectively capture long-range dependencies. In the spatial domain, we introduce the spatial receptive field expansion module, which enhances the receptive field and improves global perception capabilities. Additionally, to reduce the sensitivity of optical flow estimation to noise, we propose an unsupervised optical flow distillation mechanism that refines fine-grained inter-frame interactions during optical flow alignment. Our method demonstrates superior performance across both synthetic and real-world video denoising datasets.
Xiaowan Hu, Huaqiu Li, Haoqian Wang
AAAI6
2025 Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single Image Denoising
abstract
Many studies have concentrated on constructing supervised models utilizing paired datasets for image denoising, which proves to be expensive and time-consuming. Current self-supervised and unsupervised approaches typically rely on blind-spot networks or sub-image pairs sampling, resulting in pixel information loss and destruction of detailed structural information, thereby significantly constraining the efficacy of such methods. In this paper, we introduce Prompt-SID, a prompt-learning-based single image denoising framework that emphasizes the preservation of structural details. This approach is trained in a self-supervised manner using downsampled image pairs. It captures original-scale image information through structural encoding and integrates this prompt into the denoiser. To achieve this, we propose a structural representation generation model based on the latent diffusion process and design a structural attention module within the transformer-based denoiser architecture to decode the prompt. Additionally, we introduce a scale replay training mechanism, which effectively mitigates the scale gap from images of different resolutions. We conduct comprehensive experiments on synthetic, real-world, and fluorescence imaging datasets, showcasing the remarkable effectiveness of Prompt-SID.
Huaqiu Li, Xiaowan Hu, Haoqian Wang
AAAI6
2025 MVReward: Better Aligning and Evaluating Multi-View Diffusion Models with Human Preferences
abstract
Recent years have witnessed remarkable progress in 3D content generation. However, corresponding evaluation methods struggle to keep pace. Automatic approaches have proven challenging to align with human preferences, and the mixed comparison of text- and image-driven methods often leads to unfair evaluations. In this paper, we present a comprehensive framework to better align and evaluate multi-view diffusion models with human preferences. To begin with, we first collect and filter a standardized image prompt set from DALL·E and Objaverse, which we then use to generate multi-view assets with several multi-view diffusion models. Through a systematic ranking pipeline on these assets, we obtain a human annotation dataset with 16k expert pairwise comparisons and train a reward model, coined MVReward, to effectively encode human preferences. With MVReward, image-driven 3D methods can be evaluated against each other in a more fair and transparent manner. Building on this, we further propose Multi-View Preference Learning (MVP), a plug-and-play multi-view diffusion tuning strategy. Extensive experiments demonstrate that MVReward can serve as a reliable metric and MVP consistently enhances the alignment of multi-view diffusion models with human preferences.
Jun Meng, Haoqian Wang
AAAI6
2025 TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with Transformers
abstract
Compared with previous 3D reconstruction methods like Nerf, recent Generalizable 3D Gaussian Splatting (G-3DGS) methods demonstrate impressive efficiency even in the sparse-view setting. However, the promising reconstruction performance of existing G-3DGS methods relies heavily on accurate multi-view feature matching, which is quite challenging. Especially for the scenes that have many non-overlapping areas between various views and contain numerous similar regions, the matching performance of existing methods is poor and the reconstruction precision is limited. To address this problem, we develop a strategy that utilizes a predicted depth confidence map to guide accurate local feature matching. In addition, we propose to utilize the knowledge of existing monocular depth estimation models as prior to boost the depth estimation precision in non-overlapping areas between views. Combining the proposed strategies, we present a novel G-3DGS method named TranSplat, which obtains the best performance on both the RealEstate10K and ACID benchmarks while maintaining competitive speed and presenting strong cross-dataset generalization ability.
Chuanrui Zhang, Yingshuang Zou, Zhuoling Li, Minmin Yi, Haoqian Wang
AAAI5
2025 MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
abstract
Masked Image Modeling (MIM) with Vector Quantization (VQ) has achieved great success in both self-supervised pre-training and image generation. However, most existing methods struggle to address the trade-off in shared latent space for generation quality vs. representation learning and efficiency. To push the limits of this paradigm, we propose MergeVQ, which incorporates token merging techniques into VQ-based generative models to bridge the gap between image generation and visual representation learning in a unified architecture. During pre-training, MergeVQ decouples top-k semantics from latent space with the token merge module after self-attention blocks in the encoder for subsequent Look-up Free Quantization (LFQ) and global alignment and recovers their fine-grained details through cross-attention in the decoder for reconstruction. As for second-stage generation, we introduce MergeAR, which performs KV Cache compression for efficient raster-order prediction. Extensive experiments on ImageNet verify that MergeVQ as an AR generative model achieves competitive performance in both visual representation learning and image generation tasks while maintaining favorable token efficiency and inference speed. Code and model will be available at https://apexgen-x.github.io/MergeVQ.
Siyuan Li 0002, Luyuan Zhang, Zedong Wang, Juanxi Tian, Cheng Tan 0012, Zicheng Liu 0006, Chang Yu 0001, Qingsong Xie, Haonan Lu, Haoqian Wang, Zhen Lei 0001
CVPR10
2025 Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture
abstract
Existing methods for capturing multi-person holistic human motions from a monocular video usually involve integrating the detector, the tracker, and the human pose & shape estimator into a cascaded system. Differently, we develop a one-stage multi-person holistic human motion capture system, which 1) employs only one network, enabling significant benefits from the end-to-end training on a large-scale dataset; 2) enables performance improving of the tracking module during training, avoiding being limited by a pre-trained tracker; 3) captures the motions of all individuals within a single shot, rather than tracking and estimating each person sequentially. In this system, each query within a temporal cross-attention module is responsible for the long motion of a specific individual, implicitly aggregating individual-specific information throughout the entire video. To further boost the proposed system from end-to-end training, we also construct a synthetic human video dataset, with multi-person and whole-body annotations. Extensive experiments across different datasets demonstrate both the efficacy and the efficiency of both the proposed method and the dataset. Codes are avaiable at https://github.com/KenkunLiu/MaQ.
Kenkun Liu, Yurong Fu, Weihao Yuan 0001, Peihao Li 0003, Xiaodong Gu 0004, Lingteng Qiu, Haoqian Wang, Zilong Dong, Xiaoguang Han 0001
CVPR8
2025 HRAvatar: High-Quality and Relightable Gaussian Head Avatar
abstract
Reconstructing animatable and high-quality 3D head avatars from monocular videos, especially with realistic relighting, is a valuable task. However, the limited information from single-view input, combined with the complex head poses and facial movements, makes this challenging. Previous methods achieve real-time performance by combining 3D Gaussian Splatting with a parametric head model, but the resulting head quality suffers from inaccurate face tracking and limited expressiveness of the deformation model. These methods also fail to produce realistic effects under novel lighting conditions. To address these issues, we propose HRAvatar, a 3DGS-based method that reconstructs high-fidelity, relightable 3D head avatars. HRA-vatar reduces tracking errors through end-to-end optimization and better captures individual facial deformations using learnable blendshapes and learnable linear blend skinning. Additionally, it decomposes head appearance into several physical properties and incorporates physically-based shading to account for environmental lighting. Extensive experiments demonstrate that HRAvatar not only reconstructs superior-quality heads but also achieves realistic visual effects under varying lighting conditions. Video results and code are available at the project page.
Dongbin Zhang, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Kangjie Chen, Minghan Qin, Yu Li 0003, Haoqian Wang
CVPR8
2025 HumanMM: Global Human Motion Recovery from Multi-shot Videos
abstract
In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understanding, but are of great challenge to be recovered due to abrupt shot transitions, partial occlusions, and dynamic backgrounds presented in such videos. Existing methods primarily focus on single-shot videos, where continuity is maintained within a single camera view, or simplify multi-shot alignment in camera space only. In this work, we tackle the challenges by integrating an enhanced camera pose estimation with Human Motion Recovery (HMR) by incorporating a shot transition detector and a robust alignment module for accurate pose and orientation continuity across shots. By leveraging a custom motion integrator, we effectively mitigate the problem of foot sliding and ensure temporal consistency in human pose. Extensive evaluations on our created multi-shot dataset from public 3D human datasets demonstrate the robustness of our method in reconstructing realistic human motion in world coordinates.
Guanlin Wu, Zhuokai Zhao, Xiaoke Jiang, Zhuoheng Li, Hao (Frank) Yang, Haoqian Wang, Lei Zhang 0001
CVPR10
2025 DFNeRF: Disentangled Facial Neural Radiance Fields for Text-based Editing of Free-view Talking Head
abstract
In this paper, we propose a text-based approach that can edit the speech content of a free-view talking head based on its transcript. The core of our method is to establish the relationship between phonemes and head attributes. To avoid discontinuities in head pose and facial expressions caused by editing mouth shape. We design the disentangled facial neural radiance fields (DFNeRF) to automatically disentangle these attributes controlled by the learned latent codes. Using the free-view talking head synthesized by DFNeRF as the base corpus, we could re-assemble the latent codes based on the new content using the phoneme search method to produce a seamless edited result. Our phoneme search method with a discriminator could find the best-matched phonemes in the sequence and ensure a smooth transition of mouth shape between the adjacent phonemes. Extensive experiments demonstrate the effectiveness of our method both qualitatively and quantitatively.
Benwang Chen, Qi Zhang 0082, Haoqian Wang
ICASSP5
2025 LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling
abstract
Unified image restoration is a significantly challenging task in low-level vision. Existing methods either make tailored designs for specific tasks, limiting their generalizability across various types of degradation, or rely on training with paired datasets, thereby suffering from closed-set constraints. To address these issues, we propose a novel, dataset-free, and unified approach through recurrent posterior sampling utilizing a pretrained latent diffusion model. Our method incorporates the multimodal understanding model to provide sematic priors for the generative model under a task-blind condition. Furthermore, it utilizes a lightweight module to align the degraded input with the generated preference of the diffusion model, and employs recurrent refinement for posterior sampling. Extensive experiments demonstrate that our method outperforms state-of-the-art methods, validating its effectiveness and robustness. Our code and data are available at https://github.com/AMAP-ML/LD-RPS.
Huaqiu Li, Tongwen Huang, Hailang Huang, Haoqian Wang, Xiangxiang Chu
ICCV5
2025 DPoser-X: Diffusion Model as Robust 3D Whole-Body Human Pose Prior
abstract
We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limitations, we introduce a Diffusion model as body Pose prior (DPoser) and extend it to DPoser-X for expressive whole-body human pose modeling. Our approach unifies various pose-centric tasks as inverse problems, solving them through variational diffusion sampling. To enhance performance on downstream applications, we introduce a novel truncated timestep scheduling method specifically designed for pose data characteristics. We also propose a masked training mechanism that effectively combines whole-body and part-specific datasets, enabling our model to capture interdependencies between body parts while avoiding overfitting to specific actions. Extensive experiments demonstrate DPoser-X's robustness and versatility across multiple benchmarks for body, hand, face, and full-body pose modeling. Our model consistently outperforms state-of-the-art alternatives, establishing a new benchmark for whole-body human pose prior modeling.
Junzhe Lu 0001, Hongkun Dou, Ailing Zeng, Yue Deng 0001, Zhongang Cai, Lei Yang 0059, Yulun Zhang 0001, Haoqian Wang, Ziwei Liu 0002
ICCV10
2025 GUAVA: Generalizable Upper Body 3D Gaussian Avatar
abstract
Reconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual IDs, which is both complex and time-consuming. Furthermore, limited by SMPLX's expressiveness, these methods often focus on body motion but struggle with facial expressions. To address these challenges, we first introduce an expressive human model (EHM) to enhance facial expression capabilities and develop an accurate tracking method. Based on this template model, we propose GUAVA, the first framework for fast animatable upper-body 3D Gaussian avatar reconstruction. We leverage inverse texture mapping and projection sampling techniques to infer Ubody (upper-body) Gaussians from a single image. The rendered images are refined through a neural refiner. Experimental results demonstrate that GUAVA significantly outperforms previous methods in rendering quality and offers significant speed improvements, with reconstruction times in the sub-second range (0.1s), and supports real-time animation and rendering.
Dongbin Zhang, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Minghan Qin, Yu Li 0003, Haoqian Wang
ICCV8
2025 Interpretable Unsupervised Joint Denoising and Enhancement for Real-World low-light Scenarios
abstract
Real-world low-light images often suffer from complex degradations such as local overexposure, low brightness, noise, and uneven illumination. Supervised methods tend to overfit to specific scenarios, while unsupervised methods, though better at generalization, struggle to model these degradations due to the lack of reference images. To address this issue, we propose an interpretable, zero-reference joint denoising and low-light enhancement framework tailored for real-world scenarios. Our method derives a training strategy based on paired sub-images with varying illumination and noise levels, grounded in physical imaging principles and retinex theory. Additionally, we leverage the Discrete Cosine Transform (DCT) to perform frequency domain decomposition in the sRGB space, and introduce an implicit-guided hybrid representation strategy that effectively separates intricate compounded degradations. In the backbone network design, we develop retinal decomposition network guided by implicit degradation representation mechanisms. Extensive experiments demonstrate the superiority of our method. Code will be available at https://github.com/huaqlili/unsupervised-light-enhance-ICLR2025.
Huaqiu Li, Xiaowan Hu, Haoqian Wang
ICLR3
2025 DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic Refinement
abstract
Reconstructing high-quality images under low bitrates conditions presents a challenge, and previous methods have made this task feasible by leveraging the priors of diffusion models. However, the effective exploration of pre-trained latent diffusion models and semantic information integration in image compression tasks still needs further study. To address this issue, we introduce Diffusion-based High Perceptual Fidelity Image Compression with Semantic Refinement (DiffPC), a two-stage image compression framework based on stable diffusion. DiffPC efficiently encodes low-level image information, enabling the highly realistic reconstruction of the original image by leveraging high-level semantic features and the prior knowledge inherent in diffusion models. Specifically, DiffPC utilizes a multi-feature compressor to represent crucial low-level information with minimal bitrates and employs pre-embedding to acquire more robust hybrid semantics, thereby providing additional context for the decoding end. Furthermore, we have devised a control module tailored for image compression tasks, ensuring structural and textural consistency in reconstruction even at low bitrates and preventing decoding collapses induced by condition leakage. Extensive experiments demonstrate that our method achieves state-of-the-art perceptual fidelity and surpasses previous perceptual image compression methods by a significant margin in statistical fidelity.
Yichong Xia, Yimin Zhou 0011, Jinpeng Wang 0002, Baoyi An 0002, Haoqian Wang, Yaowei Wang 0001, Bin Chen 0011
ICLR5
2025 OFF3D:Object-Centric Feature Field for 3D Scene Segmentation
abstract
3D scene segmentation is a fundamental but challenging task for 3D scene understanding in computer vision. However, the majority of existing methods rely heavily on labor-intensive and costly human-annotated 2D or 3D labels. To address this issue, we propose a novel Object-Centric Feature Field for 3d scene segmentation, called OFF3D. Given the machine-generated inconsistent 2D masks, OFF3D can build an implicit object feature field to achieve consistent segmentation for all objects in the 3D scene. Specifically, OFF3D utilizes the continuous object features to represent object properties of each 3D point and proposes a Pseudo-Segment Clustering module to roughly locate objects in the 3D scene at the segment level and design a Confidence-Weighted Contrastive loss to achieve precise object segmentation by pixel-level feature optimization. Extensive experiments demonstrate the effectiveness of our method compared to state-of-the-art methods both quantitatively and qualitatively.
Qinwei Lin, Bing Wang 0013, Jun Xu 0019, Haoqian Wang
ICME5
2025 Leveraging 2D Annotations for Cost-Effective Dynamic Urban Scene Reconstruction
abstract
Accurate reconstruction of dynamic objects in urban scenes remains a challenging task. Existing methods typically rely on complex 3D annotations to identify dynamic objects, which are expensive and difficult to obtain. In this paper, we propose a novel approach to address this challenge by leveraging accessible and effective 2D annotations. We utilize a 2D foundation model to identify dynamic objects within the image and lift this knowledge to 3D space, obtaining the 3D point cloud of dynamic objects. To capture the rigid motion patterns of dynamic objects, we introduce an object-level motion model together with the local aggregation strategy. Additionally, we propose a structural consistency constraint to maintain the consistency of dynamic objects during motion, significantly improving the accuracy of scene reconstruction. Extensive experiments demonstrate that our method achieves comparable reconstruction performance to methods relying on 3D annotations, providing a cost-effective and accurate solution for dynamic urban scene reconstruction.
Chuming Wang, Yingshuang Zou, Haoqian Wang
ICME3
2025 DAGait: Generalized Skeleton-Guided Data Alignment for Gait Recognition
abstract
Gait recognition is emerging as a promising and innovative area within the field of computer vision, widely applied to remote person identification. Although existing gait recognition methods have achieved substantial success in controlled laboratory datasets, their performance often declines significantly when transitioning to wild datasets. We argue that the performance gap can be primarily attributed to the spatio-temporal distribution inconsistencies present in wild datasets, where subjects appear at varying angles, positions, and distances across the frames. To achieve accurate gait recognition in the wild, we propose a skeleton-guided silhouette alignment strategy, which uses prior knowledge of the skeletons to perform affine transformations on the corresponding silhouettes. To the best of our knowledge, this is the first study to explore the impact of data alignment on gait recognition. We conducted extensive experiments across multiple datasets and network architectures, and the results demonstrate the significant advantages of our proposed alignment strategy. Specifically, on the challenging Gait3D dataset, our method achieved an average performance improvement of 7.9% across all evaluated networks. Furthermore, our method achieves substantial improvements on cross-domain datasets, with accuracy improvements of up to 24.0%.Code is available at: https://github.com/DingWu1021/DAGait
Zhengxian Wu, Chuanrui Zhang, Hangrui Xu, Peng Jiao, Haoqian Wang
ICME5
2025 NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation
abstract
3D AI-generated content (AIGC) has made it increasingly accessible for anyone to become a 3D content creator. While recent methods leverage Score Distillation Sampling to distill 3D objects from pretrained image diffusion models, they often suffer from inadequate 3D priors, leading to insufficient multi-view consistency. In this work, we introduce NOVA3D, an innovative single-image-to-3D generation framework. Our key insight lies in leveraging strong 3D priors from a pretrained video diffusion model and integrating geometric information during multi-view video fine-tuning. To facilitate information exchange between color and geometric domains, we propose the Geometry-Temporal Alignment (GTA) attention mechanism, thereby improving generalization and multi-view consistency. Moreover, we introduce the deconflict geometry fusion algorithm, which improves texture fidelity by addressing multi-view inaccuracies and resolving discrepancies in pose alignment. Extensive experiments validate the superiority of NOVA3D over existing baselines.
Peihao Li 0003, Junzhe Lu 0001, Xianglong He, Minghan Qin, Haoqian Wang
ICME8
2025 Measuring and Controlling the Spectral Bias for Self-Supervised Image Denoising
abstract
Current self-supervised denoising methods for paired noisy images typically involve mapping one noisy image through the network to the other noisy image. However, after measuring the spectral bias of such methods using our proposed Image Pair Frequency-Band Similarity, it suffers from two practical limitations. Firstly, the high-frequency structural details in images are not preserved well enough. Secondly, during the process of fitting high frequencies, the network learns high-frequency noise from the mapped noisy images. To address these challenges, we introduce a Spectral Controlling network (SCNet) to optimize self-supervised denoising of paired noisy images. First, we propose a selection strategy to choose frequency band components for noisy images, to accelerate the convergence speed of training. Next, we present a parameter optimization method that restricts the learning ability of convolutional kernels to high-frequency noise using the Lipschitz constant, without changing the network structure. Finally, we introduce the Spectral Separation and low-rank Reconstruction module (SSR module), which separates noise and high-frequency details through frequency domain separation and low-rank space reconstruction, to retain the high-frequency structural details of images. Experiments performed on synthetic and real-world datasets verify the effectiveness of SCNet. The code will be released soon.
Huaqiu Li, Xiaowan Hu, Haoqian Wang
ICME6
2025 Geometrically-Inspired Irregular Expansion Techniques for Graph-based Point Cloud Learning
abstract
Advanced deep learning methodologies have made notable progress in 3D point cloud tasks. Nevertheless, the absence of significant long-range correlation measurement in irregular point cloud entities limits the representation of 3D geometries. Several methodologies have been proposed to address this issue. Still, they fall short of altering the fundamental aggregation paradigm inherent in Graph Convolutional Networks (GCNs), which suffer the exclusive use of the summation operator to encapsulate the information from adjacent nodes, resulting in limited expressiveness and inefficient computation. This paper presents an efficient geometrically irregular expansion on Graph (GIEG) convolution network for point cloud analysis. More specifically, spatial features, formulated with a localized graph representation derived from multiple sequence expansions, are comprehensively exploited by capturing local point semantics while avoiding dense point trappings, which extend the traditional neighborhood with a path-based neighborhood better than native point cloud methods. Thanks to the novel convolution module, the GIEG can extract comprehensive and influential feature semantics for individual points. Extensive experimental results validate the effectiveness of our method against challenging benchmarks.
Qi Zhang 0082, Haoqian Wang, Yuanxi Peng, Teng Li 0011
ICME2
2025 SLGaussian: Fast Language Gaussian Splatting in Sparse Views
abstract
3D semantic field learning is crucial for applications like autonomous navigation, AR/VR, and robotics, where accurate comprehension of 3D scenes from limited viewpoints is essential. Existing methods struggle under sparse view conditions, relying on inefficient per-scene multi-view optimizations, which are impractical for many real-world tasks. To address this, we propose SLGaussian, a feed-forward method for constructing 3D semantic fields from sparse viewpoints, allowing direct inference of 3DGS-based scenes. By ensuring consistent SAM segmentations through video tracking and using low-dimensional indexing for high-dimensional CLIP features, SLGaussian efficiently embeds language information in 3D space, offering a robust solution for accurate 3D scene understanding under sparse view conditions. In experiments on two-view sparse 3D object querying and segmentation in the LERF and 3D-OVS datasets, SLGaussian outperforms existing methods in chosen IoU, Localization Accuracy, and mIoU. Moreover, our model achieves scene inference in under 30 seconds and open-vocabulary querying in just 0.011 seconds per query.
Kangjie Chen, BingQuan Dai, Minghan Qin, Dongbin Zhang, Peihao Li 0003, Yingshuang Zou, Haoqian Wang
ACM Multimedia7
2025 TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
abstract
Diffusion models have recently achieved outstanding results in the field of image super-resolution. These methods typically inject low-resolution (LR) images via ControlNet. In this paper, we first explore the temporal dynamics of information infusion through ControlNet, revealing that the input from LR images predominantly influences the initial stages of the denoising process. Leveraging this insight, we introduce a novel timestep-aware diffusion model that adaptively integrates features from both ControlNet and the pre-trained Stable Diffusion (SD). Our method enhances the transmission of LR information in the early stages of diffusion to guarantee image fidelity and stimulates the generation ability of the SD model itself more in the later stages to enhance the detail of generated images. To train this method, we propose a timestep-aware training strategy that adopts distinct losses at varying timesteps and acts on disparate modules. Experiments on benchmark datasets demonstrate the effectiveness of our method.
Qinwei Lin, Xiaopeng Sun 0001, Yu Gao 0027, Zheng Zhao 0004, Dengjie Li, Haoqian Wang
ACM Multimedia7
2025 Towards Fine-Grained Human Motion Video Captioning
abstract
Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semantically inconsistent captions. In this work, we introduce the Motion-Augmented Caption Model (M-ACM), a novel generative framework that enhances caption quality by incorporating motion-aware decoding. At its core, M-ACM leverages motion representations derived from human mesh recovery to explicitly highlight human body dynamics, thereby reducing hallucinations and improving both semantic fidelity and spatial alignment in the generated captions. To support research in this area, we present the Human Motion Insight (HMI) Dataset, comprising 115K video-description pairs focused on human movement, along with HMI-Bench, a dedicated benchmark for evaluating motion-focused video captioning. Experimental results demonstrate that M-ACM significantly outperforms previous methods in accurately describing complex human motions and subtle temporal variations, setting a new standard for motion-centric video captioning.
Guorui Song, Guocun Wang, Xuefei Zhe, Haoqian Wang
ACM Multimedia7
2025 Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has demonstrated impressive performance in novel view synthesis under dense-view settings. However, in sparse-view scenarios, despite the realistic renderings in training views, 3DGS occasionally manifests appearance artifacts in novel views. This paper investigates the appearance artifacts in sparse-view 3DGS and uncovers a core limitation of current approaches: the optimized Gaussians are overly-entangled with one another to aggressively fit the training views, which leads to a neglect of the real appearance distribution of the underlying scene and results in appearance artifacts in novel views. The analysis is based on a proposed metric, termed Co-Adaptation Score (CA), which quantifies the entanglement among Gaussians, i.e., co-adaptation, by computing the pixel-wise variance across multiple renderings of the same viewpoint, with different random subsets of Gaussians. The analysis reveals that the degree of co-adaptation is naturally alleviated as the number of training views increases. Based on the analysis, we propose two lightweight strategies to explicitly mitigate the co-adaptation in sparse-view 3DGS: (1) random gaussian dropout; (2) multiplicative noise injection to the opacity. Both strategies are designed to be plug-and-play, and their effectiveness is validated across various methods and benchmarks. We hope that our insights into the co-adaptation effect will inspire the community to achieve a more comprehensive understanding of sparse-view 3DGS.
Kangjie Chen, Yingji Zhong, Youyu Chen, Minghan Qin, Haoqian Wang
NeurIPS7
2025 MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds
abstract
Reconstructing 3D objects into editable programs is pivotal for applications like reverse engineering and shape editing. However, existing methods often rely on limited domain-specific languages (DSLs) and small-scale datasets, restricting their ability to model complex geometries and structures. To address these challenges, we introduce MeshLLM, a novel framework that reconstructs complex 3D objects from point clouds into editable Blender Python scripts. We develop a comprehensive set of expressive Blender Python APIs capable of synthesizing intricate geometries. Leveraging these APIs, we construct a large-scale paired object-code dataset, where the code for each object is decomposed into distinct semantic parts. Subsequently, we train a multimodal large language model (LLM) that translates 3D point cloud into executable Blender Python scripts. Our approach not only achieves superior performance in shape-to-code reconstruction tasks but also facilitates intuitive geometric and topological editing through convenient code modifications. Furthermore, our code-based representation enhances the reasoning capabilities of LLMs in 3D shape understanding tasks. Together, these contributions establish MeshLLM as a powerful and flexible solution for programmatic 3D shape reconstruction and understanding.
Bingquan Dai, Li Ray Luo, Qihong Tang, Xinyu Lian, Minghan Qin, Xudong Xu, Bo Dai 0002, Haoqian Wang, Zhaoyang Lyu, Jiangmiao Pang
NeurIPS10
2025 VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task Awareness
abstract
Recent advances in visual tokenizers have demonstrated their effectiveness for multimodal large language models and autoregressive generative models. However, most existing visual tokenizers rely on a fixed downsampling rate at a given visual resolution, and consequently produce a constant number of visual tokens, ignoring the fact that visual information of varying complexity warrant different token budgets. Motivated by this observation, we propose an adaptive video tokenizer "VaporTok" with two core contributions:Probabilistic Taildrop: We introduce a novel taildrop mechanism that learns a truncation index sampling distribution conditioned on visual complexity of the video. During both training and inference, the decoder reconstructs videos at adaptive token lengths, allocating more tokens to complex videos and fewer to simpler ones. Parallel Sample GRPO with Vapor Reward: By leveraging the probability distribution produced by probabilistic taildrop, we reformulate the visual tokenization pipeline as a sequential decision process. To optimize this process, we propose a variant of GRPO and a composite reward encompassing token efficiency, reconstruction fidelity, and generative quality, thus enabling metrics-aware adaptive tokenization across diverse objectives. Extensive experiments on standard video generation benchmarks confirm our analysis, showing that our adaptive approach matches or outperforms fixed‐rate baselines and naive taildrop while using fewer tokens.
Zechen Bai, Haoqian Wang, Alex Jinpeng Wang
NeurIPS4
2025 CULC-Net: A Recipe for Tailored Creative Selection in Online Advertising
Baosheng Zhang, Liufang Sang, Wei Wang 0103, Changping Peng, Zhangang Lin, Jingping Shao, Jie He 0005, Haoqian Wang
ECML/PKDD (5)10
2025 BBATProt: a framework predicting biological function with enhanced feature extraction via interpretable deep learning
abstract
Accurate prediction of protein and peptide functions from amino acid sequences is essential for understanding biological processes and advancing biomolecular engineering. Due to the limitations of experimental methods, computational approaches, particularly machine learning, have gained significant attention. However, many existing tools are task-specific and lack adaptability. Here, we propose a BERT-BiLSTM-Attention-TCN Protein Function Prediction Framework (BBATProt), a versatile framework for predicting protein and peptide functions. BBATProt leverages transfer learning with a pretrained bidirectional encoder representations from transformer model to capture high-dimensional features. The custom network integrates bidirectional long short-term memory and temporal convolutional network to align with proteins' spatial characteristics, combining local and global feature extraction via attention mechanisms to achieve more precise predictions. Evaluations demonstrate that BBATProt consistently outperforms state-of-the-art models in tasks such as hydrolytic catalysis, peptide bioactivity, and post-translational modification (PTM) site prediction. Specifically, BBATProt improves accuracy by 2.96%-41.96% in antimicrobial peptide (AMP) prediction and by 0.64%-23.54% in PTM prediction tasks. In terms of area under the receiver operating characteristic curve, improvements range from 0.71% to 40.51% for AMP prediction and 0.62%-27.82% for PTM prediction. Visualizations of feature evolution and refinement via attention mechanisms validate the framework's interpretability, providing transparency into the feature-extraction process and offering deeper insights into the basis of property prediction.
Youqing Wang, Xukai Ye, Haoqian Wang, Xin Ma 0012
Briefings Bioinform.4
2025 FCA-Net: Accelerating stereo image compression through cascade alignment of side information
Yichong Xia, Yujun Huang, Bin Chen 0011, Genping Wang, Haoqian Wang, Yaowei Wang 0001
Pattern Recognit.5
2025 Jarque-Bera-Based Artificial Neural Correlation Analysis for Nonlinear and Non-Gaussian Process Monitoring
abstract
Nonlinear and non-Gaussian characteristics are common in industrial processes. Artificial neural correlation analysis (ANCA) is a good nonlinear process monitoring algorithm, which combines classical correlation analysis with artificial neural networks. However, its performance is not very satisfactory for industrial processes with non-Gaussian characteristics. To solve non-Gaussian problems, almost all the existing process monitoring algorithms only consider the effect of kurtosis. Nevertheless, both kurtosis and skewness affect the data distribution. To improve the limitations of existing algorithms, this study proposes a new process monitoring algorithm named Jarque–Bera-based ANCA. This new algorithm makes many improvements to ANCA scheme, and the designed loss function combines the influence of both kurtosis and skewness on the data distribution, which not only maintains the advantages of the ANCA algorithm in solving nonlinear problems, but also provides superior monitoring performance in non-Gaussian processes. Furthermore, the superior performance of the proposed new algorithm is verified through simulations using non-Gaussian and nonlinear numerical examples, the Tennessee Eastman process, and catalytic cracking units.
Youqing Wang, Haoqian Wang, Tongze Hou, Xukai Ye, Silvio Simani, Xin Ma 0012
IEEE Trans. Syst. Man Cybern. Syst.2
2024 High-Fidelity 3D Head Avatars Reconstruction through Spatially-Varying Expression Conditioned Neural Radiance Field
abstract
One crucial aspect of 3D head avatar reconstruction lies in the details of facial expressions. Although recent NeRF-based photo-realistic 3D head avatar methods achieve high-quality avatar rendering, they still encounter challenges retaining intricate facial expression details because they overlook the potential of specific expression variations at different spatial positions when conditioning the radiance field. Motivated by this observation, we introduce a novel Spatially-Varying Expression (SVE) conditioning. The SVE can be obtained by a simple MLP-based generation network, encompassing both spatial positional features and global expression information. Benefiting from rich and diverse information of the SVE at different positions, the proposed SVE-conditioned NeRF can deal with intricate facial expressions and achieve realistic rendering and geometry details of high-fidelity 3D head avatars. Additionally, to further elevate the geometric and rendering quality, we introduce a new coarse-to-fine training strategy, including a geometry initialization strategy at the coarse stage and an adaptive importance sampling strategy at the fine stage. Extensive experiments indicate that our method outperforms other state-of-the-art (SOTA) methods in rendering and geometry quality on mobile phone-collected and public datasets. Code and data can be found at https://github.com/minghanqin/AvatarSVE.
Minghan Qin, Yuelang Xu, Xiaochen Zhao, Yebin Liu, Haoqian Wang
AAAI6
2024 TexVocab: Texture Vocabulary-Conditioned Human Avatars
abstract
To adequately utilize the available image evidence in multi-view video-based avatar modeling, we propose TexVocab, a novel avatar representation that constructs a texture vocabulary and associates body poses with texture maps for animation. Given multi-view RGB videos, our method initially back-projects all the available images in the training videos to the posed SMPL surface, producing texture maps in the SMPL UV domain. Then we construct pairs of human poses and texture maps to establish a texture vocabulary for encoding dynamic human appearances under various poses. Unlike the commonly used joint-wise manner, we further design a body-part-wise encoding strategy to learn the structural effects of the kinematic chain. Given a driving pose, we query the pose feature hierarchically by decomposing the pose vector into several body parts and interpolating the texture features for synthesizing fine-grained human dynamics. Overall, our method is able to create animatable avatars with detailed and dynamic appearances from RGB videos, and the experiments show that our method outperforms state-of-the-art approaches. The project page can be found at https://texvocab.github.io/.
Zhe Li 0027, Yebin Liu, Haoqian Wang
CVPR4
2024 LangSplat: 3D Language Gaussian Splatting
abstract
Humans live in a 3D world and commonly use natural language to interact with a 3D scene. Modeling a 3D language field to support open-ended language queries in 3D has gained increasing attention recently. This paper introduces LangSplat, which constructs a 3D language field that enables precise and efficient open-vocabulary querying within 3D spaces. Unlike existing methods that ground CLIP language embeddings in a NeRF model, LangSplat advances the field by utilizing a collection of 3D Gaussians, each encoding language features distilled from CLIP, to represent the language field. By employing a tile-based splatting technique for rendering language features, we circumvent the costly rendering process inherent in NeRF. Instead of directly learning CLIP embeddings, LangSplat first trains a scene-wise language autoencoder and then learns language features on the scene-specific latent space, thereby alleviating substantial memory demands imposed by explicit modeling. Existing methods struggle with imprecise and vague 3D language fields, which fail to discern clear boundaries between objects. We delve into this issue and propose to learn hierarchical semantics using SAM, thereby eliminating the need for extensively querying the language field across various scales and the regularization of DINO features. Extensive experimental results show that LangSplat significantly outperforms the previous state-of-the-art method LERF by a large margin. Notably, LangSplat is extremely efficient, achieving a 199 x speedup compared to LERF at the resolution of 1440 x 1080. We strongly recommend readers to check out our video results at https://langsplat.github.io/
Minghan Qin, Wanhua Li 0001, Haoqian Wang, Hanspeter Pfister
CVPR4
2024 Category-Level Object Detection, Pose Estimation and Reconstruction from Stereo Images
Chuanrui Zhang, Yonggen Ling, Minglei Lu, Minghan Qin, Haoqian Wang
ECCV (34)5
2024 Gaussian in the Wild: 3D Gaussian Splatting for Unconstrained Image Collections
Dongbin Zhang, Chuming Wang, Peihao Li 0003, Minghan Qin, Haoqian Wang
ECCV (76)6
2024 M2Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation
Yingshuang Zou, Yikang Ding, Xi Qiu, Haoqian Wang
ECCV (46)4
2024 DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields For High-Fidelity Talking Portrait Synthesis
abstract
In this paper, we present the decomposed triplane-hash neural radiance fields (DT-NeRF), a framework that significantly improves the photorealistic rendering of talking faces and achieves state-of-the-art results on key evaluation datasets. Our architecture decomposes the facial region into two specialized triplanes: one specialized for representing the mouth, and the other for the broader facial features. We introduce audio features as residual terms and integrate them as query vectors into our model through an audio-mouth-face transformer. Additionally, our method leverages the capabilities of Neural Radiance Fields (NeRF) to enrich the volumetric representation of the entire face through additive volumetric rendering techniques. Comprehensive experimental evaluations corroborate the effectiveness and superiority of our proposed approach.
Yaoyu Su, Shaohui Wang, Haoqian Wang
ICASSP3
2024 ITportrait: Image-Text Coupled 3D Portrait Domain Adaptation
abstract
Domain adaptation of 3D portraits has gained more and more attention. However, the transfer mechanism of existing methods is mainly based on vision or language, which ignores the potential of vision-language combined guidance. In this paper, we propose an Image-Text multi-modal framework, namely Image and Text portrait (ITportrait), for 3D portrait domain adaptation. ITportrait relies on a two-stage alternating training strategy. In the first stage, we employ a 3D Artistic Paired Transfer (APT) method for image-guided style transfer. APT constructs paired photo-realistic portraits to obtain accurate artistic poses, which helps ITportrait to achieve high-quality 3D style transfer. In the second stage, we propose a 3D Image-Text Embedding (ITE) approach in the CLIP space. ITE uses a threshold function to self-adaptively control the optimization direction of images or texts in the CLIP space. Comprehensive experiments prove that our ITportrait achieves state-of-the-art (SOTA) results. All source codes and pre-trained models will be released to the public.
Xiangwen Deng, Yuanhao Cai, Jingxiang Sun, Yebin Liu, Haoqian Wang
ICME6
2024 EyebrowNet: High-Precision Eyebrow Reconstruction and Matting
abstract
Eyebrows play a crucial role in facial reconstruction. However, due to their fine details and the lack of precise eyebrow datasets, traditional methods struggle to extract eyebrow features. To address this, we propose EyebrowNet, a high-precision eyebrow matting network consisting of an optimization phase and an inference phase, to extract high-quality eyebrow masks from blurry inputs. In the optimization phase, we develop a diffusion-based super-resolution network to enhance low-quality eyebrow images. To overcome the limited real eyebrow data, we employ a progressive fine-tuning strategy. In the inference phase, we devise a matting network based on conditional generative adversarial networks. To capture the rich details in real eyebrow images, we employ a synthetic training strategy. Comprehensive experiments prove that our EyebrowNet achieves state-of-the-art results. All source codes and datasets will be released to the public.
Wensen Feng, Haoqian Wang
ICME3
2024 Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
Xiaowan Hu, Yan Li 0043, Minquan Wang, Haoqian Wang, Quan Chen 0006, Han Li 0005, Peng Jiang 0002
ACM Multimedia5
2024 Animatable 3D Gaussian: Fast and High-Quality Reconstruction of Multiple Human Avatars
abstract
Neural radiance fields are capable of reconstructing high-quality drivable human avatars but are expensive to train and render and not suitable for multi-human scenes with complex shadows. To reduce consumption, we propose Animatable 3D Gaussian, which learns human avatars from input images and poses. We extend 3D Gaussians to dynamic human scenes by modeling a set of skinned 3D Gaussians and a corresponding skeleton in canonical space and deforming 3D Gaussians to posed space according to the input poses. We introduce a multi-head hash encoder for pose-dependent shape and appearance and a time-dependent ambient occlusion module to achieve high-quality reconstructions in scenes containing complex motions and dynamic shadows. On both novel view synthesis and novel pose synthesis tasks, our method achieves higher reconstruction quality than InstantAvatar with less training time (1/60), less GPU memory (1/4), and faster rendering speed (7x). Our method can be easily extended to multi-human scenes and achieve comparable novel view synthesis results on a scene with ten people in only 25 seconds of training.
Minghan Qin, Qinwei Lin, Haoqian Wang
ACM Multimedia5
2024 OPAL: Occlusion Pattern Aware Loss for Unsupervised Light Field Disparity Estimation
abstract
Light field disparity estimation is an essential task in computer vision. Currently, supervised learning-based methods have achieved better performance than both unsupervised and optimization-based methods. However, the generalization capacity of supervised methods on real-world data, where no ground truth is available for training, remains limited. In this paper, we argue that unsupervised methods can achieve not only much stronger generalization capacity on real-world data but also more accurate disparity estimation results on synthetic datasets. To fulfill this goal, we present the Occlusion Pattern Aware Loss, named OPAL, which successfully extracts and encodes general occlusion patterns inherent in the light field for calculating the disparity loss. OPAL enables: i) accurate and robust disparity estimation by teaching the network how to handle occlusions effectively and ii) significantly reduced network parameters required for accurate and efficient estimation. We further propose an EPI transformer and a gradient-based refinement module for achieving more accurate and pixel-aligned disparity estimation results. Extensive experiments demonstrate our method not only significantly improves the accuracy compared with SOTA unsupervised methods, but also possesses stronger generalization capacity on real-world data compared with SOTA supervised methods. Last but not least, the network training and inference efficiency are much higher than existing learning-based methods. Our code will be made publicly available.
Jiayin Zhao, Jingyao Wu 0003, Chao Deng 0005, Yuqi Han, Haoqian Wang, Tao Yu 0007
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Multi-scale architectures matter: Examining the adversarial robustness of flow-based lossless compression
Yichong Xia, Bin Chen 0011, Tianshuo Ge, Yujun Huang, Haoqian Wang, Yaowei Wang 0001
Pattern Recognit.6
2024 Edge-Aware Attention Transformer for Image Super-Resolution
abstract
In this study, we explore poor edge reconstruction in image super-resolution (SR) tasks, emphasizing the significance of enhancing edge details identified through visual analysis. Existing SR networks typically optimize their network architectures, enabling complete feature extraction from feature maps. This is because the management of spatial and channel information during SR is often pivotal to the network's feature extraction capacity. Despite continuous improvements, directly comparing SR and high-resolution (HR) images through differential mapping reveals the suboptimal performance of these methods in edge reconstruction. In this paper, we introduce a edgey-aware attention transformer (EAT), which focuses on edge reconstruction while maintaining the effective original low frequency information retrieval. Our framework utilizes deformable convolution (DC) to adaptively extract edge features. Then feature enhancement techniques are employed to intensify edge-sensitive features. Furthermore, extensive experiments demonstrate our EAT's exceptional quantitative and visual results, which surpass most benchmarks. This validates the EAT's effectiveness when compared to state-of-the-art models. The code is available athttps://github.com/ImWangHaoqian/EAT.
Haoqian Wang, Zhongyang Xing, Zhongjie Xu, Xiangai Cheng, Teng Li 0011
IEEE Signal Process. Lett.1
2023 Calibrated Teacher for Sparsely Annotated Object Detection
abstract
Fully supervised object detection requires training images in which all instances are annotated. This is actually impractical due to the high labor and time costs and the unavoidable missing annotations. As a result, the incomplete annotation in each image could provide misleading supervision and harm the training. Recent works on sparsely annotated object detection alleviate this problem by generating pseudo labels for the missing annotations. Such a mechanism is sensitive to the threshold of the pseudo label score. However, the effective threshold is different in different training stages and among different object detectors. Therefore, the current methods with fixed thresholds have sub-optimal performance, and are difficult to be applied to other detectors. In order to resolve this obstacle, we propose a Calibrated Teacher, of which the confidence estimation of the prediction is well calibrated to match its real precision. In this way, different detectors in different training stages would share a similar distribution of the output confidence, so that multiple detectors could share the same fixed threshold and achieve better performance. Furthermore, we present a simple but effective Focal IoU Weight (FIoU) for the classification loss. FIoU aims at reducing the loss weight of false negative samples caused by the missing annotation, and thus works as the complement of the teacher-student paradigm. Extensive experiments show that our methods set new state-of-the-art under all different sparse settings in COCO. Code will be available at https://github.com/Whileherham/CalibratedTeacher.
Haohan Wang, Liang Liu 0007, Boshen Zhang, Jiangning Zhang, Wuhao Zhang, Zhenye Gan, Yabiao Wang, Chengjie Wang 0001, Haoqian Wang
AAAI9
2023 One-Stage 3D Whole-Body Mesh Recovery with Component Aware Transformer
abstract
Whole-body mesh recovery aims to estimate the 3D human body, face, and hands parameters from a single image. It is challenging to perform this task with a single network due to resolution issues, i.e., the face and hands are usually located in extremely small regions. Existing works usually detect hands and faces, enlarge their resolution to feed in a specific network to predict the parameter, and finally fuse the results. While this copy-paste pipeline can capture the fine-grained details of the face and hands, the connections between different parts cannot be easily recovered in late fusion, leading to implausible 3D rotation and unnatural pose. In this work, we propose a one-stage pipeline for expressive whole-body mesh recovery, named OSX, without separate networks for each part. Specifically, we design a Component Aware Transformer (CAT) composed of a global body encoder and a local face/hand decoder. The encoder predicts the body parameters and provides a high-quality feature map for the decoder, which performs a feature-level upsample-crop scheme to extract highresolution part-specific features and adopt keypointguided deformable attention to estimate hand and face precisely. The whole pipeline is simple yet effective without any manual post-processing and naturally avoids implausible prediction. Comprehensive experiments demonstrate the effectiveness of OSX. Lastly, we build a large-scale Upper-Body dataset (UBody) with high-quality 2D and 3D whole-body annotations. It contains persons with partially visible bodies in diverse real-life scenarios to bridge the gap between the basic task and downstream applications.
Ailing Zeng, Haoqian Wang, Lei Zhang 0001, Yu Li 0003
CVPR3
2023 Learning Visibility Field for Detailed 3D Human Reconstruction and Relighting
abstract
Detailed 3D reconstruction and photo-realistic relighting of digital humans are essential for various applications. To this end, we propose a novel sparse-view 3d human reconstruction framework that closely incorporates the occupancy field and albedo field with an additional visibility field-it not only resolves occlusion ambiguity in multi-view feature aggregation, but can also be used to evaluate light attenuation for self-shadowed relighting. To enhance its training viability and efficiency, we discretize visibility onto a fixed set of sample directions and supply it with coupled geometric 3D depth feature and local 2D image feature. We further propose a novel rendering-inspired loss, namely TransferLoss, to implicitly enforce the alignment between visibility and occupancy field, enabling end-to-end joint training. Results and extensive experiments demonstrate the effectiveness of the proposed method, as it surpasses state-of-the-art in terms of reconstruction accuracy while achieving comparably accurate relighting to ray-traced ground truth.
Ruichen Zheng, Haoqian Wang, Tao Yu 0007
CVPR3
2023 Prior-Enhanced Temporal Action Localization Using Subject-Aware Spatial Attention
abstract
Temporal action localization (TAL) aims to detect the boundary and identify the class of each action instance in a long untrimmed video. Current approaches treat video frames homogeneously, and tend to give background and key objects excessive attention. This limits their sensitivity to localize action boundaries. To this end, we propose a prior-enhanced temporal action localization method (PETAL), which only takes in RGB input and incorporates action subjects as priors. This proposal leverages action subjects’ information with a plug-and-play subject-aware spatial attention module (SA-SAM) to generate an aggregated and subject-prioritized representation. Experimental results on THUMOS-14 and ActivityNet-1.3 datasets demonstrate that the proposed PETAL achieves competitive performance using only RGB features, e.g., boosting mAP by 2.41% or 0.25% over the state-of-the-art approach that uses RGB features or with additional optical flow features on the THUMOS-14 dataset.
Youbao Tang, Ruei-Sung Lin, Haoqian Wang
ICASSP5
2023 Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement
abstract
When enhancing low-light images, many deep learning algorithms are based on the Retinex theory. However, the Retinex model does not consider the corruptions hidden in the dark or introduced by the light-up process. Besides, these methods usually require a tedious multi-stage training pipeline and rely on convolutional neural networks, showing limitations in capturing long-range dependencies. In this paper, we formulate a simple yet principled One-stage Retinex-based Framework (ORF). ORF first estimates the illumination information to light up the low-light image and then restores the corruption to produce the enhanced image. We design an Illumination-Guided Transformer (IGT) that utilizes illumination representations to direct the modeling of non-local interactions of regions with different lighting conditions. By plugging IGT into ORF, we obtain our algorithm, Retinexformer. Comprehensive quantitative and qualitative experiments demonstrate that our Retinexformer significantly outperforms state-of-the-art methods on thirteen benchmarks. The user study and application on low-light object detection also reveal the latent practical values of our method. Code is available at https://github.com/caiyuanhao1998/Retinexformer
Yuanhao Cai, Hao Bian, Haoqian Wang, Radu Timofte, Yulun Zhang 0001
ICCV4
2023 Template-guided Hierarchical Feature Restoration for Anomaly Detection
abstract
Targeting for detecting anomalies of various sizes for complicated normal patterns, we propose a Template-guided Hierarchical Feature Restoration method, which introduces two key techniques, bottleneck compression and template-guided compensation, for anomaly-free feature restoration. Specially, our framework compresses hierarchical features of an image by bottleneck structure to preserve the most crucial features shared among normal samples. We design template-guided compensation to restore the distorted features towards anomaly-free features. Particularly, we choose the most similar normal sample as the template, and leverage hierarchical features from the template to compensate the distorted features. The bottleneck could partially filter out anomaly features, while the compensation further converts the reminding anomaly features towards normal with template guidance. Finally, anomalies are detected in terms of the cosine distance between the pre-trained features of an inference image and the corresponding restored anomaly-free features. Experimental results demonstrate the effectiveness of our approach, which achieves the state-of-the-art performance on the MVTec LOCO AD dataset.
Hewei Guo, Liping Ren, Jingjing Fu, Yuwang Wang, Zhizheng Zhang 0004, Cuiling Lan, Haoqian Wang, Xinwen Hou
ICCV7
2023 NeRF-MS: Neural Radiance Fields with Multi-Sequence
abstract
Neural radiance fields (NeRF) achieve impressive performance in novel view synthesis when trained on only single sequence data. However, leveraging multiple sequences captured by different cameras at different times is essential for better reconstruction performance. Multi-sequence data takes two main challenges: appearance variation due to different lighting conditions and non-static objects like pedestrians. To address these issues, we propose NeRF-MS, a novel approach to training NeRF with multi-sequence data. Specifically, we utilize a triplet loss to regularize the distribution of per-image appearance code, which leads to better high-frequency texture and consistent appearance, such as specular reflections. Then, we explicitly model non-static objects to reduce floaters. Extensive results demonstrate that NeRF-MS not only outperforms state-of-the-art view synthesis methods on outdoor and synthetic scenes, but also achieves 3D consistent rendering and robust appearance controlling. Project page: https://nerf-ms.github.io/.
Peihao Li 0003, Shaohui Wang, Chen Yang 0023, Weichao Qiu, Haoqian Wang
ICCV6
2023 LNPL-MIL: Learning from Noisy Pseudo Labels for Promoting Multiple Instance Learning in Whole Slide Image
abstract
Gigapixel Whole Slide Images (WSIs) aided patient diagnosis and prognosis analysis are promising directions in computational pathology. However, limited by expensive and time-consuming annotation costs, WSIs usually only have weak annotations, including 1) WSI-level Annotations (WA) and 2) Limited Patch-level Annotations (LPA). Currently, Multiple Instance Learning (MIL) often exploits WA, while LPA usually assign pseudo-labels for unlabeled data. Intuitively, pseudo-labels can serve as a practical guide for MIL, but the unreliable prediction caused by LPA inevitably introduce noise. Furthermore, WA-supervised MIL training inevitably suffers from the semantical unalignment between instances and bag-level labels. To address these problems, we design a framework called Learning from Noisy Pseudo Labels for promoting Multiple Instance Learning (LNPL-MIL), which considers both types of weak annotation. Specifically, for the LPA-trained weak classifier, we design a Super-Patch-based LNPL (SP-LNPL) method to reduce false positives in the noisy pseudo-labels and then select more accurate Top-K key instances. In MIL, we propose a Transformer aware of instance Order and Distribution (TOD-MIL) that strengthens instances correlation and weakens semantical unalignment in the bag. We validate our LNPL-MIL on Tumor Diagnosis and Survival Prediction, achieving state-of-the-art performance with at least 2.7%/2.9% AUC and 2.6%/2.3% C-Index improvement with the patches labeled for two scale. Ablation study and visualization analysis further verify the effectiveness.
Zhuchen Shao, Yifeng Wang 0001, Yang Chen 0036, Hao Bian, Shaohui Liu, Haoqian Wang, Yongbing Zhang 0002
ICCV6
2023 Visual and Spatial Context Fusion for Implicit Human Reconstruction
abstract
3D human reconstruction aims to recover the 3D mesh of clothed-human from multi-view images. Recently, deep implicit function methods have won great success in this task for their detailed modeling. However, these efforts typically learn the implicit function in a point-wise manner, which ignores local context, resulting in shape artifacts. In this paper, we propose a Visual and Spatial Context fusion Implicit Function network, named VSC-IF. Specifically, we design two key modules: (i) a transformer-based encoder to model local geometry and learn global shape dependencies from images, and (ii) a feature fusion module to provide spatial context information for reconstruction. We validate our method and evaluate the generalization performance on two common datasets. Experiments show that our model achieves a new state-of-the-art performance, especially, its visual results exhibit less shape distortion and broken limbs than previous methods.
Haoqian Wang
ICIP4
2023 Binarized Spectral Compressive Imaging
abstract
Existing deep learning models for hyperspectral image (HSI) reconstruction achieve good performance but require powerful hardwares with enormous memory and computational resources. Consequently, these methods can hardly be deployed on resource-limited mobile devices. In this paper, we propose a novel method, Binarized Spectral-Redistribution Network (BiSRNet), for efficient and practical HSI restoration from compressed measurement in snapshot compressive imaging (SCI) systems. Firstly, we redesign a compact and easy-to-deploy base model to be binarized. Then we present the basic unit, Binarized Spectral-Redistribution Convolution (BiSR-Conv). BiSR-Conv can adaptively redistribute the HSI representations before binarizing activation and uses a scalable hyperbolic tangent function to closer approximate the Sign function in backpropagation. Based on our BiSR-Conv, we customize four binarized convolutional modules to address the dimension mismatch and propagate full-precision information throughout the whole network. Finally, our BiSRNet is derived by using the proposed techniques to binarize the base model. Comprehensive quantitative and qualitative experiments manifest that our proposed BiSRNet outperforms state-of-the-art binarization algorithms. Code and models are publicly available at https://github.com/caiyuanhao1998/BiSCI
Yuanhao Cai, Xin Yuan 0002, Yulun Zhang 0001, Haoqian Wang
NeurIPS6
2023 Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
abstract
In this paper, we present Motion-X, a large-scale 3D expressive whole-body motion dataset. Existing motion datasets predominantly contain body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions. Moreover, they are primarily collected from limited laboratory scenes with textual descriptions manually labeled, which greatly limits their scalability. To overcome these limitations, we develop a whole-body motion and text annotation pipeline, which can automatically annotate motion from either single- or multi-view videos and provide comprehensive semantic labels for each video and fine-grained whole-body pose descriptions for each frame. This pipeline is of high precision, cost-effective, and scalable for further research. Based on it, we construct Motion-X, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes. Besides, Motion-X provides 15.6M frame-level whole-body pose descriptions and 81.1K sequence-level semantic labels. Comprehensive experiments demonstrate the accuracy of the annotation pipeline and the significant benefit of Motion-X in enhancing expressive, diverse, and natural motion generation, as well as 3D whole-body human mesh recovery.
Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, Lei Zhang 0001
NeurIPS6
2023 Decoupling multi-task causality for improved skin lesion segmentation and classification
Lei Song 0003, Haoqian Wang, Z. Jane Wang 0001
Pattern Recognit.2
2023 dMIL-Transformer: Multiple Instance Learning Via Integrating Morphological and Spatial Information for Lymph Node Metastasis Classification
abstract
Automated classification of lymph node metastasis (LNM) plays an important role in the diagnosis and prognosis. However, it is very challenging to achieve satisfactory performance in LNM classification, because both the morphology and spatial distribution of tumor regions should be taken into account. To address this problem, this article proposes a two-stage dMIL-Transformer framework, which integrates both the morphological and spatial information of the tumor regions based on the theory of multiple instance learning (MIL). In the first stage, a double Max-Min MIL (dMIL) strategy is devised to select the suspected top-K positive instances from each input histopathology image, which contains tens of thousands of patches (primarily negative). The dMIL strategy enables a better decision boundary for selecting the critical instances compared with other methods. In the second stage, a Transformer-based MIL aggregator is designed to integrate all the morphological and spatial information of the selected instances from the first stage. The self-attention mechanism is further employed to characterize the correlation between different instances and learn the bag-level representation for predicting the LNM category. The proposed dMIL-Transformer can effectively deal with the thorny classification in LNM with great visualization and interpretability. We conduct various experiments over three LNM datasets, and achieve 1.79%-7.50% performance improvement compared with other state-of-the-art methods.
Yang Chen 0036, Zhuchen Shao, Hao Bian, Zijie Fang, Yifeng Wang 0001, Yuanhao Cai, Haoqian Wang, GuoJun Liu, Yongbing Zhang 0002
IEEE J. Biomed. Health Informatics7
2022 Unpaired Multi-Domain Stain Transfer for Kidney Histopathological Images
abstract
As an essential step in the pathological diagnosis, histochemical staining can show specific tissue structure information and, consequently, assist pathologists in making accurate diagnoses. Clinical kidney histopathological analyses usually employ more than one type of staining: H&E, MAS, PAS, PASM, etc. However, due to the interference of colors among multiple stains, it is not easy to perform multiple staining simultaneously on one biological tissue. To address this problem, we propose a network based on unpaired training data to virtually generate multiple types of staining from one staining. Our method can preserve the content of input images while transferring them to multiple target styles accurately. To efficiently control the direction of stain transfer, we propose a style guided normalization (SGN). Furthermore, a multiple style encoding (MSE) is devised to represent the relationship among different staining styles dynamically. An improved one-hot label is also proposed to enhance the generalization ability and extendibility of our method. Vast experiments have demonstrated that our model can achieve superior performance on a tiny dataset. The results exhibit not only good performance but also great visualization and interpretability. Especially, our method also achieves satisfactory results over cross-tissue, cross-staining as well as cross-task. We believe that our method will significantly influence clinical stain transfer and reduce the workload greatly for pathologists. Our code and Supplementary materials are available at https://github.com/linyiyang98/UMDST.
Yiyang Lin, Bowei Zeng, Yifeng Wang 0001, Yang Chen 0036, Zijie Fang, Jian Zhang 0018, Xiangyang Ji, Haoqian Wang, Yongbing Zhang 0002
AAAI8
2022 Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction
abstract
Hyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactions is beneficial for HSI reconstruction. However, existing CNN-based methods show limitations in capturing spectral-wise similarity and long-range dependencies. Besides, the HSI information is modulated by a coded aperture (physical mask) in CASSI. Nonetheless, current algorithms have not fully explored the guidance effect of the mask for HSI restoration. In this paper, we propose a novel framework, Mask-guided Spectral-wise Transformer (MST), for HSI reconstruction. Specifically, we present a Spectral-wise Multi-head Self-Attention (S-MSA) that treats each spectral feature as a token and calculates self-attention along the spectral dimension. In addition, we customize a Mask-guided Mechanism (MM) that directs S- MSA to pay attention to spatial regions with high-fidelity spectral representations. Extensive experiments show that our MST significantly outperforms state-of-the-art (SOTA) methods on simulation and real HSI datasets while requiring dramatically cheaper computational and memory costs. https://github.com/caiyuanhao1998/MST/
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
CVPR4
2022 HDNet: High-resolution Dual-domain Learning for Spectral Compressive Imaging
abstract
The rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance against complexity, losing fine-grained high-resolution (HR) features. Secondly, even if the optimization focusing on spatial-spectral domain learning (SDL) converges to the ideal solution, there is still a significant visual difference between the reconstructed HSI and the truth. So we propose a high-resolution dual-domain learning network (HDNet) for HSI reconstruction. On the one hand, the proposed HR spatial-spectral attention module with its efficient feature fusion provides continuous and fine pixel-level features. On the other hand, frequency domain learning (FDL) is introduced for HSI reconstruction to narrow the frequency domain discrepancy. Dynamic FDL supervision forces the model to reconstruct fine-grained frequencies and compensate for excessive smoothing and distortion caused by pixel-level losses. The HR pixel-level attention and frequency-level refinement in our HDNet mutually promote HSI perceptual quality. Extensive quantitative and qualitative experiments show that our method achieves SOTA performance on simulated and real HSI datasets. https://github.com/Huxiaowan/HDNet
Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
CVPR4
2022 Diversity Matters: Fully Exploiting Depth Clues for Reliable Monocular 3D Object Detection
abstract
As an inherently ill-posed problem, depth estimation from single images is the most challenging part of monocular 3D object detection (M3OD). Many existing methods rely on preconceived assumptions to bridge the missing spatial information in monocular images, and predict a sole depth value for every object of interest. However, these assumptions do not always hold in practical applications. To tackle this problem, we propose a depth solving system that fully explores the visual clues from the subtasks in M3OD and generates multiple estimations for the depth of each target. Since the depth estimations rely on different assumptions in essence, they present diverse distributions. Even if some assumptions collapse, the estimations established on the remaining assumptions are still reliable. In addition, we develop a depth selection and combination strategy. This strategy is able to remove abnormal estimations caused by collapsed assumptions, and adaptively combine the remaining estimations into a single one. In this way, our depth solving system becomes more precise and robust. Exploiting the clues from multiple subtasks of M3OD and without introducing any extra information, our method surpasses the current best method by more than 20% relatively on the Moderate level of test split in the KITTI 3D object detection benchmark, while still maintaining real-time efficiency.
Zhuoling Li, Jianzhuang Liu, Haoqian Wang, Lihui Jiang
CVPR5
2022 Adaptive Sparse and Monotonic Attention for Transformer-based Automatic Speech Recognition
abstract
The Transformer architecture model, based on self-attention and multi-head attention, has achieved remarkable success in offline end-to-end Automatic Speech Recognition (ASR). However, self-attention and multi-head attention cannot be easily applied for streaming or online ASR. For self-attention in Transformer ASR, the softmax normalization function-based attention mechanism makes it impossible to highlight important speech information. For multi-head attention in Transformer ASR, it is not easy to model monotonic alignments in different heads. To overcome these two limits, we integrate sparse attention and monotonic attention into Transformer-based ASR. The sparse mechanism introduces a learned sparsity scheme to enable each self-attention structure to fit the corresponding head better. The monotonic attention deploys regularization to prune redundant heads for the multi-head attention structure. The experiments show that our method can effectively improve the attention mechanism on widely used benchmarks of speech recognition.
Chendong Zhao, Jianzong Wang, Xiaoyang Qu, Haoqian Wang, Jing Xiao 0006
DSAA5
2022 Cross-Modal Image-Text Matching via Coupled Projection Learning Hashing
abstract
Hashing, which aims to explore the correlations between modalities for multimedia data, shows numerous ad-vantages in cross-modal image-text matching tasks. However, most studies simply consider inter-modal correlations and ignore the semantic information contained within individual modalities, which produces suboptimal correlation representations between modalities. Based on this, we propose a Coupled Projection Learning Hashing method (CPLH). It directly maps the image-text data pairs into the different Hamming spaces to preserve as much of the original feature information as possible. Next, the CPLH joins the heterogeneous data over the Hamming spaces to maximize the correlation between two distinct modalities. In this way, the discriminative hash functions applied to the image-text matching process have semantic information from all modalities, which improves the matching accuracy. Besides, to further reduce the computational complexity of the proposed method, we propose a variant called CPLH-ts. It separates the learning process of the hash functions from the objective function by exploiting a two-step hashing strategy. Adequate experiments on three public datasets illustrate the superiority of our method and its variant compared to several state-of-the-art methods in matching accuracy and training efficiency.
Huan Zhao 0003, Haoqian Wang, Xupeng Zha, Song Wang 0016
DSAA2
2022 Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
ECCV (17)4
2022 r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled Noise Introducing and Contextual Information Incorporation
abstract
Grapheme-to-phoneme (G2P) conversion is the process of converting the written form of words to their pronunciations. It has an important role for text-to-speech (TTS) synthesis and automatic speech recognition (ASR) systems. In this paper, we aim to evaluate and enhance the robustness of G2P models. We show that neural G2P models are extremely sensitive to orthographical variations in graphemes like spelling mistakes. To solve this problem, we propose three controlled noise introducing methods to synthesize noisy training data. Moreover, we incorporate the contextual information with the baseline and propose a robust training strategy to stabilize the training process. The experimental results demonstrate that our proposed robust G2P model (r-G2P) outperforms the baseline significantly (-2.73% WER on Dict-based benchmarks and -9.09% WER on Real-world sources).
Chendong Zhao, Jianzong Wang, Xiaoyang Qu, Haoqian Wang, Jing Xiao 0006
ICASSP4
2022 Pyramid Knowledge Distillation for Efficient Human Pose Estimation
abstract
Human pose estimation is an important task in many real-time applications. Existing methods directly slim the CNN by deploying well-designed lightweight modules. However, these methods lack privileged information guidance and the knowledge distillation technique stays less explored. In this work, we propose a novel method, namely Pyramid Knowledge Distillation (PKD) for efficient human pose estimation. Specifically, PKD composes of Pyramid Structured Map Distillation (PSMD) and Pyramid Feature Map Distillation (PFMD). In PSMD, we formulate a structured map encoding robust interjoint correlation. Based on structured map, the spatial dependencies between keypoints can be better transferred from a cumbersome teacher network to a compact student model. To further promote the efficiency of student, PFMD is used to distill rich local and global features from teacher. Experiments demonstrate that PKD achieves an optimal trade-off between cost and accuracy on COCO and MPII benchmarks, even with a much faster inference speed.
Peng Jiao, Haoqian Wang
ICIP3
2022 Flow-Guided Sparse Transformer for Video Deblurring
abstract
Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Transformer (FGST), for video deblurring. In FGST, we customize a self-attention module, Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA). For each $query$ element on the blurry reference frame, FGSW-MSA enjoys the guidance of the estimated optical flow to globally sample spatially sparse yet highly related $key$ elements corresponding to the same scene patch in neighboring frames. Besides, we present a Recurrent Embedding (RE) mechanism to transfer information from past frames and strengthen long-range temporal dependencies. Comprehensive experiments demonstrate that our proposed FGST outperforms state-of-the-art (SOTA) methods on both DVD and GOPRO datasets and yields visually pleasant results in real video deblurring. https://github.com/linjing7/VR-Baseline
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Youliang Yan, Xueyi Zou, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
ICML4
2022 Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video Restoration
abstract
How to properly model the inter-frame relation within the video sequence is an important but unsolved challenge for video restoration (VR). In this work, we propose an unsupervised flow-aligned sequence-to-sequence model (S2SVR) to address this problem. On the one hand, the sequence-to-sequence model, which has proven capable of sequence modeling in the field of natural language processing, is explored for the first time in VR. Optimized serialization modeling shows potential in capturing long-range dependencies among frames. On the other hand, we equip the sequence-to-sequence model with an unsupervised optical flow estimator to maximize its potential. The flow estimator is trained with our proposed unsupervised distillation loss, which can alleviate the data discrepancy and inaccurate degraded optical flow issues of previous flow-based methods. With reliable optical flow, we can establish accurate correspondence among multiple frames, narrowing the domain difference between 1D language and 2D misaligned frames and improving the potential of the sequence-to-sequence model. S2SVR shows superior performance in multiple VR tasks, including video deblurring, video super-resolution, and compressed video quality enhancement. https://github.com/linjing7/VR-Baseline
Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Youliang Yan, Xueyi Zou, Yulun Zhang 0001, Luc Van Gool
ICML4
2022 Iterative Few-shot Semantic Segmentation from Image Label Text
abstract
Few-shot semantic segmentation aims to learn to segment unseen class objects with the guidance of only a few support images. Most previous methods rely on the pixel-level label of support images. In this paper, we focus on a more challenging setting, in which only the image-level labels are available. We propose a general framework to firstly generate coarse masks with the help of the powerful vision-language model CLIP, and then iteratively and mutually refine the mask predictions of support and query images. Extensive experiments on PASCAL-5i and COCO-20i datasets demonstrate that our method not only outperforms the state-of-the-art weakly supervised approaches by a significant margin, but also achieves comparable or better results to recent supervised methods. Moreover, our method owns an excellent generalization ability for the images in the wild and uncommon classes. Code will be available at https://github.com/Whileherham/IMR-HSNet.
Haohan Wang, Liang Liu 0007, Wuhao Zhang, Jiangning Zhang, Zhenye Gan, Yabiao Wang, Chengjie Wang 0001, Haoqian Wang
IJCAI8
2022 DAPID: A Differential-adaptive PID Optimization Strategy for Neural Network Training
abstract
Derived from automatic control theory, the PID optimizer for neural network training can effectively inhibit the overshoot phenomenon of conventional optimization algorithms such as SGD-Momentum. However, its differential term may unexpectedly have a relatively large scale during iteration, which may amplify the inherent noise of input samples and deteriorate the training process. In this paper, we adopt a self-adaptive iterating rule for the PID optimizer's differential term, which uses both first-order and second-order moment estimation to calculate the differential's unbiased statistical value approximately. Such strategy prevents the differential term from being divergent and accelerates the iteration without increasing much computational cost. Empirical results on several popular machine learning datasets demonstrate that the proposed optimization strategy achieves favorable acceleration of convergence as well as competitive accuracy compared with other stochastic optimization approaches.
Yulin Cai, Haoqian Wang
IJCNN2
2022 Multiple Instance Learning with Mixed Supervision in Gleason Grading
Hao Bian, Zhuchen Shao, Yang Chen 0036, Yifeng Wang 0001, Haoqian Wang, Jian Zhang 0018, Yongbing Zhang 0002
MICCAI (8)5
2022 Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging
abstract
In coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer from two issues. Firstly, they do not estimate the degradation patterns and ill-posedness degree from CASSI to guide the iterative learning. Secondly, they are mainly CNN-based, showing limitations in capturing long-range dependencies. In this paper, we propose a principled Degradation-Aware Unfolding Framework (DAUF) that estimates parameters from the compressed image and physical mask, and then uses these parameters to control each iteration. Moreover, we customize a novel Half-Shuffle Transformer (HST) that simultaneously captures local contents and non-local dependencies. By plugging HST into DAUF, we establish the first Transformer-based deep unfolding method, Degradation-Aware Unfolding Half-Shuffle Transformer (DAUHST), for HSI reconstruction. Experiments show that DAUHST surpasses state-of-the-art methods while requiring cheaper computational and memory costs. Code and models are publicly available at https://github.com/caiyuanhao1998/MST
Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
NeurIPS3
2022 Effective Backdoor Defense by Exploiting Sensitivity of Poisoned Samples
abstract
Poisoning-based backdoor attacks are serious threat for training deep models on data from untrustworthy sources. Given a backdoored model, we observe that the feature representations of poisoned samples with trigger are more sensitive to transformations than those of clean samples. It inspires us to design a simple sensitivity metric, called feature consistency towards transformations (FCT), to distinguish poisoned samples from clean samples in the untrustworthy training set. Moreover, we propose two effective backdoor defense methods. Built upon a sample-distinguishment module utilizing the FCT metric, the first method trains a secure model from scratch using a two-stage secure training module. And the second method removes backdoor from a backdoored model with a backdoor removal module which alternatively unlearns the distinguished poisoned samples and relearns the distinguished clean samples. Extensive results on three benchmark datasets demonstrate the superior defense performance against eight types of backdoor attacks, to state-of-the-art backdoor defenses. Codes are available at: https://github.com/SCLBD/Effectivebackdoordefense.
Baoyuan Wu, Haoqian Wang
NeurIPS3
2022 Learning Invariant Representation and Risk Minimized for Unsupervised Accent Domain Adaptation
abstract
Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the unsupervised paradigm needs to be carefully designed and little is known about what properties these representations acquire. There is no guarantee that the model learns meaningful representations for valuable information for recognition. Moreover, the adaptation ability of the learned representations to other domains still needs to be estimated. In this work, we explore learning domain-invariant representations via a direct mapping of speech representations to their corresponding high-level linguistic informations. Results prove that the learned latents not only capture the articulatory feature of each phoneme but also enhance the adaptation ability, outperforming the baseline largely on accented benchmarks.
Chendong Zhao, Jianzong Wang, Xiaoyang Qu, Haoqian Wang, Jing Xiao 0006
SLT4
2022 Affective feature knowledge interaction for empathetic conversation generation
abstract
A popular chatbot can generate natural and human-like responses, and the crucial technology is the ability to understand and appreciate the emotions and demands expressed from the perspective of the user. However, some empathetic dialogue generation models only specialise in commonsense and neglect emotion, which can only get a one-sided understanding of the user's situation and makes the model unable to express emotion better. In this paper, we propose a novel affective feature knowledge interactive model named AFKI, to enhance response generation performance, which enriches conversation history to obtain emotional interactive context by leveraging fine-grained emotional features and commonsense knowledge. Furthermore, we utilise an emotional interactive context encoder to learn higher-level affective interaction information and distill the emotional state feature to guide the empathetic response generation. The emotional features are to well capture the subtle differences of the user's emotional expression, and the commonsense knowledge improves the representation of affective information on generated responses. Extensive experiments on the empathetic conversation task demonstrate that our model generates multiple responses with higher emotion accuracy and stronger empathetic ability compared with baseline model approaches for empathetic response generation.
Ensi Chen, Huan Zhao 0003, Xupeng Zha, Haoqian Wang, Song Wang 0016
Connect. Sci.5
2022 Wide Weighted Attention Multi-Scale Network for Accurate MR Image Super-Resolution
abstract
High-quality magnetic resonance (MR) images afford more detailed information for reliable diagnoses and quantitative image analyses. Given low-resolution (LR) images, the deep convolutional neural network (CNN) has shown its promising ability for image super-resolution (SR). The LR MR images usually share some visual characteristics: structural textures of different sizes, edges with high correlation, and less informative background. However, multi-scale structural features are informative for image reconstruction, while the background is more smooth. Most previous CNN-based SR methods use a single receptive field and equally treat the spatial pixels (including the background). It neglects to sense the entire space and get diversified features from the input, which is critical for high-quality MR image SR. We propose a wide weighted attention multi-scale network ($\text{W}^{2}$AMSN) for accurate MR image SR to address these problems. On the one hand, the features of varying sizes can be extracted by the wide multi-scale branches. On the other hand, we design a non-reduction attention mechanism to recalibrate feature responses adaptively. Such attention preserves continuous cross-channel interaction and focuses on more informative regions. Meanwhile, the learnable weighted factors fuse extracted features selectively. The encapsulated wide weighted attention multi-scale block ($\text{W}^{2}$AMSB) is integrated through a recurrent framework and global attention mechanism. Extensive experiments and diversified ablation studies show the effectiveness of our proposed$\text{W}^{2}$AMSN, which surpasses state-of-the-art methods on most popular MR image SR benchmarks quantitatively and qualitatively. And our method still offers superior accuracy and adaptability on real MR images.
Haoqian Wang, Xiaowan Hu, Xiaole Zhao, Yulun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Memory Recall: A Simple Neural Network Training Framework Against Catastrophic Forgetting
abstract
It is widely acknowledged that biological intelligence is capable of learning continually without forgetting previously learned skills. Unfortunately, it has been widely observed that many artificial intelligence techniques, especially (deep) neural network (NN)-based ones, suffer from catastrophic forgetting problem, which severely forgets previous tasks when learning a new one. How to train NNs without catastrophic forgetting, which is termed continual learning, is emerging as a frontier topic and attracting considerable research interest. Inspired by memory replay and synaptic consolidation mechanism in brain, in this article, we propose a novel and simple framework termed memory recall (MeRec) for continual learning with deep NNs. In particular, we first analyze the feature stability across tasks in NN and show that NN can yield task stable features in certain layers. Then, based on this observation, we use a memory module to keep the feature statistics (mean and std) for each learned task. Based on the memory and statistics, we show that a simple replay strategy with Gaussian distribution-based feature regeneration can recall and recover the knowledge from previous tasks. Together with the weight regularization, MeRec preserves weights learned from previous tasks. Based on this simple framework, MeRec achieved leading performance with extremely small memory budget (only two feature vectors for each class) for continual learning on CIFAR-10 and CIFAR-100 datasets, with at least 50% accuracy drop reduction after several tasks compared to previous state-of-the-art approaches.
Baosheng Zhang, Haoqian Wang, Qionghai Dai
IEEE Trans. Neural Networks Learn. Syst.5
2021 Pseudo 3D Auto-Correlation Network for Real Image Denoising
abstract
The extraction of auto-correlation in images has shown great potential in deep learning networks, such as the self-attention mechanism in the channel domain and the self-similarity mechanism in the spatial domain. However, the realization of the above mechanisms mostly requires complicated module stacking and a large number of convolution calculations, which inevitably increases model complexity and memory cost. Therefore, we propose a pseudo 3D auto-correlation network (P3AN) to explore a more efficient way of capturing contextual information in image de-noising. On the one hand, P3AN uses fast 1D convolution instead of dense connections to realize criss-cross interaction, which requires less computational resources. On the other hand, the operation does not change the feature size and makes it easy to expand. It means that only a simple adaptive fusion is needed to obtain contextual information that includes both the channel domain and the spatial domain. Our method built a pseudo 3D auto-correlation attention block through 1D convolutions and a lightweight 2D structure for more discriminative features. Extensive experiments have been conducted on three synthetic and four real noisy datasets. According to quantitative metrics and visual quality evaluation, the P3AN shows great superiority and surpasses state-of-the-art image denoising methods.
Xiaowan Hu, Ruijun Ma 0001, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001, Haoqian Wang
CVPR7
2021 PoseDet: Fast Multi-Person Pose Estimation Using Pose Embedding
abstract
Current methods of multi-person pose estimation typically treat the localization and the association of body joints separately. It is convenient but inefficient, leading to additional computation and a waste of time. This paper, however, presents a novel framework PoseDet (Estimating Pose by Detection) to localize and associate body joints simultaneously at higher inference speed. Moreover, we propose the keypoint-aware pose embedding to represent an object in terms of the locations of its keypoints. The proposed pose embedding contains semantic and geometric information, allowing us to efficiently access discriminative and informative features. It is utilized for candidate classification and body joint localization in PoseDet, leading to robust predictions of various poses. This simple framework achieves an unprecedented speed and a competitive accuracy on the COCO benchmark compared with state-of-the-art methods. Extensive experiments on the CrowdPose benchmark show the robustness in the crowd scenes. Code is available at https://github.com/IIGROUP/PoseDet.
Weihao Xia 0001, Haoqian Wang, Yujiu Yang 0001
FG5
2021 Margin Loss Based On Adaptive Metric For Image Recognition
abstract
Cross-entropy (CE) loss is one of the most commonly used supervision in image recognition. However, the features trained by CE loss are not discriminative enough. There are many methods using the angular-softmax CE loss to learn angularly discriminative features. But these methods still use manually designed metric, which cannot deal with complicated distribution of high-dimensional features, and there are no appropriate inter-class restraint. Therefore, we propose a novel method to restrain the distribution of features in high-dimensional. We construct trainable centers, and design adaptive metric to express the distance between features. Specifically, we design AdaMetricLoss which can manipulate inter-class distance and intra-class distance simultaneously. We evaluate the effectiveness of AdaMetricLoss on Cifar10 and Cifar100 datasets, and our method shows preferable classification performance. We also visualize the discriminative distribution of features on MNIST, proves the proposed method is more suitable for image recognition.
Lei Song 0003, Xiaowan Hu, Haoqian Wang
ICIP4
2021 Pyramid Orthogonal Attention Network based on Dual Self-Similarity for Accurate Mr Image Super-Resolution
abstract
For magnetic resonance (MR) images sharing visual characteristics, the internal structure repetitions of different scales are considerable image-specific priors. Following the traditional algorithms, we try to combine external dataset-driven learning with the internal self-similarity for MR image super-resolution (SR). We propose a pyramid orthogonal attention network (POAN) based on dual self-similarity. On the one hand, by combining the point-similarity and the pyramid-similarity, sufficient spatial autocorrelation is explored to alleviate less training data limitation. On the other hand, the non-reduction channel attention mechanism maximizes inter-channel dependence. It increases the probability of the high-frequency region (e.g., structural textures and edges) being activated while suppresses low-frequency regions (e.g., background) adaptively. Out proposed POAN reconstructs the MR image under the guidance of pyramid orthogonal attention. Extensive experiments demonstrate that our method obtains the best results compared with state-of-the-art MR image SR methods quantitatively and visually.
Xiaowan Hu, Haoqian Wang, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001
ICME2
2021 ST-NAS: Efficient Optimization of Joint Neural Architecture and Hyperparameter
Jinhang Cai, Yimin Ou, Xiu Li 0001, Haoqian Wang
ICONIP (5)4
2021 Multi-Scale Selective Feedback Network with Dual Loss for Real Image Denoising
abstract
The feedback mechanism in the human visual system extracts high-level semantics from noisy scenes. It then guides low-level noise removal, which has not been fully explored in image denoising networks based on deep learning. The commonly used fully-supervised network optimizes parameters through paired training data. However, unpaired images without noise-free labels are ubiquitous in the real world. Therefore, we proposed a multi-scale selective feedback network (MSFN) with the dual loss. We allow shallow layers to access valuable contextual information from the following deep layers selectively between two adjacent time steps. Iterative refinement mechanism can remove complex noise from coarse to fine. The dual regression is designed to reconstruct noisy images to establish closed-loop supervision that is training-friendly for unpaired data. We use the dual loss to optimize the primary clean-to-noisy task and the dual noisy-to-clean task simultaneously. Extensive experiments prove that our method achieves state-of-the-art results and shows better adaptability on real-world images than the existing methods.
Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Yulun Zhang 0001
IJCAI4
2021 Learning to Generate Realistic Noisy Images via Pixel-level Noise-aware Adversarial Training
abstract
Existing deep learning real denoising methods require a large amount of noisy-clean image pairs for supervision. Nonetheless, capturing a real noisy-clean dataset is an unacceptable expensive and cumbersome procedure. To alleviate this problem, this work investigates how to generate realistic noisy images. Firstly, we formulate a simple yet reasonable noise model that treats each real noisy pixel as a random variable. This model splits the noisy image generation problem into two sub-problems: image domain alignment and noise domain alignment. Subsequently, we propose a novel framework, namely Pixel-level Noise-aware Generative Adversarial Network (PNGAN). PNGAN employs a pre-trained real denoiser to map the fake and real noisy images into a nearly noise-free solution space to perform image domain alignment. Simultaneously, PNGAN establishes a pixel-level adversarial training to conduct noise domain alignment. Additionally, for better noise fitting, we present an efficient architecture Simple Multi-scale Network (SMNet) as the generator. Qualitative validation shows that noise generated by PNGAN is highly similar to real noise in terms of intensity and distribution. Quantitative experiments demonstrate that a series of denoisers trained with the generated noisy images achieve state-of-the-art (SOTA) results on four real denoising benchmarks.
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Yulun Zhang 0001, Hanspeter Pfister, Donglai Wei 0001
NeurIPS3
2021 Semi-supervised self-growing generative adversarial networks for image recognition
Haoqian Wang
Multim. Tools Appl.2
2021 Bridging the Gap Between 2D and 3D Contexts in CT Volume for Liver and Tumor Segmentation
abstract
Automatic liver and tumor segmentation remain a challenging topic, which subjects to the exploration of 2D and 3D contexts in CT volume. Existing methods are either only focus on the 2D context by treating the CT volume as many independent image slices (but ignore the useful temporal information between adjacent slices), or just explore the 3D context lied in many little voxels (but damage the spatial detail in each slice). These factors lead an inadequate context exploration together for automatic liver and tumor segmentation. In this paper, we propose a novel full-context convolution neural network to bridge the gap between 2D and 3D contexts. The proposed network can utilize the temporal information along the Z axis in CT volume while retaining the spatial detail in each slice. Specifically, a 2D spatial network for intra-slice features extraction and a 3D temporal network for inter-slice features extraction are proposed separately and then are guided by the squeeze-and-excitation layer that allows the flow of 2D context and 3D temporal information. To address the severe class imbalance issue in the CT volume and meanwhile improve the segmentation performance, a loss function consisting of weighted cross-entropy and jaccard distance is proposed. During the network training, the 2D and 3D contexts are learned jointly in an end-to-end way. The proposed network achieves competitive results on the Liver Tumor Segmentation Challenge (LiTS) and the 3D-IRCADB datasets. This method should be a new promising paradigm to explore the contexts for liver and tumor segmentation.
Lei Song 0003, Haoqian Wang, Z. Jane Wang 0001
IEEE J. Biomed. Health Informatics2
2021 Precise No-Reference Image Quality Evaluation Based on Distortion Identification
abstract
The difficulty of no-reference image quality assessment (NR IQA) often lies in the lack of knowledge about the distortion in the image, which makes quality assessment blind and thus inefficient. To tackle such issue, in this article, we propose a novel scheme for precise NR IQA, which includes two successive steps, i.e., distortion identification and targeted quality evaluation. In the first step, we employ the well-known Inception-ResNet-v2 neural network to train a classifier that classifies the possible distortion in the image into the four most common distortion types, i.e., Gaussian white noise (WN), Gaussian blur (GB), jpeg compression (JPEG), and jpeg2000 compression (JP2K). Specifically, the deep neural network is trained on the large-scale Waterloo Exploration database, which ensures the robustness and high performance of distortion classification. In the second step, after determining the distortion type of the image, we then design a specific approach to quantify the image distortion level, which can estimate the image quality specially and more precisely. Extensive experiments performed on LIVE, TID2013, CSIQ, and Waterloo Exploration databases demonstrate that (1) the accuracy of our distortion classification is higher than that of the state-of-the-art distortion classification methods, and (2) the proposed NR IQA method outperforms the state-of-the-art NR IQA methods in quantifying the image quality.
Chenggang Yan 0001, Tong Teng, Yutao Liu 0002, Yongbing Zhang 0002, Haoqian Wang, Xiangyang Ji
ACM Trans. Multim. Comput. Commun. Appl.5
2020 Learning Delicate Local Representations for Multi-person Pose Estimation
Yuanhao Cai, Zhicheng Wang 0001, Zhengxiong Luo 0001, Binyi Yin, Angang Du, Haoqian Wang, Xiangyu Zhang 0005, Erjin Zhou, Jian Sun 0001
ECCV (3)6
2020 All-in-depth via Cross-baseline Light Field Camera
abstract
Light-field (LF) camera holds great promise for passive/general depth estimation benefited from high angular resolution, yet suffering small baseline for distanced region. While stereo solution with large baseline is superior to handle distant scenarios, the problem of limited angular resolution becomes bothering for near objects. Aiming for all-in-depth solution, we propose a cross-baseline LF camera using a commercial LF camera and a monocular camera, which naturally form a 'stereo camera' enabling compensated baseline for LF camera. The idea is simple yet non-trivial, due to the significant angular resolution gap and baseline gap between LF and stereo cameras.
Dingjian Jin, Anke Zhang, Gaochang Wu, Haoqian Wang, Lu Fang 0001
ACM Multimedia5
2020 Simple accurate model-based phase diversity phase retrieval algorithm for wavefront sensing in high-resolution optical imaging systems
abstract
In optical imaging systems, the aberration is an important factor that impedes realising diffraction‐limited imaging. Accurate wavefront sensing and control play important role in modern high‐resolution optical imaging systems nowadays. In this study, a simple model‐based phase retrieval algorithm is proposed for accurate efficient wavefront sensing with high dynamic range. In the authors’ algorithm, a wavefront is represented by the Zernike polynomials, and the Zernike coefficients are solved by the least‐squares‐based non‐linear optimisation method, i.e. the Lederberg–Marquardt algorithm, with multiple phase‐diversity images. The numerical results show that the proposed algorithm is capable of retrieving wavefront with a large dynamic range up to seven wavelength and robust to noise. In comparison, the proposed algorithm is more efficient than the existing model‐based technique and more accurate than existing Fourier ‐ transformation‐based iterative techniques.
Shun Qin, Yongbing Zhang 0002, Haoqian Wang, Wai Kin Chan
IET Image Process.3
2020 Color-Guided Depth Image Recovery With Adaptive Data Fidelity and Transferred Graph Laplacian Regularization
abstract
Depth images play an important role and are prevalently used in many computer vision and computational imaging tasks. However, due to the limitation of active sensing technology, the captured depth images in practice usually suffer from low resolution and noise, which prevents its further applications. To remedy this problem, in this paper, we first propose an adaptive data fidelity formulation to optimally generate each depth pixel from a mixture probability distribution, characterizing the similarity both in the depth map and the corresponding high-resolution guided color image. The proposed method is able to fit the distribution of the input depth signal as an optimization problem by maximizing the mixture probability. Furthermore, to promote the piecewise property that depth images exhibit, we propose a transferred graph Laplacian model as a regularization term, which is general and able to handle various depth recovery tasks such as super-resolution and denoising well. Specifically, each pixel within the recovered depth image is represented as a vertex in a graph with weights in connected edges representing the similarity between vertices. By minimizing the squared variations of the image signal, the task of depth image recovery can be converted to the problem of graph-based image filtering. Since the proposed graph Laplacian regularization model is able to fully exploit a priori information about the depth image, a much more accurate and robust estimation of the underlying depth can be obtained. Extensive experiment evaluations verify that the proposed method obtains recovered depth with higher quality in terms of both objective and subjective criteria, compared with most of the state-of-the-art methods.
Yongbing Zhang 0002, Yihui Feng, Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.6
2020 STAR: A Structure and Texture Aware Retinex Model
abstract
Retinex theory is developed mainly to decompose an image into the illumination and reflectance components by analyzing local image derivatives. In this theory, larger derivatives are attributed to the changes in reflectance, while smaller derivatives are emerged in the smooth illumination. In this paper, we utilize exponentiated local derivatives (with an exponent γ) of an observed image to generate its structure map and texture map. The structure map is produced by been amplified with γ > 1, while the texture map is generated by been shrank with γ < 1. To this end, we design exponential filters for the local derivatives, and present their capability on extracting accurate structure and texture maps, influenced by the choices of exponents γ. The extracted structure and texture maps are employed to regularize the illumination and reflectance components in Retinex decomposition. A novel Structure and Texture Aware Retinex (STAR) model is further proposed for illumination and reflectance decomposition of a single image. We solve the STAR model by an alternating optimization algorithm. Each sub-problem is transformed into a vectorized least squares regression, with closed-form solutions. Comprehensive experiments on commonly tested datasets demonstrate that, the proposed STAR model produce better quantitative and qualitative performance than previous competing methods, on illumination and reflectance decomposition, low-light image enhancement, and color correction. The code is publicly available at https://github.com/csjunxu/STAR.
Jun Xu 0019, Yingkun Hou, Dongwei Ren, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Haoqian Wang, Ling Shao 0001
IEEE Trans. Image Process.7
2020 An End-to-End Multi-Task Deep Learning Framework for Skin Lesion Analysis
abstract
Automatic skin lesion analysis of dermoscopy images remains a challenging topic. In this paper, we propose an end-to-end multi-task deep learning framework for automatic skin lesion analysis. The proposed framework can perform skin lesion detection, classification, and segmentation tasks simultaneously. To address the class imbalance issue in the dataset (as often observed in medical image datasets) and meanwhile to improve the segmentation performance, a loss function based on the focal loss and the jaccard distance is proposed. During the framework training, we employ a three-phase joint training strategy to ensure the efficiency of feature learning. The proposed framework outperforms state-of-the-art methods on the benchmarks ISBI 2016 challenge dataset towards melanoma classification and ISIC 2017 challenge dataset towards melanoma segmentation, especially for the segmentation task. The proposed framework should be a promising computer-aided tool for melanoma diagnosis.
Lei Song 0003, Jianzhe Lin, Z. Jane Wang 0001, Haoqian Wang
IEEE J. Biomed. Health Informatics4
2020 PID Controller-Based Stochastic Optimization Acceleration for Deep Neural Networks
abstract
Deep neural networks (DNNs) are widely used and demonstrated their power in many applications, such as computer vision and pattern recognition. However, the training of these networks can be time consuming. Such a problem could be alleviated by using efficient optimizers. As one of the most commonly used optimizers, stochastic gradient descent-momentum (SGD-M) uses past and present gradients for parameter updates. However, in the process of network training, SGD-M may encounter some drawbacks, such as the overshoot phenomenon. This problem would slow the training convergence. To alleviate this problem and accelerate the convergence of DNN optimization, we propose a proportional-integral-derivative (PID) approach. Specifically, we investigate the intrinsic relationships between the PID-based controller and SGD-M first. We further propose a PID-based optimization algorithm to update the network parameters, where the past, current, and change of gradients are exploited. Consequently, our proposed PID-based optimization alleviates the overshoot problem suffered by SGD-M. When tested on popular DNN architectures, it also obtains up to 50% acceleration with competitive accuracy. Extensive experiments about computer vision and natural language processing demonstrate the effectiveness of our method on benchmark data sets, including CIFAR10, CIFAR100, Tiny-ImageNet, and PTB. We have released the code at https://github.com/tensorboy/PIDOptimizer.
Haoqian Wang, Wangpeng An, Qingyun Sun, Jun Xu 0019, Lei Zhang 0006
IEEE Trans. Neural Networks Learn. Syst.1
2019 SPI-Optimizer: An Integral-Separated PI Controller for Stochastic Optimization
abstract
To overcome the oscillation problem in the classical momentum-based optimizer, recent work associates it with the proportional-integral (PI) controller, and artificially adds D term producing a PID controller. It suppresses oscillation with the sacrifice of introducing extra hyper-parameter. In this paper, we analyze that the fluctuation problem relates to the lag effect of the integral (I) term, and propose SPI-Optimizer, an integral-Separated PI controller based optimizer WITHOUT introducing extra hyper-parameter. It separates momentum term adaptively when the inconsistency of current and historical gradient direction occurs. Extensive experiments demonstrate that SPI-Optimizer generalizes well on popular network architectures to eliminate the oscillation, and owns competitive performance with faster convergence speed (up to 40% epochs reduction ratio) and more accurate classification result on MNIST, CIFAR10, and CIFAR100 (up to 27.5% error reduction ratio) than state-of-the-art methods.
Mengqi Ji, Haoqian Wang, Lu Fang 0001
ICIP4
2018 A PID Controller Approach for Stochastic Optimization of Deep Networks
abstract
Deep neural networks have demonstrated their power in many computer vision applications. State-of-the-art deep architectures such as VGG, ResNet, and DenseNet are mostly optimized by the SGD-Momentum algorithm, which updates the weights by considering their past and current gradients. Nonetheless, SGD-Momentum suffers from the overshoot problem, which hinders the convergence of network training. Inspired by the prominent success of proportional-integral-derivative (PID) controller in automatic control, we propose a PID approach for accelerating deep network optimization. We first reveal the intrinsic connections between SGD-Momentum and PID based controller, then present the optimization algorithm which exploits the past, current, and change of gradients to update the network parameters. The proposed PID method reduces much the overshoot phenomena of SGD-Momentum, and it achieves up to 50% acceleration on popular deep network architectures with competitive accuracy, as verified by our experiments on the benchmark datasets including CIFAR10, CIFAR100, and Tiny-ImageNet.
Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu 0019, Qionghai Dai, Lei Zhang 0006
CVPR2
2018 CrossNet: An End-to-End Reference-Based Super Resolution Network Using Cross-Scale Warping
Haitian Zheng, Mengqi Ji, Haoqian Wang, Yebin Liu, Lu Fang 0001
ECCV (6)3
2018 A Natural Shape-Preserving Stereoscopic Image Stitching
abstract
This paper presents a method for stereoscopic image stitching, which can make stereoscopic images look as natural as possible. Our method combines a constrained projective warp and a shape-preserving warp to reduce the projective distortion and the vertical disparity of the stitched image. In addition to provide a good alignment accuracy and maintain the consistency of input stereoscopic images, we add a specific restriction into the projective warp, which establishes the connection between target left and right images. To optimize the whole warp, a energy term is designed. It can constrain the shape of straight line and vertical disparity. Experimental results on a variety of stereoscopic images can ensure the efficiency of the proposed method.
Haoqian Wang, YaZing Zhou, Xingzheng Wang, Lu Fang 0001
ICASSP1
2018 Magnify-Net for Multi-Person 2D Pose Estimation
abstract
We propose a novel method for multi-person 2D pose estimation. Our model zooms in the image gradually, which we refer to as the Magnify-Net, to solve the bottleneck problem of mean average precision (mAP) versus pixel error. Moreover, we squeeze the network efficiently by an inspired design that increases the mAP while saving the processing time. It is a simple, yet robust, bottom-up approach consisting of one stage. The architecture is designed to detect the part position and their association jointly via two branches of the same sequential prediction process, resulting in a remarkable performance and efficiency rise. Our method outcompetes the previous state-of-the-art results on the challenging COCO key-points task and MPII Multi-Person Dataset.
Haoqian Wang, W. P. An, Xingzheng Wang, Lu Fang 0001, Jiahui Yuan
ICME1
2018 Fast, Robust, and Accurate Image Denoising via Very Deeply Cascaded Residual Networks
abstract
Patch based image modelings have shown great potential in image denoising. They mainly exploit the nonlocal self-similarity (NSS) of either input degraded images or clean natural ones when training models, while failing to learn the mappings between them. More seriously, these algorithms have very high time complexity and poor robustness when handling images with different noise variances and resolutions. To address these problems, in this paper, we propose very deeply cascaded residual networks (VDCRN) to build the precise relationships between the noisy images and their corresponding noise-free ones. It adopts a new residual unit with an identity skip connection (shortcut) to make training easy and improve generalization. The introduction of shortcut is helpful to avoid the problem of gradient vanishing and preserve more image details. By cascading three such residual units, we build the VDCRN to deploy deeper and larger convolutional networks. Based on such a residual network, our VDCRN achieves very fast speed and good robustness. Experimental results demonstrate that our model outperforms a lot of state-of-the-art denoising algorithms quantitively and qualitively.
Yongbing Zhang 0002, Xingzheng Wang, Haoqian Wang, Qionghai Dai
MMSP4
2018 Accurate saliency detection based on depth feature of 3D images
Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
Multim. Tools Appl.1
2018 Reversible data hiding in JPEG image based on DCT frequency and block selection
Dongdong Hou, Haoqian Wang, Weiming Zhang 0001, Nenghai Yu
Signal Process.2
2018 RevHashNet: Perceptually de-hashing real-valued image hashes for similarity retrieval
abstract
Image hashing has attracted increasing popularity in recent years. Some off-the-shelf image hashing methods are able to generate more compact and robust hashes for fast indexing and content-based similarity retrieval. However, the ability to infer original image contents from their real-valued image hashes has seldom been examined. Inherited from cryptographic hashing for image privacy protection, general image hashing is supposed to be a non-revertible function. Should there be a way to revert (or perceptually reconstruct) images from the corresponding real-valued image hashes? This paper explores the feasibility of perceptually image hashing reversion, and fill this gap by proposing a deep learning based framework, entitled RevHashNet. Given real-valued image hashes from certain image hashing methods, the proposed RevHashNet can automatically reconstruct perceptually similar images with respect to the original ones with high visual quality. Experiments and simulations on real image datasets support the de-hashing effectiveness of the proposed RevHashNet.
Hamid Palangi, Z. Jane Wang 0001, Haoqian Wang
Signal Process. Image Commun.4
2017 Learning Cross-scale Correspondence and Patch-based Synthesis for Reference-based Super-Resolution
Haitian Zheng, Mengqi Ji, Ziwei Xu 0001, Haoqian Wang, Yebin Liu, Lu Fang 0001
BMVC5
2017 An accurate saliency prediction method based on generative adversarial networks
abstract
In this paper, we propose a saliency prediction algorithm utilizing generative adversarial networks. The proposed system contains two parts: saliency network and adversarial networks. The saliency network is the basis for saliency prediction, which calculates an Euclidean cost function on the grayscale values between the predicted saliency map and the ground truth. In order to improve the accuracy of the algorithm, adversarial networks are subsequently utilized to extract the features of input data by coordinating the learning rates of the two sub-networks contained in the networks. Experimental results validate the high accuracy of the proposed approach compared with the state-of-the-art models on three public datasets, SALICON, MIT1003 and Cerf.
Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
ICIP2
2017 Exponential decay sine wave learning rate for fast deep neural network training
abstract
Most state-of-the-art results on image classification tasks were obtained by residual neural networks, which use stochastic gradient descent (SGD) with momentum for training. In most cases, the learning rate drops by a constant factor every pre-defined number of epochs. However, it is difficult and time-consuming to estimate how many epochs to drop the learning rate. To tackle this problem, cyclical learning rate is gaining popularity in gradient-based optimization to improve the convergence speed in accelerated gradient schemes. But cyclical learning rate scheme scans a broad range of learning rate, some of which are not suitable for deep neural network training. In this paper, we propose a simple yet effective exponential decay sine wave like learning rate technique for SGD to improve its convergence speed. In the training process, the learning rate would vary in sine wave way. While the maximum value of sine wave would decay exponentially along with training epochs. An ensemble of wide residual nets with our proposed learning scheme achieves 3.01% and 16.03% errors on CIFAR-10 and CIFAR-100 respectively. Furthermore, our proposed method uses far less number of epochs than most recent learning rate strategies, accelerating neural network training tremendously.
Wangpeng An, Haoqian Wang, Yulun Zhang 0001, Qionghai Dai
VCIP2
2017 Light-Field Depth Estimation via Epipolar Plane Image Analysis and Locally Linear Embedding
abstract
In this paper, we propose a novel method for 4D light-field (LF) depth estimation exploiting the special linear structure of an epipolar plane image (EPI) and locally linear embedding (LLE). Without high computational complexity, depth maps are locally estimated by locating the optimal slope of each line segmentation on the EPIs, which are projected by the corresponding scene points. For each pixel to be processed, we build and then minimize the matching cost that aggregates the intensity pixel value, gradient pixel value, spatial consistency, as well as reliability measure to select the optimal slope from a predefined set of directions. Next, a subangle estimation method is proposed to further refine the obtained optimal slope of each pixel. Furthermore, based on a local reliability measure, all the pixels are classified into reliable and unreliable pixels. For the unreliable pixels, LLE is employed to propagate the missing pixels by the reliable pixels based on the assumption of manifold preserving property maintained by natural images. We demonstrate the effectiveness of our approach on a number of synthetic LF examples and real-world LF data sets, and show that our experimental results can achieve higher performance than the typical and recent state-of-the-art LF stereo matching methods.
Yongbing Zhang 0002, Huijin Lv, Yebin Liu, Haoqian Wang, Xingzheng Wang, Qian Huang 0008, Xinguang Xiang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4
2016 Deep Convolutional Neural Network for Decompressed Video Enhancement
abstract
Block-wise intra/inter prediction, transformation and quantization used in block-based hybrid video coding will inevitably result in blocking artifacts, especially at the low bit rate. To address this problem, this paper employs a deep convolutional neural network (CNN) to approximate the reverse function of video compression, motived by the great success of deep learning in computer vision fields recently. The proposed method establishes an end-to-end mapping, represented as the CNN, which takes the decompressed frame as input and outputs the enhanced one. Employing numerous sequences compressed by H.264 and HEVC reference software, the proposed CNN learns the connections between the lossy frame and the original one in an implicit way under different quantization parameters (QP). Figure 1 shows the architecture of our CNN and the pipeline of the network training. We build our network with convolution layers and ReLU layer and the weights and biases of all the convolution layers in our model are updated by minimizing the loss using stochastic gradient descent with the standard backpropagation. We implement the CNN as a post-loop deblocking filter and explore varying CNN parameters for different QPs. Various experimental results demonstrate that the proposed method is able to significantly improve the quality of enhanced frames in terms of both objective and subjective criterions.
Rongqun Lin, Yongbing Zhang 0002, Haoqian Wang, Xingzheng Wang, Qionghai Dai
DCC3
2016 Depth Feature Based Accurate Saliency Detection for 3D Images
abstract
In this paper, we present an accurate saliency detection algorithm based on depth feature for 3D images. We first calculate depth cue based on the sharp regions' positions within the depth ranges. Then, the coarse saliency map is computed based on the background and location prior. Finally, we employ the contrast information in the coarse saliency map to obtain the final result. Experimental evaluation by comparison with existed methods verifies the effectiveness of our proposed algorithm in terms of precision, recall and F-Measure.
Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
PDCAT2
2016 Decompressed video enhancement via accurate regression prior
abstract
There is an increasing need for high-quality multimedia applications based on block-based hybrid video coding. Inevitably, the frame will degrade during the process of block-wise intra/inter prediction, transformation, and quantization, especially when the bit rate is low. In this paper, we propose an efficient decompressed video enhancement algorithm based on the adjusted anchored neighborhood regression (A+) method. In our work, first, we learn offline linear regressors, i.e. projection matrices from the decompressed to original video frames in the training phase. For grouping anchored neighborhoods more accurately, we adopt MI-KSVD rather than KSVD to learn the dictionary. Moreover, we exploit the mutual coherence between dictionary atoms and training samples to find the nearest neighbors. Second, in the enhancement phase, we boost the quality of input decompressed videos offline by learned regression priors. To verify the robustness of our enhancement method, extensive experiments are conducted. As shown in our experimental results, the proposed enhancement method yields superior performance both objectively and subjectively.
Yulun Zhang 0001, Yongbing Zhang 0002, Xingzheng Wang, Haoqian Wang, Qionghai Dai
VCIP5
2016 Fast and High Quality Highlight Removal From a Single Image
abstract
Specular reflection exists widely in photography and causes the recorded color deviating from its true value, thus, fast and high quality highlight removal from a single nature image is of great importance. In spite of the progress in the past decades in highlight removal, achieving wide applicability to the large diversity of nature scenes is quite challenging. To handle this problem, we propose an analytic solution to highlight removal based on an L2chromaticity definition and corresponding dichromatic model. Specifically, this paper derives a normalized dichromatic model for the pixels with identical diffuse color: a unit circle equation of projection coefficients in two subspaces that are orthogonal to and parallel with the illumination, respectively. In the former illumination orthogonal subspace, which is specular-free, we can conduct robust clustering with an explicit criterion to determine the cluster number adaptively. In the latter, illumination parallel subspace, a property called pure diffuse pixels distribution rule helps map each specular-influenced pixel to its diffuse component. In terms of efficiency, the proposed approach involves few complex calculation, and thus can remove highlight from high resolution images fast. Experiments show that this method is of superior performance in various challenging cases.
Jin-Li Suo, Dongsheng An, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Image Process.4
2015 A novel light field super-resolution framework based on hybrid imaging system
abstract
We propose a novel light field super-resolution framework based on hybrid imaging system, which combines two different imaging mechanisms: conventional imaging and current art-of-the-state imaging - light field imaging. We take advantage of conventional imaging in spatial resolution to make up light field and reconstruct a higher quality light field. In our method, we classify the points of the 3D scene: First, for highlight and occlusion, dictionary learning based interpolation is utilized, Second, for other areas, an improved patch matching algorithm is applied. As shown in experimental results, compared with four methods, which include the art-of-the-state algorithms, our approach is effective.
Judong Wu, Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
VCIP2
2015 Accurate image specular highlight removal based on light field imaging
abstract
Specular reflection removal is indispensable to many computer vision tasks. However, most existing methods fail or degrade in complex real scenarios for their individual drawbacks. Benefiting from the light field imaging technology, this paper proposes a novel and accurate approach to remove specularity and improve image quality. We first capture images with specularity by the light field camera (Lytro ILLUM). After accurately estimating the image depth, a simple and concise threshold strategy is adopted to cluster the specular pixels into "unsaturated" and "saturated" category. Finally, a color variance analysis of multiple views and a local color refinement are individually conducted on these two categories to recover diffuse color information. Experimental evaluation by comparison with existed methods verifies the effectiveness of our proposed algorithm.
Chenxue Xu, Xingzheng Wang, Haoqian Wang, Yongbing Zhang 0002
VCIP3
2015 Adaptive local nonparametric regression for fast single image super-resolution
abstract
We propose a fast single image super-resolution algorithm based on adaptive local nonparametric regression. Making use of dictionary learning and regression, we learn multiple projection matrices mapping low-resolution features to their corresponding high-resolution ones directly. Different from previous linear regression that needs some constant parameters, our method would not use extra parameters for regression. We use the mutual coherence between dictionary atom and low-resolution feature as a label to reconstruct more sophisticated high-resolution feature. As we use the same form of mutual coherence as labels in both training and testing phases, our method would lead to an adaptive local linear regression model. Moreover, we investigate the statistical property of the dictionary atoms from the training features. Utilizing the learned statistical priors, our method would not only obtain more useful dictionary atoms, but also further decrease the computational time. As shown in our experimental results, the proposed method yields high-quality super-resolution images quantitatively and visually against state-of-the-art methods.
Yulun Zhang 0001, Yongbing Zhang 0002, Jian Zhang 0018, Haoqian Wang, Xingzheng Wang, Qionghai Dai
VCIP4
2015 Structuring Lecture Videos by Automatic Projection Screen Localization and Analysis
abstract
We present a fully automatic system for extracting the semantic structure of a typical academic presentation video, which captures the whole presentation stage with abundant camera motions such as panning, tilting, and zooming. Our system automatically detects and tracks both the projection screen and the presenter whenever they are visible in the video. By analyzing the image content of the tracked screen region, our system is able to detect slide progressions and extract a high-quality, non-occluded, geometrically-compensated image for each slide, resulting in a list of representative images that reconstruct the main presentation structure. Afterwards, our system recognizes text content and extracts keywords from the slides, which can be used for keyword-based video retrieval and browsing. Experimental results show that our system is able to generate more stable and accurate screen localization results than commonly-used object tracking methods. Our system also extracts more accurate presentation structures than general video summarization methods, for this specific type of video.
Kai Li 0016, Jue Wang 0001, Haoqian Wang, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Depth estimation from a single defocused image using multi-scale kernels
abstract
Depth estimation from defocus (DFD) has proved to be an efficient way to recover depth information based on the blur amount of defocus images. By introducing a multi-scale strategy into DFD, a novel depth estimation method from a single defocused image is proposed in this paper. The original input image is re-blurred using Gaussian kernels with different scale parameters, then a robust estimation of defocus blur amount at edge locations could be obtained by calculating the gradient magnitude ratio according to the original and re-blurred images. Dense defocus maps are generated via global interpolation and refinement and hence depth can be obtained under certain camera parameters. Experimental results demonstrate the effectiveness of the proposed method on obtaining high quality dense defocus and depth maps.
Haoqian Wang, Yushi Tian, Xingzheng Wang
ICARCV1
2014 Automatic foreground extraction in video
abstract
This paper presents an automatic and efficient system for extracting dynamic objects of interest from videos. We take advantage of a saliency map and an optimization-based segmentation algorithm to extract the foreground objects automatically in some key frames. Then, the segmentation results in those key frames are propagated to other frames via an error map-based propagation scheme. Finally, a Bayesian matting-based refinement approach is employed to to handle the topology changes. Experiments show that our system is able to generate high quality results at a low computation cost.
Haoqian Wang, Kai Li 0016, Yongbing Zhang 0002, Lei Zhang 0006
ICASSP1
2014 Real-time air quality estimation based on color image processing
abstract
This paper address the problem of efficient, realtime estimation of the particulate mass concentration, exactly PM2.5 (particles with aerodynamic diameters less than 2.5 μm) from a superb view image. And the proposed method is to achieve high degree of accuracy at the cost of only modest user's effort by analyzing the relationship between the PM2.5 and the degradation of the observed image. With the fitting algorithm with experimental data, the PM2.5 could be real-time estimated by a general camera with little artificial participation, and the correlation coefficient produced by our data set and the standard observation will be as high as 0.8219, as the MSE (Mean Squared Error) value 51.2324 μg/m3.
Haoqian Wang, Xin Yuan 0002, Xingzheng Wang, Yongbing Zhang 0002, Qionghai Dai
VCIP1
2013 A novel depth propagation algorithm with color guided motion estimation
abstract
Depth propagation is an effective and efficient way to produce depth maps for a video sequence. Motion estimation in most existing depth propagation schemes is performed only based on the estimated depth maps without consideration for color information. This paper presents a novel key frame depth propagation algorithm combining bilateral filtering and motion estimation. A color guided motion estimation process is proposed by taking both color and depth information into account when estimating the motion vectors. In addition, a bidirectional propagation strategy is adopted to reduce the accumulation of depth errors. Experimental results show that the proposed algorithm outperforms most of the existing techniques in obtaining high quality depth maps indicating a better effect of the synthesized stereoscopic video.
Haoqian Wang, Yushi Tian, Yongbing Zhang 0002
VCIP1
2013 Effective stereo matching using reliable points based graph cut
abstract
In this paper, we propose an effective stereo matching algorithm using reliable points and region-based graph cut. Firstly, the initial disparity maps are calculated via local windowbased method. Secondly, the unreliable points are detected according to the DSI(Disparity Space Image) and the estimated disparity values of each unreliable point are obtained by considering its surrounding points. Then, the scheme of reliable points is introduced in region-based graph cut framework to optimize the initial result. Finally, remaining errors in the disparity results are effectively handled in a multi-step refinement process. Experiment results show that the proposed algorithm achieves a significant reduction in computation cost and guarantee high matching quality.
Haoqian Wang, Yongbing Zhang 0002, Lei Zhang 0006
VCIP1
2013 A Progressive Tri-level Segmentation Approach for Topology-Change-Aware Video Matting
abstract
Abstract Previous video matting approaches mostly adopt the “binary segmentation + matting” strategy, i.e., first segment each frame into foreground and background regions, then extract the fine details of the foreground boundary using matting techniques. This framework has several limitations due to the fact that binary segmentation is employed. In this paper, we propose a new supervised video matting approach. Instead of applying binary segmentation, we explicitly model segmentation uncertainty in a novel tri‐level segmentation procedure. The segmentation is done progressively, enabling us to handle difficult cases such as large topology changes, which are challenging to previous approaches. The tri‐level segmentation results can be naturally fed into matting techniques to generate the final alpha mattes. Experimental results show that our system can generate high quality results with less user inputs than the state‐of‐theart methods.
Jinlong Ju, Jue Wang 0001, Yebin Liu, Haoqian Wang, Qionghai Dai
Comput. Graph. Forum4
2013 Up-sampling oriented frame rate reduction
Yongbing Zhang 0002, Haoqian Wang, Debin Zhao
Signal Process. Image Commun.2
2013 Stereo Interleaving Video Coding With Content Adaptive Image Subsampling
abstract
Stereo interleaving video coding, in which both left and right view frames are subsampled into half size and multiplexed into one single frame before being encoded by a traditional 2-D video encoder, is an efficient encoding scenario for stereoscopic video. Many existing stereo interleaving video coding methods subsample each frame by utilizing fixed subsampling filter coefficients. Such methods are easy to implement; however, the varying property of the frame signal is ignored. By jointly considering the influences of subsampling and compression, a rate and distortion analysis about stereo interleaving video coding is proposed. The final distortion in stereo interleaving video coding is the summation of errors caused by subsampling (causing distortion between subsampling-interpolated image and the original full resolution one) and by quantization during compression. Based on the provided rate distortion analysis, a content adaptive image subsampling (CAIS) is also proposed. In CAIS, the half-size frames are generated by the optimal subsampling filters, which are calculated based on frame contents and the targeted interpolation coefficients. Experimental results demonstrate that the proposed CAIS is able to greatly improve compression efficiency of stereo interleaving video coding.
Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.3
2013 Statistical Analysis of Tongue Images for Feature Extraction and Diagnostics
abstract
In this paper, an in-depth analysis on the statistical distribution characteristics of human tongue color that aims to propose a mathematically described tongue color space for diagnostic feature extraction is presented. Three characteristics of tongue color space, i.e., tongue color gamut that defines the range of colors, color centers of 12 tongue color categories, and color distribution of typical image features in the tongue color gamut, are elaborately investigated in this paper. Based on a large database, which contains over 9000 tongue images collected by a specially designed noncontact colorimetric imaging system using a digital camera, the tongue color gamut is established in the CIE chromaticity diagram by an innovatively proposed color gamut boundary descriptor using one-class SVM algorithm. Thereafter, centers of 12 tongue color categories are defined accordingly. Furthermore, color distributions of several typical tongue features, such as red points and petechial points, are obtained to build a relationship between the tongue color space and color distributions of various tongue features. With the obtained tongue color space, a new color feature extraction method is proposed for diagnostic classification purposes, with experimental results validating its effectiveness.
Xingzheng Wang, Bob Zhang 0001, Zhimin Yang, Haoqian Wang, David Zhang 0001
IEEE Trans. Image Process.4
2012 Content Adaptive Subsampling for Stereo Interleaving Video Coding
abstract
Stereo interleaving video coding receives considerable attention due to its desirable property of being compatible with 2D video coding standards. The errors caused by sub sampling (causing distortion between subsampling interpolated image and the original full resolution one) and by quantization during compression lead to the final distortion in stereo interleaving video coding. In this paper, the rate and distortion analysis in stereo interleaving video coding is provided. It proves that appropriate sub sampling in stereo interleaving video coding is able to obtain good compression performance. Subsequently, a content adaptive sub sampling (CAS) is proposed. In CAS, the half resolution frames are generated by decimation, where the down sampling filter coefficients are calculated based on frame contents and the targeted interpolation coefficients. Experiment results demonstrate that the CAS is able to achieve high compression efficiency of stereo interleaving encoding scheme for stereoscopic videos.
Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Lei Zhang 0006, Qionghai Dai
DCC3
2012 An effective mode decision method for multi-view video coding
abstract
Multi-view video coding (MVC) is an essential technique in video processing, while the huge computational complexity of mode decision limits its widespread application. This paper presents a fast mode decision algorithm to significantly reduce the computational burden. The mode is predicted through various schemes, and the best prediction is determined by a new proposed reliable method. The best coding mode can be obtained through the given accurate prediction. Comparing with existed mode prediction decision algorithms, our algorithm will present precise prediction. Simulation results shows the effectiveness in reducing computational complexity without any obvious PSNR loss.
Chengli Du, Haoqian Wang
ICARCV2
2012 Stereo matching using graph cuts: A 3D-Hough transformation approach
abstract
Based on the assumption that the scene structure consists of a set of planar surface patches and the algorithm of planar graph cuts, a novel stereo matching algorithm is presented in this paper using 3D-Hough transformation. Unlike current stereo matching algorithm, our method provides a quite fast and robust plane fitting resolution that can enhance the reliability of each disparity model. Experimental results demonstrate the effectiveness of the proposed method in obtaining precise and smooth disparity image.
Haoqian Wang
ICARCV2
2011 Up-sampling Dependent Frame Rate Reduction for Low Bit-Rate Video Coding
abstract
Summary form only given. In low bit rate video coding, the frame rate of input sequence can be reduced to the half or even smaller portion by skipping or deleting frames before compression, and then the temporal resolution is restored via up-sampling at the decoder side. Numerous algorithms have been developed to address the problem of temporal resolution improvement. Actually, the quality of up-sampled frames depends on not only the performance of up-sampling method but also the information maintained in the down-sampled video sequence. To improve the quality of up-sampled frames and smooth the quality between the up-sampled and decompressed frames, this paper proposes an up-sampling dependent frame rate reduction, which is shown in Fig. 1. The proposed low bit rate video coding scheme is composed of up-sampling dependent frame rate reduction, compression, decompression and up-sampling components. The proposed frame rate reduction method is hinged to the temporal up-sampling. It is noted that there is a feedback between frame rate reduction and up-sampling in the proposed up-sampling dependent frame rate reduction, of which the goal is to obtain a down-sampled sequence maintaining more information about the frames to be up-sampled at the decoder side.
Yongbing Zhang 0002, Haoqian Wang, Debin Zhao
DCC2
2004 Kalman filtering for descriptor systems with current and delayed measurements
abstract
A class of discrete-time Kalman filtering problem for the descriptor time-varying systems with current and delayed measurements is considered. Using the known maximum likelihood (ML) estimation results and the method of measurements reorganization, the optimal Kalman filter and corresponding Riccati equations for descriptor systems involving current and delayed measurements are derived. Our solution does not require system augmentation or system transformation, and the estimator is given in terms of two Riccati equations of the same order as that of the system state. A simple algorithm is presented for the problem.
Haoqian Wang, Huanshui Zhang, Guangren Duan 0001
ICARCV1