Junhui Hou

dblp:122/2673 · DBLP profile ↗
← Back
246ranked-venue papers
19as first author
167since 2021 · last 2026
0000-0003-3431-2021ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 159 · 16 first-author · 106 since 2021Artificial intelligence and machine learning · 95 · 77 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 since 2021Systems, architecture and hardware · 3 · 3 first-authorComputer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 DiCaP: Distribution-Calibrated Pseudo-labeling for Semi-Supervised Multi-Label Learning
abstract
Semi-supervised multi-label learning (SSMLL) aims to address the challenge of limited labeled data in multi-label learning (MLL) by leveraging unlabeled data to improve the model’s performance. While pseudo-labeling has become a dominant strategy in SSMLL, most existing methods assign equal weights to all pseudo-labels regardless of their quality, which can amplify the impact of noisy or uncertain predictions and degrade the overall performance. In this paper, we theoretically verify that the optimal weight for a pseudo-label should reflect its correctness likelihood. Empirically, we observe that on the same dataset, the correctness likelihood distribution of unlabeled data remains stable, even as the number of labeled training samples varies. Building on this insight, we propose Distribution-Calibrated Pseudo-labeling (DiCaP), a correctness-aware framework that estimates posterior precision to calibrate pseudo-label weights. We further introduce a dual-thresholding mechanism to separate confident and ambiguous regions: confident samples are pseudo-labeled and weighted accordingly, while ambiguous ones are explored by unsupervised contrastive learning. Experiments conducted on multiple benchmark datasets verify that our method achieves consistent improvements, surpassing state-of-the-art methods by up to 4.27%.
Bo Han 0017, Zhuoming Li, Yaxin Hou, Hui Liu 0032, Junhui Hou, Yuheng Jia
AAAI6
2026 Towards Better IncomLDL: We Are Unaware of Hidden Labels in Advance
abstract
Label distribution learning (LDL) is a novel paradigm that describe the samples by label distribution of a sample. However, acquiring LDL dataset is costly and time-consuming, which leads to the birth of incomplete label distribution learning (IncomLDL). All the previous IncomLDL methods set the description degrees of "missing" labels in an instance to 0, but remains those of other labels unchanged. This setting is unrealistic because when certain labels are missing, the degrees of the remaining labels will increase accordingly. We fix this unrealistic setting in IncomLDL and raise a new problem: LDL with hidden labels (HidLDL), which aims to recover a complete label distribution from a real-world incomplete label distribution where certain labels in an instance are omitted during annotation. To solve this challenging problem, we discover the significance of proportional information of the observed labels and capture it by an innovative constraint to utilize it during the optimization process. We simultaneously use local feature similarity and the global low-rank structure to reveal the mysterious veil of hidden labels. Moreover, we **theoretically** give the recovery bound of our method, proving the feasibility of our method in learning from hidden labels. Extensive recovery and predictive experiments on various datasets prove the superiority of our method to state-of-the-art LDL and IncomLDL methods.
Jiecheng Jiang, Hui Liu 0032, Junhui Hou, Yuheng Jia
AAAI5
2026 ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering
abstract
Typical deep clustering methods, while achieving notable progress, can only provide one clustering result per dataset. This limitation arises from their assumption of a fixed underlying data distribution, which may fail to meet user needs and provide unsatisfactory clustering outcomes. Our work investigates how multi-modal large language models (MLLMs) can be leveraged to achieve user-driven clustering, emphasizing their adaptability to user-specified semantic requirements. However, directly using MLLM output for clustering has risks for producing unstructured and generic image descriptions instead of feature-specific and concrete ones. To address these issues, our method first discovers that MLLMs' hidden states of text tokens are strongly related to the corresponding features, and leverages these embeddings to perform clusterings from any user-defined criteria. We also employ a lightweight clustering head augmented with pseudo-label learning, significantly enhancing clustering accuracy. Extensive experiments demonstrate its competitive performance on diverse datasets and metrics.
Yuheng Jia, Hui Liu 0032, Junhui Hou
AAAI4
2026 Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Chaofan Gan, Mingxi Lv, Ziran Qin, Li Shen 0008, Junhui Hou, Weiyao Lin
Int. J. Comput. Vis.9
2026 Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?
abstract
Abstract Cross-modal contrastive distillation has recently been explored for learning effective 3D representations. However, existing methods focus primarily on modality-shared features, neglecting the modality-specific features during the pre-training process, which leads to suboptimal representations. In this paper, we theoretically analyze the limitations of current contrastive methods for 3D representation learning and propose a new framework, namely CMCR (Cross-Modal Comprehensive Representation Learning), to address these shortcomings. Our approach improves upon traditional methods by better integrating both modality-shared and modality-specific features. Specifically, we introduce masked image modeling and occupancy estimation tasks to guide the network in learning more comprehensive modality-specific features. Furthermore, we introduce a novel multi-modal unified codebook that learns an embedding space shared across different modalities. Besides, we propose geometry-enhanced masked image modeling to further boost 3D representation learning. Extensive experiments demonstrate that our method mitigates the challenges faced by traditional approaches and consistently outperforms existing image-to-LiDAR contrastive distillation methods in downstream tasks. Code will be available at https://github.com/Eaphan/CMCR .
Yifan Zhang 0036, Junhui Hou
Int. J. Comput. Vis.2
2026 FlexPara: Flexible Neural Surface Parameterization
abstract
Surface parameterization is a fundamental geometry processing task, laying the foundations for the visual presentation of 3D assets and numerous downstream shape analysis scenarios. Conventional parameterization approaches demand high-quality mesh triangulation and are restricted to certain simple topologies unless additional surface cutting and decomposition are provided. In practice, the optimal configurations (e.g., type of parameterization domains, distribution of cutting seams, number of mapping charts) may vary drastically with different surface structures and task characteristics, thus requiring more flexible and controllable processing pipelines. To this end, this paper introduces FlexPara, an unsupervised neural optimization framework to achieve both global and multi-chart surface parameterizations by establishing point-wise mappings between 3D surface points and adaptively-deformed 2D UV coordinates. We ingeniously design and combine a series of geometrically-interpretable sub-networks, with specific functionalities of cutting, deforming, unwrapping, and wrapping, to construct a bi-directional cycle mapping framework for global parameterization without the need for manually specified cutting seams. Furthermore, we construct a multi-chart parameterization framework with adaptively-learned chart assignment. Extensive experiments demonstrate the universality, superiority, and inspiring potential of our neural surface parameterization paradigm.
Qijian Zhang, Junhui Hou, Jiazhi Xia, Wenping Wang 0001, Ying He 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Semantic Contrast for Domain-Robust Underwater Image Quality Assessment
abstract
Underwater image quality assessment (UIQA) is hindered by complex degradation and domain shifts across aquatic environments. Existing no-reference IQA methods rely on costly and subjective mean opinion scores (MOS), which limit their generalization to unseen domains. To overcome these challenges, we propose SCUIA, an unsupervised UIQA framework leveraging semantic contrastive learning for quality prediction without human annotations. Specifically, we introduce a vision-language contrastive learning strategy that aligns image features with textual embeddings in a unified semantic space, capturing implicit degradation-quality correlations. We further enhance quality discrimination with a hierarchical contrastive learning mechanism that combines image-specific statistical priors and semantic prompts. A triplet-based inter-group contrastive loss explicitly models relative quality relationships. To tackle cross-domain variations, we develop an unsupervised domain adaptation module that uses local statistical features to guide CLIP fine-tuning to disentangle domain-invariant quality representations from domain-specific noise. This enables zero-shot cross-domain quality prediction without labeled data. Extensive experiments on public UIQA benchmarks demonstrate significant improvements over existing methods, highlighting superior generalization and domain adaptability.
Jingchun Zhou, Chunjiang Liu, Qiuping Jiang, Xianping Fu, Junhui Hou, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Progressive local self-attention for content-aligned super-resolution
Detian Huang, Xiancheng Zhu, Fei Shen 0004, Taotao Lai, Huanqiang Zeng, Junhui Hou
Pattern Recognit.7
2026 Exploring non-local spatial-angular correlations with a hybrid Mamba-Transformer framework for light field super-resolution
Haosong Liu, Xiancheng Zhu, Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Junhui Hou
Pattern Recognit.6
2026 Point cloud compression using graph neural networks
Xiangjie Zhang, Huanqiang Zeng, Xin-Rong Gong, Junhui Hou, Jianqing Zhu
Pattern Recognit.5
2026 Feedforward Compression of Static and Streamable 3D Gaussian Splatting
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity novel view synthesis, yet their substantial storage cost remains a major barrier to practical deployment. Although several compression techniques have been explored, they share a common limitation:each existing 3DGS requires per-scene optimization to achieve compression, making the compressionslow and inefficient. In this work, we present Fast Compression of 3D Gaussian Splatting (FCGS), an optimization-free approach that compresses existing 3DGS in a single feed-forward pass, reducing compression time from minutes to seconds. To enhance compression efficiency, we design a multi-path entropy module that routes Gaussian attributes through separate entropy-constrained paths, achieving a better trade-off between size and fidelity. Furthermore, we introduce both inter- and intra-Gaussian context models to effectively remove redundancies for the unstructured Gaussian representation. Experimental results show that FCGS achieves over 20× compression while maintaining high fidelity, outperforming most State-of-The-Art (SoTA) per-scene optimization-based methods. Beyond static scenes, we further extend FCGS to a streamable setting which eliminates redundant temporal information, demonstrating its strong potential for compressing streamable 3DGS data.
Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Junhui Hou, Mehrtash Harandi, Jianfei Cai 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 RoSe: Robust Self-Supervised Stereo Matching Under Adverse Weather Conditions
abstract
Recent self-supervised stereo matching methods have made significant progress, but their performance significantly degrades under adverse weather conditions such as night, rain, and fog. We identify two primary weaknesses contributing to this performance degradation. First, adverse weather introduces noise and reduces visibility, making CNN-based feature extractors struggle with degraded regions like reflective and textureless areas. Second, these degraded regions can disrupt accurate pixel correspondences, leading to ineffective supervision based on the photometric consistency assumption. To address these challenges, we propose injecting robust priors derived from the visual foundation model into the CNN-based feature extractor to improve feature representation under adverse weather conditions. We then introduce scene correspondence priors to construct robust supervisory signals rather than relying solely on the photometric consistency assumption. Specifically, we create synthetic stereo datasets with realistic weather degradations. These datasets feature clear and adverse image pairs that maintain the same semantic context and disparity, preserving the scene correspondence property. With this knowledge, we propose a robust self-supervised training paradigm, consisting of two key steps: robust self-supervised scene correspondence learning and adverse weather distillation. Both steps aim to align underlying scene results from clean and adverse image pairs, thus improving model disparity estimation under adverse weather effects. Extensive experiments demonstrate the effectiveness and versatility of our proposed solution, which outperforms existing state-of-the-art self-supervised methods. Codes are available at https://github.com/cocowy1/RoSe-Robust-Self-supervised-Stereo-Matching-under-Adverse-Weather-Conditions.
Yun Wang 0053, Junjie Hu 0003, Junhui Hou, Chenghao Zhang 0003, Renwei Yang, Dapeng Oliver Wu
IEEE Trans. Circuits Syst. Video Technol.3
2026 Bridging the Gap Between Implicit and Explicit Representation for Efficient Image Compression
abstract
Neural Image Compression (NIC) has achieved superior compression performance by modeling images as implicit feature representations, yet its practical deployment is severely hindered by computational overhead. Recently, GaussianImage was proposed as a computationally efficient explicit image representation paradigm, which renders images from 2D Gaussians. However, it suffers from inferior compression performance relative to mainstream NIC frameworks, mainly due to low compressibility and representation capability of explicit Gaussian parameters. To this end, we propose Pixel-Aligned Generalized 2D Splatting (PA-G2DS) as a computation-efficient and compression-friendly image representation format. Specifically, we deploy a learnable rendering function with implicit coefficients to enhance the reconstructed image quality and improve the compression ratio over explicit Gaussian coefficients. Incorporating the proposed PA-G2DS as a computationally efficient decoder, we further develop a suite of image codecs optimized for either compression ratios or flexible deployment scenarios. Experiments prove that the proposed codecs could achieve 30ms compression latency and millisecond-level decompression latency, reducing the performance and efficiency gap between implicit and explicit image representation. Furthermore, the proposed codec opens potential applications for NIC such as JPEG-like sequential decompression and random-access during decompression.
Wenrui Dai, Chern Hong Lim, Carl J. Debono, Thittaporn Ganokratanaa, Junhui Hou, Weiyao Lin
IEEE Trans. Circuits Syst. Video Technol.7
2026 ResFlow: Fine-Tuning Residual Optical Flow for Event-Based High Temporal Resolution Motion Estimation
Qianang Zhou, Junhui Hou, Yongjian Deng, Youfu Li 0001, Junlin Xiong
IEEE Trans. Circuits Syst. Video Technol.3
2026 Reflectance Prediction-Based Knowledge Distillation for Robust 3D Object Detection in Compressed Point Clouds
abstract
Regarding intelligent transportation systems, low-bitrate transmission via lossy point cloud compression is vital for facilitating real-time collaborative perception among connected agents, such as vehicles and infrastructures, under restricted bandwidth. In existing compression transmission systems, the sender lossily compresses point coordinates and reflectance to generate a transmission code stream, which faces transmission burdens from reflectance encoding and limited detection robustness due to information loss. To address these issues, this paper proposes a 3D object detection framework with reflectance prediction-based knowledge distillation (RPKD). We compress point coordinates while discarding reflectance during low-bitrate transmission, and feed the decoded non-reflectance compressed point clouds into a student detector. The discarded reflectance is then reconstructed by a geometry-based reflectance prediction (RP) module within the student detector for precise detection. A teacher detector with the same structure as the student detector is designed for performing reflectance knowledge distillation (RKD) and detection knowledge distillation (DKD) from raw to compressed point clouds. Our cross-source distillation training strategy (CDTS) equips the student detector with robustness to low-quality compressed data while preserving the accuracy benefits of raw data through transferred distillation knowledge. Experimental results on the KITTI and DAIR-V2X-V datasets demonstrate that our method can boost detection accuracy for compressed point clouds across multiple code rates. We will release the code publicly at https://github.com/HaoJing-SX/RPKD.
Anhong Wang, Yifan Zhang 0036, Donghan Bu, Junhui Hou
IEEE Trans. Image Process.5
2026 Enhancing Underwater Light Field Images via Global Geometry-Aware Diffusion Process
abstract
This work studies the challenging problem of acquiring high-quality underwater images via 4-D light field (LF) imaging. To this end, we propose GeoDiff-LF, a novel diffusion-based framework built upon SD-Turbo to enhance underwater 4-D LF imaging by leveraging its spatial-angular structure. GeoDiff-LF consists of three key adaptations: 1) a modified U-Net architecture with convolutional and attention adapters to model geometric cues, 2) a geometry-guided loss function using tensor decomposition and progressive weighting to regularize global structure, and 3) an optimized sampling strategy with noise prediction to improve efficiency. By integrating diffusion priors and LF geometry, GeoDiff-LF effectively mitigates color distortion in underwater scenes. Extensive experiments demonstrate that our framework outperforms existing methods across both visual fidelity and quantitative performance, advancing the state-of-the-art in enhancing underwater imaging. The code will be publicly available at https://github.com/linlos1234/GeoDiff-LF.
Yuji Lin, Qian Zhao 0002, Zongsheng Yue, Junhui Hou, Deyu Meng
IEEE Trans. Image Process.4
2026 FD-SCU: Frequency Decomposition-Based Spectrum Collaborative Upsampling for Point Cloud Color Attribute
abstract
Existing point cloud color upsampling methods typically treat color upsampling as an interpolation problem within a local color or implicit feature domain. This largely overlooks the ability of the frequency domain to capture color correlations in local point sets. To address this limitation, we propose a spectrum collaborative strategy that uses frequency decomposition on voxel blocks (VBs) to enhance point cloud color reconstruction. We first voxelize the low-resolution (LR) color point cloud to generate multiple VBs and introduce a virtual filling strategy that adaptively assigns colors to empty voxels in each VB, ensuring that the irregularly distributed color information fully occupies the VB. We then apply the discrete cosine transform, known for its strong frequency-domain representation of locally smooth signals, to each color-filled VB to obtain frequency coefficients. These frequency coefficients are separated into high-frequency (HF) and low-frequency (LF) components. The LF coefficients, together with the LR color point cloud, are fed into a multi-scale cross-domain feature extraction module to capture deep features. Next, a Gaussian perturbation-based feature expansion generates upsampled color features, which are used to regress a coarse upsampled color point cloud. Finally, a high-frequency-guided residual refinement module uses the HF coefficients to refine the coarse upsampled result and produce a high-fidelity color point cloud. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art methods. Our code will be publicly available at https://github.com/wangwenchaoxx/FD-SCU.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan, Junhui Hou
IEEE Trans. Image Process.6
2026 RAW-CLIP Fusion: Unleashing Semantic-Aware Denoising for Sensor-Agnostic Low-Light Imaging
abstract
Denoising images captured under extreme low-light conditions remains a persistent challenge in computational photography, primarily due to low signal-to-noise ratios and sensor-specific noise characteristics. These variations often require per-sensor noise calibration to achieve effective denoising. Although recent calibration-free methods aim to reduce this dependency through synthetic noise modeling or few-shot fine-tuning, their performance often degrades in extreme low-light scenarios across different sensors due to mismatches between synthetic and real-world noise. To address this gap, we introduce CLIP-Guided Denoising (CLD), the first framework to leverage large-scale vision models pretrained on sRGB images for cross-domain feature fusion, effectively guiding RAW image denoising across diverse sensors. Although not trained on RAW data, CLIP embeddings offer semantically robust and noise-invariant features that help guide the denoising network to focus on the underlying image content rather than fitting to specific noise distributions. Extensive experiments on the SID and ELD datasets demonstrate that CLD achieves state-of-the-art performance in calibration-free settings, significantly outperforming prior methods under extreme low-light conditions and achieving robust generalization across unseen sensor domains.
Mingde Qiao, Junjun Jiang, Zhanghong Zhao, Junhui Hou, Jiayi Ma 0001
IEEE Trans. Image Process.5
2026 PCCRender: Joint Learning of Point Cloud Compression and Gaussian Splatting Rendering
abstract
Point cloud is a crucial 3D representation that plays significant role in fields such as VR/AR and digital museums. With the increasing amount of point cloud data, many compression and rendering methods are proposed, which often focus on either 3D or 2D visual quality. However, with the rapid advancements in display technology, there is a growing expectation for improved 2D and 3D visual quality while using lower bandwidth. To address these challenges, we present PCCRender, a novel end-to-end framework that achieves high-quality attribute preservation and rendering capabilities while maintaining low bandwidth requirements. To achieve efficient variable rate compression, we introduce Geometry-Invariant Rate Adjustment (GIRA) module that mitigates the influence of point cloud density during rate adjustment. Additionally, to improve decoding speed, we develop Uneven Four-Group Context Model (UFCM), achieving a trade-off between the accuracy and complexity of entropy parameter prediction. Moreover, we design Voxel to Gaussian Primitive Converter (V2G-Converter) which generates Gaussian primitives from decoded point clouds, enabling differentiable rendering. Our unified optimization framework jointly minimizes bit rate while maximizing both 2D rendered image quality and point cloud attribute fidelity. Experimental results demonstrate that PCCRender achieves state-of-the-art performance. Compared to the GPCC v23 (GS Render) method, our framework achieves 10.58% BD-Rate gain in attribute compression and 1.5 dB BD-PSNR increase in 2D rendered visual quality. These compelling results, combined with support for variable rate and free viewpoint rendering, establish PCCRender as a practical solution for the next generation point cloud applications.
Kangli Wang, Ronggang Wang, Ge Li 0002, Junhui Hou, Wei Gao 0003
IEEE Trans. Image Process.5
2026 Spatially-Guided Temporal Aggregation for Robust Event-RGB Optical Flow Estimation
abstract
Current optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel in challenging scenarios. These complementary characteristics underscore the potential of integrating frame and event data for optical flow estimation. However, most cross-modal approaches fail to fully utilize the complementary advantages, relying instead on simply stacking information. This study introduces a novel approach that uses a spatially dense modality to guide the aggregation of the temporally dense event modality, achieving effective cross-modal fusion. Specifically, we propose an event-enhanced frame representation that preserves the rich texture of frames and the basic structure of events. We use the enhanced representation as the guiding modality and employ events to capture temporally dense motion information. The robust motion features derived from the guiding modality direct the aggregation of motion information from events. To further enhance fusion, we propose a transformer-based module that complements sparse event motion features with spatially rich frame information and enhances global information propagation. Additionally, a mix-fusion encoder is designed to extract comprehensive spatiotemporal contextual features from both modalities. Extensive experiments on the MVSEC and DSEC-Flow datasets demonstrate the effectiveness of our framework. Leveraging the complementary strengths of frames and events, our method achieves leading performance on the DSEC-Flow dataset. Compared to the event-only model, frame guidance improves accuracy by 10%. Furthermore, it outperforms the state-of-the-art fusion-based method with a 4% accuracy gain and a 45% reduction in inference time. The code is publicly available athttps://github.com/ZhouQianang/STFlow.
Qianang Zhou, Junhui Hou, Yongjian Deng, Youfu Li 0001, Junlin Xiong
IEEE Trans. Multim.2
2026 SAND: Spatially Adaptive Network Depth for Fast Sampling of Neural Implicit Surfaces
abstract
Implicit neural representations are powerful for geometric modeling, but their practical use is often limited by the high computational cost of network evaluations. We observe that implicit representations require progressively lower accuracy as query points move farther from the target surface, and that even within the same iso-surface, representation difficulty varies spatially with local geometric complexity. However, conventional neural implicit models evaluate all query points with the same network depth and computational cost, ignoring this spatial variation and thereby incurring substantial computational waste. Motivated by this observation, we propose an efficient neural implicit geometry representation framework with spatially adaptive network depth (SAND). SAND leverages a volumetric network-depth map together with a tailed multi-layer perceptron (T-MLP) to model implicit representation. The volumetric depth map records, for each spatial region, the network depth required to achieve sufficient accuracy, while the T-MLP is a modified MLP designed to learn implicit functions such as signed distance functions, where an output branch, referred to as a tail, is attached to each hidden layer. This design allows network evaluation to terminate adaptively without traversing the full network and directs computational resources to geometrically important and complex regions, improving efficiency while preserving high-fidelity representations. Extensive experimental results demonstrate that our approach can significantly improve the inference-time query speed of implicit neural representations.
Chuanxiang Yang, Junhui Hou, Yuan Liu 0025, Guangshun Wei, Taku Komura, Yuanfeng Zhou, Wenping Wang 0001
ACM Trans. Graph.2
2026 DecoRec: Decomposed 3D Scene Reconstruction From Single-View Images via Object-Level Diffusion
abstract
In this paper, we introduce DecoRec, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for single-view scene reconstruction typically rely on object retrieval or the regression of coarse 3D voxels or surfaces, leading to inaccuracies in capturing the appearance and geometry of the input image. The lack of high-quality large-scale scene-level datasets further complicates direct 3D scene generation from single-view images. To achieve high-quality 3D scene generation from a single-view image, DecoRec takes advantage of recent diffusion-based single-view object reconstruction methods to reconstruct individual objects separately. Subsequently, a refinement pipeline is proposed to effectively merge these reconstructed objects, enhancing appearance and geometry through a differentiable rendering technique and diffusion-guided refinement. Our results demonstrate that DecoRec facilitates high-quality single-view scene reconstruction in both geometry and novel synthesis, offering significant benefits for downstream applications like room interior design.
Yuhan Ping, Yuan Liu 0025, Xiaoxiao Long, Peng Wang 0099, Junhui Hou, Jianyi Zheng, Jia Pan 0001, Xin Li 0003, Cheng Lin 0001
IEEE Trans. Vis. Comput. Graph.5
2026 HuGDiffusion: Generalizable Single-Image Human Rendering via 3D Gaussian Diffusion
abstract
We present HuGDiffusion, a generalizable 3D Gaussian splatting (3DGS) learning pipeline to achieve novel view synthesis (NVS) of human characters from single-view input images. Existing approaches typically require monocular videos or calibrated multi-view images as inputs, whose applicability could be weakened in real-world scenarios with arbitrary and/or unknown camera poses. In this paper, we aim to generate the set of 3DGS attributes via a diffusion-based framework conditioned on human priors extracted from a single image. Specifically, we begin with carefully integrated human-centric feature extraction procedures to deduce informative conditioning signals. Based on our empirical observations that jointly learning the whole 3DGS attributes is challenging to optimize, we design a multi-stage generation strategy to obtain different types of 3DGS attributes. To facilitate the training process, we investigate constructing proxy ground-truth 3D Gaussian attributes as high-quality attribute-level supervision signals. Through extensive experiments, our HuGDiffusion shows significant performance improvements over the state-of-the-art methods.
Yingzhi Tang, Qijian Zhang, Junhui Hou
IEEE Trans. Vis. Comput. Graph.3
2026 SuperCarver: Texture-Consistent 3D Geometry Super-Resolution for High-Fidelity Surface Detail Generation
abstract
Conventional production workflow of high-precision mesh assets necessitates a cumbersome and laborious process of manual sculpting by specialized 3D artists/modelers. The recent years have witnessed remarkable advances in AI-empowered 3D content creation for generating plausible structures and intricate appearances from images or text prompts. However, synthesizing realistic surface details still poses great challenges, and enhancing the geometry fidelity of existing lower-quality 3D meshes (instead of image/text-to-3D generation) remains an open problem. In this paper, we introduce SuperCarver, a 3D geometry super-resolution pipeline for supplementing texture-consistent surface details onto a given coarse mesh. We start by rendering the original textured mesh into the image domain from multiple viewpoints. To achieve detail boosting, we construct a deterministic prior-guided normal diffusion model, which is fine-tuned on a carefully curated dataset of paired detail-lacking and detail-rich normal map renderings. To update mesh surfaces from potentially imperfect normal map predictions, we design a noise-resistant inverse rendering scheme through deformable distance field. Experiments demonstrate that our SuperCarver is capable of generating realistic and expressive surface details depicted by the actual texture appearance, making it a powerful tool to both upgrade historical low-quality 3D assets and reduce the workload of sculpting high-poly meshes.
Qijian Zhang, Xiaozheng Jian, Wenping Wang 0001, Junhui Hou
IEEE Trans. Vis. Comput. Graph.5
2026 GSwap: Realistic Head Swapping With Dynamic Neural Gaussian Field
abstract
We present GSwap, a novel consistent and realistic video head-swapping system empowered by dynamic neural Gaussian portrait priors, which significantly advances the state of the art in face and head replacement. Unlike previous methods that rely primarily on 2D generative models or 3D Morphable Face Models (3DMM), our approach overcomes their inherent limitations, including poor 3D consistency, unnatural facial expressions, and restricted synthesis quality. Moreover, existing techniques struggle with full head-swapping tasks due to insufficient holistic head modeling and ineffective background blending, often resulting in visible artifacts and misalignments. To address these challenges, GSwap introduces an intrinsic 3D Gaussian feature field embedded within a full-body SMPL-X surface, effectively elevating 2D portrait videos into a dynamic neural Gaussian field. This innovation ensures high-fidelity, 3D-consistent portrait rendering while preserving natural head-torso relationships and seamless motion dynamics. To facilitate training, we adapt a pretrained 2D portrait generative model to the source head domain using only a few reference images, enabling efficient domain adaptation. Furthermore, we propose a neural re-rendering strategy that harmoniously integrates the synthesized foreground with the original background, eliminating blending artifacts and enhancing realism. Extensive experiments demonstrate that GSwap surpasses existing methods in multiple aspects, including visual quality, temporal coherence, identity preservation, and 3D consistency.
Xuan Gao 0003, Dongyu Liu, Junhui Hou, Juyong Zhang
IEEE Trans. Vis. Comput. Graph.4
2025 A Lightweight UDF Learning Framework for 3D Reconstruction Based on Local Shape Functions
abstract
Unsigned distance fields (UDFs) provide a versatile framework for representing a diverse array of 3D shapes, encompassing both watertight and non-watertight geometries. Traditional UDF learning methods typically require extensive training on large 3D shape datasets, which is costly and necessitates re-training for new datasets. This paper presents a novel neural framework, LoSF-UDF, for reconstructing surfaces from 3D point clouds by leveraging local shape functions to learn UDFs. We observe that 3D shapes manifest simple patterns in localized regions, prompting us to develop a training dataset of point cloud patches characterized by mathematical functions that represent a continuum from smooth surfaces to sharp edges and corners. Our approach learns features within a specific radius around each query point and utilizes an attention mechanism to focus on the crucial features for UDF estimation. Despite being highly lightweight, with only 653 KB of trainable parameters and a modest-sized training dataset with 0.5 GB storage, our method enables efficient and robust surface reconstruction from point clouds without requiring for shape-specific training. Furthermore, our method exhibits enhanced resilience to noise and outliers in point clouds compared to existing methods. We conduct comprehensive experiments and comparisons across various datasets, including synthetic and real-scanned point clouds, to validate our method’s efficacy. Notably, our lightweight framework offers rapid and reliable initialization for other unsupervised iterative approaches, improving both the efficiency and accuracy of their reconstructions. Our project and code are available at https://jbhu67.github.io/LoSF-UDF.github.io/.
Jiangbei Hu, Yanggeng Li, Fei Hou 0001, Junhui Hou, Zhebin Zhang, Shengfa Wang, Na Lei, Ying He 0001
CVPR4
2025 Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score Distillation
abstract
We present Acc3D to tackle the challenge of accelerating the diffusion process to generate 3D models from single images. To derive high-quality reconstructions through few-step inferences, we emphasize the critical issue of regularizing the learning of score function in states of random noise. To this end, we propose edge consistency, i.e., consistent predictions across the high signal-to-noise ratio region, to enhance a pre-trained diffusion model, enabling a distillation-based refinement of the endpoint score function. Building on those distilled diffusion models, we propose an adversarial augmentation strategy to further enrich the generation detail and boost overall generation quality. The two modules complement each other, mutually reinforcing to elevate generative performance. Extensive experiments demonstrate that our Acc3D not only achieves over a 20× increase in computational efficiency but also yields notable quality improvements, compared to the state-of-the-arts.
Kendong Liu, Hui Liu 0032, Junhui Hou
CVPR4
2025 RigGS: Rigging of 3D Gaussians for Modeling Articulated Objects in Videos
abstract
This paper considers the problem of modeling articulated objects captured in 2D videos to enable novel view synthesis, while also being easily editable, drivable, and reposable. To tackle this challenging problem, we propose RigGS, a new paradigm that leverages 3D Gaussian representation and skeleton-based motion representation to model dynamic objects without utilizing additional template priors. Specifically, we first propose skeleton-aware node-controlled deformation, which deforms a canonical 3D Gaussian representation over time to initialize the modeling process, producing candidate skeleton nodes that are further simplified into a sparse 3D skeleton according to their motion and semantic information. Subsequently, based on the resulting skeleton, we design learnable skin deformations and pose-dependent detailed deformations, thereby easily deforming the 3D Gaussian representation to generate new actions and render further high-quality images from novel views. Extensive experiments demonstrate that our method can generate realistic new actions easily for objects and achieve high-quality rendering.
Yuxin Yao 0001, Junhui Hou
CVPR3
2025 VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object Tracking
Zekun Qian, Rui-Ze Han, Junhui Hou, Linqi Song, Wei Feng 0005
ICCV3
2025 COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue Fusion
Zekun Qian, Rui-Ze Han, Junhui Hou, Wei Feng 0005
ICCV4
2025 Neural Compression for 3D Geometry Sets
Junhui Hou, Weiyao Lin, Wenping Wang 0001
ICCV2
2025 Towards Calibrated Deep Clustering Network
abstract
Deep clustering has exhibited remarkable performance; however, the over confidence problem, i.e., the estimated confidence for a sample belonging to a particular cluster greatly exceeds its actual prediction accuracy, has been over looked in prior research. To tackle this critical issue, we pioneer the development of a calibrated deep clustering framework. Specifically, we propose a novel dual head (calibration head and clustering head) deep clustering model that can effectively calibrate the estimated confidence and the actual accuracy. The calibration head adjusts the overconfident predictions of the clustering head, generating prediction confidence that matches the model learning status. Then, the clustering head dynamically selects reliable high-confidence samples estimated by the calibration head for pseudo-label self-training. Additionally, we introduce an effective network initialization strategy that enhances both training speed and network robustness. The effectiveness of the proposed calibration approach and initialization strategy are both endorsed with solid theoretical guarantees. Extensive experiments demonstrate the proposed calibrated deep clustering model not only surpasses the state-of-the-art deep clustering methods by 5× on average in terms of expected calibration error, but also significantly outperforms them in terms of clustering accuracy. The code is available at https://github.com/ChengJianH/CDC.
Yuheng Jia, Jianhong Cheng, Hui Liu 0032, Junhui Hou
ICLR4
2025 MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors
abstract
In this paper, we propose MoDGS, a new pipeline to render novel-view images in dynamic scenes using only casually captured monocular videos. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid movement of input cameras to construct multiview consistency but fail to reconstruct dynamic scenes on casually captured input videos whose cameras are static or move slowly. To address this challenging task, MoDGS adopts recent single-view depth estimation methods to guide the learning of the dynamic scene. Then, a novel 3D-aware initialization method is proposed to learn a reasonable deformation field and a new robust depth loss is proposed to guide the learning of dynamic scene geometry. Comprehensive experiments demonstrate that MoDGS is able to render high-quality novel view images of dynamic scenes from just a casually captured monocular video, which outperforms baseline methods by a significant margin. Project page: https://MoDGS.github.io
Qingming Liu, Yuan Liu 0025, Jiepeng Wang 0001, Xianqiang Lyu, Peng Wang 0099, Wenping Wang 0001, Junhui Hou
ICLR7
2025 ParaSolver: A Hierarchical Parallel Integral Solver for Diffusion Models
abstract
This paper explores the challenge of accelerating the sequential inference process of Diffusion Probabilistic Models (DPMs). We tackle this critical issue from a dynamic systems perspective, in which the inherent sequential nature is transformed into a parallel sampling process. Specifically, we propose a unified framework that generalizes the sequential sampling process of DPMs as solving a system of banded nonlinear equations. Under this generic framework, we reveal that the Jacobian of the banded nonlinear equations system possesses a unit-diagonal structure, enabling further approximation for acceleration. Moreover, we theoretically propose an effective initialization approach for parallel sampling methods. Finally, we construct \textit{ParaSolver}, a hierarchical parallel sampling technique that enhances sampling speed without compromising quality. Extensive experiments show that ParaSolver achieves up to \textbf{12.1× speedup} in terms of wall-clock time. The source code is publicly available at https://github.com/Jianrong-Lu/ParaSolver.git.
Jianrong Lu, Junhui Hou
ICLR3
2025 Shape as Line Segments: Accurate and Flexible Implicit Surface Representation
abstract
Distance field-based implicit representations like signed/unsigned distance fields have recently gained prominence in geometry modeling and analysis. However, these distance fields are reliant on the closest distance of points to the surface, introducing inaccuracies when interpolating along cube edges during surface extraction. Additionally, their gradients are ill-defined at certain locations, causing distortions in the extracted surfaces. To address this limitation, we propose Shape as Line Segments (SALS), an accurate and efficient implicit geometry representation based on attributed line segments, which can handle arbitrary structures. Unlike previous approaches, SALS leverages a differentiable Line Segment Field to implicitly capture the spatial relationship between line segments and the surface. Each line segment is associated with two key attributes, intersection flag and ratio, from which we propose edge-based dual contouring to extract a surface. We further implement SALS with a neural network, producing a new neural implicit presentation. Additionally, based on SALS, we design a novel learning-based pipeline for reconstructing surfaces from 3D point clouds. We conduct extensive experiments, showcasing the significant advantages of our methods over state-of-the-art methods. The source code is available at https://github.com/rsy6318/SALS.
Junhui Hou
ICLR2
2025 NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer
abstract
By harnessing the potent generative capabilities of pre-trained large video diffusion models, we propose a new novel view synthesis paradigm that operates without the need for training. The proposed method adaptively modulates the diffusion sampling process with the given views to enable the creation of visually pleasing results from single or multiple views of static scenes or monocular videos of dynamic scenes. Specifically, built upon our theoretical modeling, we iteratively modulate the score function with the given scene priors represented with warped input views to control the video diffusion process. Moreover, by theoretically exploring the boundary of the estimation error, we achieve the modulation in an adaptive fashion according to the view pose and the number of diffusion steps. Extensive evaluations on both static and dynamic scenes substantiate the significant superiority of our method over state-of-the-art methods both quantitatively and qualitatively. The source code can be found on https://github.com/ZHU-Zhiyu/NVS_Solver.
Meng You, Hui Liu 0032, Junhui Hou
ICLR4
2025 Learning from Sample Stability for Deep Clustering
abstract
Deep clustering, an unsupervised technique independent of labels, necessitates tailored supervision for model training. Prior methods explore supervision like similarity and pseudo labels, yet overlook individual sample training analysis. Our study correlates sample stability during unsupervised training with clustering accuracy and network memorization on a per-sample basis. Unstable representations across epochs often lead to mispredictions, indicating difficulty in memorization and atypicality. Leveraging these findings, we introduce supervision signals for the first time based on sample stability at the representation level. Our proposed strategy serves as a versatile tool to enhance various deep clustering techniques. Experiments across benchmark datasets showcase that incorporating sample stability into training can improve the performance of deep clustering. The code is available at https://github.com/LZX-001/LFSS.
Yuheng Jia, Hui Liu 0032, Junhui Hou
ICML4
2025 Generalization Performance of Ensemble Clustering: From Theory to Algorithm
abstract
Ensemble clustering has demonstrated great success in practice; however, its theoretical foundations remain underexplored. This paper examines the generalization performance of ensemble clustering, focusing on generalization error, excess risk and consistency. We derive a convergence rate of generalization error bound and excess risk bound both of $\mathcal{O}(\sqrt{\frac{\log n}{m}}+\frac{1}{\sqrt{n}})$, with $n$ and $m$ being the numbers of samples and base clusterings. Based on this, we prove that when $m$ and $n$ approach infinity and $m$ is significantly larger than log $n$, i.e., $m,n\to \infty, m\gg \log n$, ensemble clustering is consistent. Furthermore, recognizing that $n$ and $m$ are finite in practice, the generalization error cannot be reduced to zero. Thus, by assigning varying weights to finite clusterings, we minimize the error between the empirical average clusterings and their expectation. From this, we theoretically demonstrate that to achieve better clustering performance, we should minimize the deviation (bias) of base clustering from its expectation and maximize the differences (diversity) among various base clusterings. Additionally, we derive that maximizing diversity is nearly equivalent to a robust (min-max) optimization model. Finally, we instantiate our theory to develop a new ensemble clustering algorithm. Compared with SOTA methods, our approach achieves average improvements of 6.0%, 7.3%, and 6.0% on 10 datasets w.r.t. NMI, ARI, and Purity. The code is available at https://github.com/xuz2019/GPEC.
Haoye Qiu, Weixuan Liang, Hui Liu 0032, Junhui Hou, Yuheng Jia
ICML5
2025 Mixed Blessing: Class-Wise Embedding guided Instance-Dependent Partial Label Learning
abstract
In partial label learning (PLL), every sample is associated with a candidate label set comprising the ground-truth label and several noisy labels. The conventional PLL assumes the noisy labels are randomly generated (instance-independent), while in practical scenarios, the noisy labels are always instance-dependent and are highly related to the sample features, leading to the instance-dependent partial label learning (IDPLL) problem. Instance-dependent noisy label is a double-edged sword. On one side, it may promote model training as the noisy labels can depict the sample to some extent. On the other side, it brings high label ambiguity as the noisy labels are quite undistinguishable from the ground-truth label. To leverage the nuances of IDPLL effectively, for the first time we create class-wise embeddings for each sample, which allow us to explore the relationship of instance-dependent noisy labels, i.e., the class-wise embeddings in the candidate label set should have high similarity, while the class-wise embeddings between the candidate label set and the non-candidate label set should have high dissimilarity. Moreover, to reduce the high label ambiguity, we introduce the concept of class prototypes containing global feature information to disambiguate the candidate label set. Extensive experimental comparisons with twelve methods on six benchmark data sets, including four fine-grained data sets, demonstrate the effectiveness of the proposed method. The code implementation is publicly available at https://github.com/Yangfc-ML/CEL.
Fuchao Yang, Jianhong Cheng, Hui Liu 0032, Yongqiang Dong, Yuheng Jia, Junhui Hou
KDD (1)6
2025 Keep It on a Leash: Controllable Pseudo-label Generation Towards Realistic Long-Tailed Semi-Supervised Learning
abstract
Current long-tailed semi-supervised learning methods assume that labeled data exhibit a long-tailed distribution, and unlabeled data adhere to a typical predefined distribution (i.e., long-tailed, uniform, or inverse long-tailed). However, the distribution of the unlabeled data is generally unknown and may follow an arbitrary distribution. To tackle this challenge, we propose a Controllable Pseudo-label Generation (CPG) framework, expanding the labeled dataset with the progressively identified reliable pseudo-labels from the unlabeled dataset and training the model on the updated labeled dataset with a known distribution, making it unaffected by the unlabeled data distribution. Specifically, CPG operates through a controllable self-reinforcing optimization cycle: (i) at each training step, our dynamic controllable filtering mechanism selectively incorporates reliable pseudo-labels from the unlabeled dataset into the labeled dataset, ensuring that the updated labeled dataset follows a known distribution; (ii) we then construct a Bayes-optimal classifier using logit adjustment based on the updated labeled data distribution; (iii) this improved classifier subsequently helps identify more reliable pseudo-labels in the next training step. We further theoretically prove that this optimization cycle can significantly reduce the generalization error under some conditions. Additionally, we propose a class-aware adaptive augmentation module to further improve the representation of minority classes, and an auxiliary branch to maximize data utilization by leveraging all labeled and unlabeled samples. Comprehensive evaluations on various commonly used benchmark datasets show that CPG achieves consistent improvements, surpassing state-of-the-art methods by up to **15.97\%** in accuracy. The code is available at https://github.com/yaxinhou/CPG.
Yaxin Hou, Bo Han 0017, Yuheng Jia, Hui Liu 0032, Junhui Hou
NeurIPS5
2025 You Can Trust Your Clustering Model: A Parameter-free Self-Boosting Plug-in for Deep Clustering
abstract
Recent deep clustering models have produced impressive clustering performance. However, a common issue with existing methods is the disparity between global and local feature structures. While local structures typically show strong consistency and compactness within class samples, global features often present intertwined boundaries and poorly separated clusters. Motivated by this observation, we propose **DCBoost, a parameter-free plug-in** designed to enhance the global feature structures of current deep clustering models. By harnessing reliable local structural cues, our method aims to elevate clustering performance effectively. Specifically, we first identify high-confidence samples through adaptive $k$-nearest neighbors-based consistency filtering, aiming to select a sufficient number of samples with high label reliability to serve as trustworthy anchors for self-supervision. Subsequently, these samples are utilized to compute a discriminative loss, which promotes both intra-class compactness and inter-class separability, to guide network optimization. Extensive experiments across various benchmark datasets showcase that our DCBoost significantly improves the clustering performance of diverse existing deep clustering models. Notably, our method improves the performance of current state-of-the-art baselines (e.g., ProPos) by more than 3\% and amplifies the silhouette coefficient by over $7\times$. **Code is available at [https://github.com/l-h-y168/DCBoost](https://github.com/l-h-y168/DCBoost).**
Yuheng Jia, Hui Liu 0032, Junhui Hou
NeurIPS4
2025 DOVTrack: Data-Efficient Open-Vocabulary Tracking
abstract
Open-Vocabulary Multi-Object Tracking (OVMOT) aims to detect and track multi-category objects including both seen and unseen categories during training. Currently, a significant challenge in this domain is the lack of large-scale annotated video data for training. To address this challenge, this work aims to effectively train the OV tracker using only the existing limited and sparsely annotated video data. We propose a comprehensive training sample space expansion strategy that addresses the fundamental limitation of sparse annotations in OVMOT training. Specifically, for the association task, we develop a diffusion-based feature generation framework that synthesizes intermediate object features between sparsely annotated frames, effectively expanding the training sample space by approximately 3× and enabling robust association learning from temporally continuous features. For the detection task, we introduce a dynamic group contrastive learning approach that generates diverse sample groups through affinity, dispersion, and adversarial grouping strategies, tripling the effective training samples for classification while maintaining sample quality. Additionally, we propose an adaptive localization loss that expands positive sample coverage by lowering IoU thresholds while mitigating noise through confidence-based weighting. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the OVMOT benchmark, surpassing existing methods by 3.8\% in TETA metric, without requiring additional data or annotations. The code will be available at https://github.com/zekunqian/DOVTrack.
Zekun Qian, Rui-Ze Han, Junhui Hou, Wei Feng 0005
NeurIPS4
2025 Message from Guest Editors of the CVM 2025 Special Issue
abstract
The Computational Visual Media (CVM) conference series provides a leading international forum for the exchange of innovative research ideas and significant computational methodologies that both underpin and advance visual media. Its primary mission is to foster cross-disciplinary research that integrates computer graphics, computer vision, machine learning, image and video processing, visualization, and geometric computing. Topics of particular interest include classification, composition, retrieval, synthesis, cognition, and understanding of visual media, encompassing images, video, and 3D geometry.
Piotr Didyk, Junhui Hou
Comput. Vis. Media2
2025 Preface
Shi-Min Hu 0001, Piotr Didyk, Junhui Hou
J. Comput. Sci. Technol.3
2025 Unveiling the Power of Self-Supervision for Multi-View Multi-Human Association and Tracking
abstract
Multi-view multi-human association and tracking (MvMHAT), is an emerging yet important problem for multi-person scene video surveillance, aiming to track a group of people over time in each view, as well as to identify the same person across different views at the same time, which is different from previous MOT and multi-camera MOT tasks only considering the over-time human tracking. This way, the videos for MvMHAT require more complex annotations while containing more information for self-learning. In this work, we tackle this problem with an end-to-end neural network in a self-supervised learning manner. Specifically, we propose to take advantage of the spatial-temporal self-consistency rationale by considering three properties of reflexivity, symmetry, and transitivity. Besides the reflexivity property that naturally holds, we design the self-supervised learning losses based on the properties of symmetry and transitivity, for both appearance feature learning and assignment matrix optimization, to associate multiple humans over time and across views. Furthermore, to promote the research on MvMHAT, we build two new large-scale benchmarks for the network training and testing of different algorithms. Extensive experiments on the proposed benchmarks verify the effectiveness of our method. We have released the benchmark and code to the public.
Wei Feng 0005, Fei Wang 0032, Rui-Ze Han, Yiyang Gan, Zekun Qian, Junhui Hou, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 DDM: A Metric for Comparing 3D Shapes Using Directional Distance Fields
abstract
Qualifying the discrepancy between 3D geometric models, which could be represented with either point clouds or triangle meshes, is a pivotal issue with board applications. Existing methods mainly focus on directly establishing the correspondence between two models and then aggregating point-wise distance between corresponding points, resulting in them being either inefficient or ineffective. In this paper, we propose DDM, an efficient, effective, robust, and differentiable distance metric for 3D geometry data. Specifically, we construct DDM based on the proposed implicit representation of 3D models, namely directional distance field (DDF), which defines the directional distances of 3D points to a model to capture its local surface geometry. We then transfer the discrepancy between two 3D geometric models as the discrepancy between their DDFs defined on an identical domain, naturally establishing model correspondence. To demonstrate the advantage of our DDM, we explore various distance metric-driven 3D geometric modeling tasks, including template surface fitting, rigid registration, non-rigid registration, scene flow estimation and human pose optimization. Extensive experiments show that our DDM achieves significantly higher accuracy under all tasks. As a generic distance metric, DDM has the potential to advance the field of 3D geometric modeling.
Junhui Hou, Xiaodong Chen 0009, Hongkai Xiong, Wenping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Human as Points: Explicit Point-Based 3D Human Reconstruction From Single-View RGB Images
abstract
The latest trends in the research field of single-view human reconstruction are devoted to learning deep implicit functions constrained by explicit body shape priors. Despite the remarkable performance improvements compared with traditional processing pipelines, existing learning approaches still exhibit limitations in terms of flexibility, generalizability, robustness, and/or representation capability. To comprehensively address the above issues, in this paper, we investigate an explicit point-based human reconstruction framework named HaP, which utilizes point clouds as the intermediate representation of the target geometric structure. Technically, our approach features fully explicit point cloud estimation (exploiting depth and SMPL), manipulation (SMPL rectification), generation (built upon diffusion), and refinement (displacement learning and depth replacement) in the 3D geometric space, instead of an implicit learning process that can be ambiguous and less controllable. Extensive experiments demonstrate that our framework achieves quantitative performance improvements of 20$\%$% to 40$\%$% over current state-of-the-art methods, and better qualitative results. Our promising results may indicate a paradigm rollback to the fully-explicit and geometry-centric algorithm design. In addition, we newly contribute a real-scanned 3D human dataset featuring more intricate geometric details.
Yingzhi Tang, Qijian Zhang, Yebin Liu, Junhui Hou
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Monge-Ampere Regularization for Learning Arbitrary Shapes From Point Clouds
abstract
As commonly used implicit geometry representations, the signed distance function (SDF) is limited to modeling watertight shapes, while the unsigned distance function (UDF) is capable of representing various surfaces. However, its inherent theoretical shortcoming, i.e., the non-differentiability at the zero-level set, would result in sub-optimal reconstruction quality. In this paper, we propose the scaled-squared distance function (S2DF), a novel implicit surface representation for modeling arbitrary surface types. S2DF does not distinguish between inside and outside regions while effectively addressing the non-differentiability issue of UDF at the zero-level set. We demonstrate that S2DF satisfies a second-order partial differential equation of Monge-Ampere-type, allowing us to develop a learning pipeline that leverages a novel MongeAmpere regularization to directly learn S2DF from raw unoriented point clouds without supervision from ground-truth S2DF values. Extensive experiments across multiple datasets show that our method significantly outperforms state-of-the-art supervised approaches that require ground-truth surface information as supervision for training. The code will be publicly available at https://github.com/chuanxiang-yang/S2DF.
Chuanxiang Yang, Yuanfeng Zhou, Guangshun Wei, Long Ma 0009, Junhui Hou, Yuan Liu 0025, Wenping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 SPARE: Symmetrized Point-to-Plane Distance for Robust Non-Rigid 3D Registration
abstract
Existing optimization-based methods for non-rigid registration typically minimize an alignment error metric based on the point-to-point or point-to-plane distance between corresponding point pairs on the source surface and target surface. However, these metrics can result in slow convergence or a loss of detail. In this paper, we propose SPARE, a novel formulation that utilizes a symmetrized point-to-plane distance for robust non-rigid registration. The symmetrized point-to-plane distance relies on both the positions and normals of the corresponding points, resulting in a more accurate approximation of the underlying geometry and can achieve higher accuracy than existing methods. To solve this optimization problem efficiently, we introduce an as-rigid-as-possible regulation term to estimate the deformed normals and propose an alternating minimization solver using a majorization-minimization strategy. Moreover, for effective initialization of the solver, we incorporate a deformation graph-based coarse alignment that improves registration quality and efficiency. Extensive experiments show that the proposed method greatly improves the accuracy of non-rigid registration problems and maintains relatively high solution efficiency.
Yuxin Yao 0001, Bailin Deng, Junhui Hou, Juyong Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Self-Supervised Learning of LiDAR 3D Point Clouds via 2D-3D Neural Calibration
abstract
This paper introduces a novel self-supervised learning framework for enhancing 3D perception in autonomous driving scenes. Specifically, our approach, namely NCLR, focuses on 2D-3D neural calibration, a novel pretext task that estimates the rigid pose aligning camera and LiDAR coordinate systems. First, we propose the learnable transformation alignment to bridge the domain gap between image and point cloud data, converting features into a unified representation space for effective comparison and matching. Second, we identify the overlapping area between the image and point cloud with the fused features. Third, we establish dense 2D-3D correspondences to estimate the rigid pose. The framework not only learns fine-grained matching from points to pixels but also achieves alignment of the image and point cloud at a holistic level, understanding the LiDAR-to-camera extrinsic parameters. We demonstrate the efficacy of NCLR by applying the pre-trained backbone to downstream tasks, such as LiDAR-based 3D semantic segmentation, object detection, and panoptic segmentation. Comprehensive experiments on various datasets illustrate the superiority of NCLR over existing self-supervised methods. The results confirm that joint learning from different modalities significantly enhances the network's understanding abilities and effectiveness of learned representation.
Yifan Zhang 0036, Junhui Hou, Jinjian Wu, Yixuan Yuan, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Learning Efficient and Effective Trajectories for Differential Equation-Based Image Restoration
abstract
The differential equation-based image restoration approach aims to establish learnable trajectories connecting high-quality images to a tractable distribution, e.g., low-quality images or a Gaussian distribution. In this paper, we reformulate the trajectory optimization of this kind of method, focusing on enhancing both reconstruction quality and efficiency. Initially, we navigate effective restoration paths through a reinforcement learning process, gradually steering potential trajectories toward the most precise options. Additionally, to mitigate the considerable computational burden associated with iterative sampling, we propose cost-aware trajectory distillation to streamline complex paths into several manageable steps with adaptable sizes. Moreover, we fine-tune a foundational diffusion model (FLUX) with 12B parameters by using our algorithms, producing a unified framework for handling 7 kinds of image restoration tasks. Extensive experiments showcase the significant superiority of the proposed method, achieving a maximum PSNR improvement of 2.1 dB over state-of-the-art methods, while also greatly enhancing visual perceptual quality.
Jinhui Hou, Hui Liu 0032, Huanqiang Zeng, Junhui Hou
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
abstract
Continual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced sequentially. For video data, the task becomes more complex than image data because it requires learning and preserving both spatial appearance and temporal action involvement. To address this challenge, we propose a novel exemplar-free framework that equips separate spatiotemporal adapters to learn new class patterns, accommodating the incremental information representation requirements unique to each class. While separate adapters are proven to mitigate forgetting and fit unique requirements, naively applying them hinders the intrinsic connection between spatial and temporal information increments, affecting the efficiency of representing newly learned class information. Motivated by this, we introduce two key innovations from a causal perspective. First, a causal distillation module is devised to maintain the relation between spatial-temporal knowledge for a more efficient representation. Second, a causal compensation mechanism is proposed to reduce the conflicts during increment and memorization between different types of information. Extensive experiments conducted on benchmark datasets demonstrate that our framework can achieve new state-of-the-art results, surpassing current example-based methods by 4.2% in accuracy on average. The codes are accessible in https://github.com/tychen-SJTU/CSTA.
Tieyuan Chen, Huabin Liu 0001, Chern Hong Lim, John See, Xing Gao 0005, Junhui Hou, Weiyao Lin
IEEE Trans. Circuits Syst. Video Technol.6
2025 Rendering-Oriented 3D Point Cloud Attribute Compression Using Sparse Tensor-Based Transformer
abstract
The evolution of 3D visualization techniques has fundamentally transformed how we interact with digital content. At the forefront of this change is point cloud technology, offering an immersive experience that surpasses traditional 2D representations. However, the massive data size of point clouds presents significant challenges in data compression. Current methods for lossy point cloud attribute compression (PCAC) generally focus on reconstructing the original point clouds with minimal error. However, for point cloud visualization scenarios, the reconstructed point clouds with distortion still need to undergo a complex rendering process, which affects the final user-perceived quality. In this paper, we propose an end-to-end deep learning framework that seamlessly integrates PCAC with differentiable rendering, denoted as rendering-oriented PCAC (RO-PCAC), directly targeting the quality of rendered multiview images for viewing. In a differentiable manner, the impact of the rendering process on the reconstructed point clouds is taken into account. Moreover, we characterize point clouds as sparse tensors and propose a sparse tensor-based transformer, called SP-Trans. By aligning with the local density of the point cloud and utilizing an enhanced local attention mechanism, SP-Trans captures the intricate relationships within the point cloud, further improving feature analysis and synthesis within the framework. Extensive experiments demonstrate that the proposed RO-PCAC achieves state-of-the-art compression performance, compared to existing reconstruction-oriented methods, including traditional, learning-based, and hybrid methods. The code will be released athttps://github.com/net-F/RO-PCAC.git.
Xiao Huo, Junhui Hou, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Boosting 3D Object Detection With Semantic-Aware Multi-Branch Framework
abstract
In autonomous driving, LiDAR sensors are vital for acquiring 3D point clouds, providing reliable geometric information. However, traditional sampling methods of preprocessing often ignore semantic features, leading to detail loss and ground point interference in 3D object detection. To address this, we propose a multi-branch two-stage 3D object detection framework using a Semantic-aware Multi-branch Sampling (SMS) module and multi-view consistency constraints. The SMS module includes random sampling, Density Equalization Sampling (DES) for enhancing distant objects, and Ground Abandonment Sampling (GAS) to focus on non-ground points. The sampled multi-view points are processed through a Consistent KeyPoint Selection (CKPS) module to generate consistent keypoint masks for efficient proposal sampling. The first-stage detector uses multi-branch parallel learning with multi-view consistency loss for feature aggregation, while the second-stage detector fuses multi-view data through a Multi-View Fusion Pooling (MVFP) module to precisely predict 3D objects. The experimental results on the KITTI dataset and Waymo Open Dataset show that our method achieves excellent detection performance improvement for a variety of backbones, especially for low-performance backbones with simple network structures. The code will be publicly available at https://github.com/HaoJing-SX/SMS.
Anhong Wang, Lijun Zhao 0002, Yakun Yang, Donghan Bu, Yifan Zhang 0036, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.8
2025 A No-Reference Quality Assessment Model for Screen Content Videos via Hierarchical Spatiotemporal Perception
abstract
In this paper, a novel deep learning-based no-reference video quality assessment (NR-VQA) model for screen content videos (SCVs) is proposed, called the hierarchical spatiotemporal perceptual quality model (HSPQ). Firstly, the human visual system (HVS) perceives SCVs hierarchically, with varying sensitivity and attention to diverse attribute regions. Secondly, the visual redundancies are copious in the spatiotemporal domain of SCVs, degrading video quality to some extent. Based on these characteristics, the SCVs are decomposed into three hierarchical levels (i.e., patch level, frame level, and video level), which contain quality-related spatiotemporal information. Specifically, the visual saliency is first utilized for more salient textual and pictorial patches selection, and then, a dual-channel convolutional neural network integrating spatial-gate feature enhancement module (SGFEM) is designed to evaluate the quality of patches based on their attributes at the patch level separately. With spatial correlation, an adaptive blur-focused visual mechanism-based weighting strategy (BFWS) is proposed for converting quality scores from patch level to frame level. Finally, the video-level quality score, which reflects the temporal perceptual quality degradation, is combined to provide a comprehensive evaluation of distorted SCV quality. Experiments conducted on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases demonstrate that our proposed HSPQ model aligns better with the visual perception of SCVs by the HVS. Moreover, it exhibits strong robustness compared to multiple classic and state-of-the-art image/video quality assessment models.
Huanqiang Zeng, Jing Chen 0001, Yifan Shi 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.6
2025 Synthetic-to-Real Video Person Re-ID
abstract
Person re-identification (Re-ID) is an important task and has significant applications for public security and information forensics, which has progressed rapidly with the development of deep learning. In this work, we investigate a novel and challenging setting of Re-ID, i.e., cross-domain video-based person Re-ID. Specifically, we utilize synthetic video datasets as the source domain for training and real-world videos for testing, notably reducing the reliance on expensive real data acquisition and annotation. To harness the potential of synthetic data, we first propose a self-supervised domain-invariant feature learning strategy for both static and dynamic (temporal) features. Additionally, to enhance person identification accuracy in the target domain, we propose a mean-teacher scheme incorporating a self-supervised ID consistency loss. Experimental results across five real datasets validate the rationale behind cross-synthetic-real domain adaptation and demonstrate the efficacy of our method. Notably, the discovery that synthetic data outperforms real data in the cross-domain scenario is a surprising outcome. The code and data are publicly available at https://github.com/XiangqunZhang/UDA_Video_ReID.
Xiangqun Zhang 0003, Rui-Ze Han, Likai Wang 0002, Linqi Song, Junhui Hou, Wei Feng 0005
IEEE Trans. Inf. Forensics Secur.5
2025 HOPE: Enhanced Position Image Priors via High-Order Implicit Representations
abstract
Deep Image Prior (DIP) has shown that networks with stochastic initialization and custom architectures can effectively address inverse imaging challenges. Despite its potential, DIP requires significant computational resources, whereas the lighter Implicit Neural Positional Image Prior (PIP) often yields overly smooth solutions due to exacerbated spectral bias. Research on lightweight, high-performance solutions for inverse imaging remains limited. This paper proposes a novel framework, Enhanced Positional Image Priors through High-Order Implicit Representations (HOPE), incorporating high-order interactions between layers within a conventional cascade structure. This approach reduces the spectral bias commonly seen in PIP, enhancing the model's ability to capture both low- and high-frequency components for optimal inverse problem performance. We theoretically demonstrate that HOPE's expanded representational space, narrower convergence range, and improved Neural Tangent Kernel (NTK) diagonal properties enable more precise frequency representations than PIP. Comprehensive experiments across tasks such as signal representation (audio, image, volume) and inverse image processing (denoising, super-resolution, CT reconstruction, inpainting) confirm that HOPE establishes new benchmarks for recovery quality and training efficiency.
Ruituo Wu, Junhui Hou, Ce Zhu, Yipeng Liu 0001
IEEE Trans. Image Process.3
2025 PVNet: Point-Voxel Interaction LiDAR Scene Upsampling via Diffusion Models
abstract
Accurate 3D scene understanding in outdoor environments heavily relies on high-quality point clouds. However, LiDAR-scanned data often suffer from extreme sparsity, severely hindering downstream 3D perception tasks. Existing point cloud upsampling methods primarily focus on individual objects, thus demonstrating limited generalization capability for complex outdoor scenes. To address this issue, we propose PVNet, a diffusion model-based point-voxel interaction framework to perform LiDAR point cloud upsampling without dense supervision. Specifically, we adopt the classifier-free guidance-based DDPMs to guide the generation, in which we employ a sparse point cloud as the guiding condition and the synthesized point clouds derived from its nearby frames as the input. Moreover, we design a voxel completion module to refine and complete the coarse voxel features for enriching the feature representation. In addition, we propose a point-voxel interaction module to integrate features from both points and voxels, which efficiently improves the environmental perception capability of each upsampled point. To the best of our knowledge, our approach is the first scene-level point cloud upsampling method supporting arbitrary upsampling rates. Extensive experiments on various benchmarks demonstrate that our method achieves state-of-the-art performance. The source code will be available at https://github.com/chengxianjing/PVNet.
Xianjing Cheng, Lintai Wu, Zuowen Wang, Junhui Hou, Jie Wen 0001, Yong Xu 0001
IEEE Trans. Image Process.4
2025 Irregular Tensor Low-Rank Representation for Hyperspectral Image Representation
abstract
Spectral variations pose a common challenge in analyzing hyperspectral images (HSI). To address this, low-rank tensor representation has emerged as a robust strategy, leveraging inherent correlations within HSI data. However, the spatial distribution of ground objects in HSIs is inherently irregular, existing naturally in tensor format, with numerous class-specific regions manifesting as irregular tensors. Current low-rank representation techniques are designed for regular tensor structures and overlook this fundamental irregularity in real-world HSIs, leading to performance limitations. To tackle this issue, we propose a novel model for irregular tensor low-rank representation tailored to efficiently model irregular 3D cubes. By incorporating a non-convex nuclear norm to promote low-rankness and integrating a global negative low-rank term to enhance the discriminative ability, our proposed model is formulated as a constrained optimization problem and solved using an alternating augmented Lagrangian method. Experimental validation conducted on four public datasets demonstrates the superior performance of our method compared to existing state-of-the-art approaches. The code is publicly available at https://github.com/hb-studying/ITLRR.
Bo Han 0017, Yuheng Jia, Hui Liu 0032, Junhui Hou
IEEE Trans. Image Process.4
2025 Structural-Spectral Graph Convolution With Evidential Edge Learning for Hyperspectral Image Clustering
abstract
Hyperspectral image (HSI) clustering groups pixels into clusters without labeled data, which is an important yet challenging task. For large-scale HSIs, most methods rely on superpixel segmentation and perform superpixel-level clustering based on graph neural networks (GNNs). However, existing GNNs cannot fully exploit the spectral information of the input HSI, and the inaccurate superpixel topological graph may lead to the confusion of different class semantics during information aggregation. To address these challenges, we first propose a structural-spectral graph convolutional operator (SSGCO) tailored for graph-structured HSI superpixels to improve their representation quality through the co-extraction of spatial and spectral features. Second, we propose an evidence-guided adaptive edge learning (EGAEL) module that adaptively predicts and refines edge weights in the superpixel topological graph. We integrate the proposed method into a contrastive learning framework to achieve clustering, where representation learning and clustering are simultaneously conducted. Experiments demonstrate that the proposed method improves clustering accuracy by 2.61%, 6.06%, 4.96% and 3.15% over the best compared methods on four HSI datasets. Our code is available at https://github.com/jhqi/SSGCO-EGAEL.
Jianhan Qi, Yuheng Jia, Hui Liu 0032, Junhui Hou
IEEE Trans. Image Process.4
2025 Modeling State Shifting via Local-Global Distillation for Event-Frame Gaze Tracking
abstract
This paper tackles the problem of passive gaze estimation using both event and frame (or 2D image) data. Considering the inherently different physiological structures, it is intractable to accurately estimate gaze purely based on a given state. Thus, we reformulate gaze estimation as the quantification of the state shifting from the current state to several prior registered anchor states. Specifically, we propose a two-stage learning-based gaze estimation framework that divides the whole gaze estimation process into a coarse-to-fine approach involving anchor state selection and final gaze location. Moreover, to improve the generalization ability, instead of learning a large gaze estimation network directly, we align a group of local experts with a student network, where a novel denoising distillation algorithm is introduced to utilize denoising diffusion techniques to iteratively remove inherent noise in event data. Extensive experiments demonstrate the effectiveness of the proposed method, which surpasses state-of-the-art methods by a large margin of 15$\%$. The code will be publicly available athttps://github.com/ZHU-Zhiyu/Event_Gaze_Tracking.
Jinhui Hou, Jiading Li, Jinjian Wu, Junhui Hou
IEEE Trans. Mob. Comput.5
2025 RBFIM: Perceptual Quality Assessment for Compressed Point Clouds Using Radial Basis Function Interpolation
abstract
One of the main challenges in point cloud compression (PCC) is how to evaluate the perceived distortion so that the codec can be optimized for perceptual quality. Current standard practices in PCC highlight a primary issue: while single-feature metrics are widely used to assess compression distortion, the classic method of searching point-to-point nearest neighbors frequently fails to adequately build precise correspondences between point clouds, resulting in an ineffective capture of human perceptual features. To overcome the related limitations, we propose a novel assessment method called RBFIM, utilizing radial basis function (RBF) interpolation to convert discrete point features into a continuous feature function for the distorted point cloud. By substituting the geometry coordinates of the original point cloud into the feature function, we obtain the bijective sets of point features. This enables an establishment of precise corresponding features between distorted and original point clouds and significantly improves the accuracy of quality assessments. Moreover, this method avoids the complexity caused by bidirectional searches. Extensive experiments on multiple subjective quality datasets of compressed point clouds demonstrate that our RBFIM excels in addressing human perception tasks, thereby providing robust support for PCC optimization efforts.
Shuai Wan, Fuzheng Yang 0001, Mengting Yu, Junhui Hou
IEEE Trans. Multim.6
2025 Leveraging Single-View Images for Unsupervised 3D Point Cloud Completion
abstract
Point clouds captured by scanning devices are often incomplete due to occlusion. To overcome this limitation, point cloud completion methods have been developed to predict the complete shape of an object based on its partial input. These methods can be broadly classified as supervised or unsupervised. However, both categories require a large number of 3D complete point clouds, which may be difficult to capture. In this paper, we propose Cross-PCC, an unsupervised point cloud completion method without requiring any 3D complete point clouds. We only utilize 2D images of the complete objects, which are easier to capture than 3D complete and clean point clouds. Specifically, to take advantage of the complementary information from 2D images, we use a single-view RGB image to extract 2D features and design a fusion module to fuse the 2D and 3D features extracted from the partial point cloud. To guide the shape of predicted point clouds, we project the predicted points of the object to the 2D plane and use the foreground pixels of its silhouette maps to constrain the position of the projected points. To reduce the outliers of the predicted point clouds, we propose a view calibrator to move the points projected to the background into the foreground by the single-view silhouette image. To the best of our knowledge, our approach is the first point cloud completion method that does not require any 3D supervision. The experimental results of our method are superior to those of the state-of-the-art unsupervised methods by a large margin. Moreover, our method even achieves comparable performance to some supervised methods. We will make the source code publicly available athttps://github.com/ltwu6/cross-pcc
Lintai Wu, Qijian Zhang, Junhui Hou, Yong Xu 0001
IEEE Trans. Multim.3
2025 PointMCD: Boosting Deep Point Cloud Encoders via Multi-View Cross-Modal Distillation for 3D Shape Recognition
abstract
As two fundamental representation modalities of 3D objects, 3D point clouds and multi-view 2D images record shape information from different domains of geometric structures and visual appearances. In the current deep learning era, remarkable progress in processing such two data modalities has been achieved through respectively customizing compatible 3D and 2D network architectures. However, unlike multi-view image-based 2D visual modeling paradigms, which have shown leading performance in several common 3D shape recognition benchmarks, point cloud-based 3D geometric modeling paradigms are still highly limited by insufficient learning capacity due to the difficulty of extracting discriminative features from irregular geometric signals. In this article, we explore the possibility of boosting deep 3D point cloud encoders by transferring visual knowledge extracted from deep 2D image encoders under a standard teacher-student distillation workflow. Generally, we propose PointMCD, a unified multi-view cross-modal distillation architecture, including a pretrained deep image encoder as the teacher and a deep point encoder as the student. To perform heterogeneous feature alignment between 2D visual and 3D geometric domains, we further investigate visibility-aware feature projection (VAFP), by which point-wise embeddings are reasonably aggregated into view-specific geometric descriptors. By pair-wisely aligning multi-view visual and geometric descriptors, we can obtain more powerful deep point encoders without exhausting and complicated network modification. Experiments on 3D shape classification, part segmentation, and unsupervised learning strongly validate the effectiveness of our method.The code and data will be publicly available athttps://github.com/keeganhk/PointMCD.
Qijian Zhang, Junhui Hou
IEEE Trans. Multim.2
2025 CS-Net: Contribution-Based Sampling Network for Point Cloud Simplification
abstract
Point cloud sampling plays a crucial role in reducing computation costs and storage requirements for various vision tasks. Traditional sampling methods, such as farthest point sampling, lack task-specific information and, as a result, cannot guarantee optimal performance in specific applications. Learning-based methods train a network to sample the point cloud for the targeted downstream task. However, they do not guarantee that the sampled points are the most relevant ones. Moreover, they may result in duplicate sampled points, which requires completion of the sampled point cloud through post-processing techniques. To address these limitations, we propose a contribution-based sampling network (CS-Net), where the sampling operation is formulated as a Top-$k$k operation. To ensure that the network can be trained in an end-to-end way using gradient descent algorithms, we use a differentiable approximation to the Top-$k$k operation via entropy regularization of an optimal transport problem. Our network consists of a feature embedding module, a cascade attention module, and a contribution scoring module. The feature embedding module includes a specifically designed spatial pooling layer to reduce parameters while preserving important features. The cascade attention module combines the outputs of three skip connected offset attention layers to emphasize the attractive features and suppress less important ones. The contribution scoring module generates a contribution score for each point and guides the sampling process to prioritize the most important ones. Experiments on the ModelNet40 and PU147 showed that CS-Net achieved state-of-the-art performance in two semantic-based downstream tasks (classification and registration) and two reconstruction-based tasks (compression and surface reconstruction). CS-Net also achieved high average precision for objection detection on the KITTI LiDAR point cloud dataset, demonstrating its effectiveness in three-dimensional object detection.
Chen Chen 0063, Hui Yuan 0001, Xiaolong Mao, Raouf Hamzaoui, Junhui Hou
IEEE Trans. Vis. Comput. Graph.6
2025 On Optimal Sampling for Learning SDF Using MLPs Equipped With Positional Encoding
abstract
Neural implicit fields, such as the neural signed distance field (SDF) of a shape, have emerged as a powerful representation for many applications, e.g., encoding a 3D shape and performing collision detection. Typically, implicit fields are encoded by Multi-layer Perceptrons (MLP) with positional encoding (PE) to capture high-frequency geometric details. However, a notable side effect of such PE-equipped MLPs is the noisy artifacts present in the learned implicit fields. While increasing the sampling rate could in general mitigate these artifacts, in this paper we aim to explain this adverse phenomenon through the lens of Fourier analysis. We devise a tool to determine the appropriate sampling rate for learning an accurate neural implicit field without undesirable side effects. Specifically, we propose a simple yet effective method to estimate the intrinsic frequency of a given network with randomized weights based on the Fourier analysis of the network's responses. It is observed that a PE-equipped MLP has an intrinsic frequency much higher than the highest frequency component in the PE layer. Sampling against this intrinsic frequency following the Nyquist-Sannon sampling theorem allows us to determine an appropriate training sampling rate. We empirically show in the setting of SDF fitting that this recommended sampling rate is sufficient to secure accurate fitting results, while further increasing the sampling rate would not further noticeably reduce the fitting error. Training PE-equipped MLPs simply with our sampling strategy leads to performances superior to the existing methods.
Guying Lin, Lei Yang 0048, Yuan Liu 0025, Congyi Zhang 0001, Junhui Hou, Xiaogang Jin 0001, Taku Komura, John Keyser, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.5
2025 Unsupervised 3D Point Cloud Completion via Multi-View Adversarial Learning
abstract
In real-world scenarios, scanned point clouds are often incomplete due to occlusion issues. The tasks of self-supervised and weakly-supervised point cloud completion involve reconstructing missing regions of these incomplete objects without the supervision of complete ground truth. Current methods either rely on multiple views of partial observations for supervision or overlook the intrinsic geometric similarity that can be identified and utilized from the given partial point clouds. In this paper, we propose MAL-UPC, a framework that effectively leverages both region-level and category-specific geometric similarities to complete missing structures. Our MAL-UPC does not require any 3D complete supervision and only necessitates single-view partial observations in the training set. Specifically, we first introduce a Pattern Retrieval Network to retrieve similar position and curvature patterns between the partial input and the predicted shape, then leverage these similarities to densify and refine the reconstructed results. Additionally, we render the reconstructed complete shape into multi-view depth maps and design an adversarial learning module to learn the geometry of the target shape from category-specific single-view depth images of the partial point clouds in the training set. To achieve anisotropic rendering, we design a density-aware radius estimation algorithm to improve the quality of the rendered images. Our MAL-UPC outperforms current state-of-the-art self-supervised methods and even some unpaired approaches.
Lintai Wu, Xianjing Cheng, Yong Xu 0001, Huanqiang Zeng, Junhui Hou
IEEE Trans. Vis. Comput. Graph.5
2025 3D Shape Completion on Unseen Categories: A Weakly-Supervised Approach
abstract
3D shapes captured by scanning devices are often incomplete due to occlusion. 3D shape completion methods have been explored to tackle this limitation. However, most of these methods are only trained and tested on a subset of categories, resulting in poor generalization to unseen categories. In this article, we propose a novel weakly-supervised framework to reconstruct the complete shapes from unseen categories. We first propose an end-to-end prior-assisted shape learning network that leverages data from the seen categories to infer a coarse shape. Specifically, we construct a prior bank consisting of representative shapes from the seen categories. Then, we design a multi-scale pattern correlation module for learning the complete shape of the input by analyzing the correlation between local patterns within the input and the priors at various scales. In addition, we propose a self-supervised shape refinement model to further refine the coarse shape. Considering the shape variability of 3D objects across categories, we construct a category-specific prior bank to facilitate shape refinement. Then, we devise a voxel-based partial matching loss and leverage the partial scans to drive the refinement process. Extensive experimental results show that our approach is superior to state-of-the-art methods by a large margin.
Lintai Wu, Junhui Hou, Linqi Song, Yong Xu 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Decoupling Dynamic Monocular Videos for Dynamic View Synthesis
abstract
The challenge of dynamic view synthesis from dynamic monocular videos, i.e., synthesizing novel views for free viewpoints given a monocular video of a dynamic scene captured by a moving camera, mainly lies in accurately modeling the dynamic objects of a scene using limited 2D frames, each with a varying timestamp and viewpoint. Existing methods usually require pre-processed 2D optical flow and depth maps by off-the-shelf methods to supervise the network, making them suffer from the inaccuracy of the pre-processed supervision and the ambiguity when lifting the 2D information to 3D. In this paper, we tackle this challenge in an unsupervised fashion. Specifically, we decouple the motion of the dynamic objects into object motion and camera motion, respectively regularized by proposed unsupervised surface consistency and patch-based multi-view constraints. The former enforces the 3D geometric surfaces of moving objects to be consistent over time, while the latter regularizes their appearances to be consistent across different viewpoints. Such a fine-grained motion formulation can alleviate the learning difficulty for the network, thus enabling it to produce not only novel views with higher quality but also more accurate scene flows and depth than existing methods requiring extra supervision.
Meng You, Junhui Hou
IEEE Trans. Vis. Comput. Graph.2
2024 Segment Any Event Streams via Weighted Adaptation of Pivotal Tokens
abstract
In this paper, we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data, with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is the precise alignment and calibration of embeddings derived from event-centric data such that they harmoniously coincide with those originating from RGB imagery. Capitalizing on the vast repositories of datasets with paired events and RGB images, our proposition is to harness and extrapolate the profound knowledge encapsulated within the pretrained SAM framework. As a cornerstone to achieving this, we introduce a multi-scale feature distillation methodology. This methodology rigorously optimizes the alignment of token embeddings originating from event data with their RGB image counterparts, thereby preserving and enhancing the robustness of the overall architecture. Considering the distinct significance that token embeddings from intermediate layers hold for higher-level embeddings, our strategy is centered on accurately calibrating the pivotal token embeddings. This targeted calibration is aimed at effectively managing the discrepancies in high-level embeddings originating from both the event and image domains. Extensive experiments on different datasets demonstrate the effectiveness of the proposed distillation method. Code in https://github.com/happychenpipi/EventSAM.
Zhiwen Chen 0002, Yifan Zhang 0036, Junhui Hou, Guangming Shi, Jinjian Wu
CVPR4
2024 DynoSurf: Neural Deformation-Based Temporally Consistent Dynamic Surface Reconstruction
Yuxin Yao 0001, Junhui Hou, Juyong Zhang, Wenping Wang 0001
ECCV (33)3
2024 FairMatch: Promoting Partial Label Learning by Unlabeled Samples
abstract
This paper studies the semi-supervised partial label learning (SSPLL) problem, which aims to improve the partial label learning (PLL) by leveraging unlabeled samples. Both the existing SSPLL methods and the semi-supervised learning methods exploit the information in unlabeled samples by selecting high-confidence unlabeled samples as the pseudo labels based on the maximum value of the model output. However, the scarcity of labeled samples and the ambiguity from partial labels skew this strategy towards an unfair selection of high-confidence samples on each class, most notably during the initial phases of training, resulting in slower training and performance degradation. In this paper, we propose a novel method FairMatch, which adopts a learning state aware self-adaptive threshold for selecting the same number of high-confidence samples on each class, and uses augmentation consistency to incorporate the unlabeled samples to promote PLL. In addition, we adopt the candidate label disambiguation to utilize the partial labeled samples and mix up the partial labeled samples and the selected high-confidence unlabeled samples to prevent the model from overfitting on partial label samples. FairMatch can achieve maximum accuracy improvements of 9.53%, 4.9%, and 16.45% on CIFAR-10, CIFAR-100, and CIFAR-100H, respectively. The codes can be found at https://github.com/jhjiangSEU/FairMatch.
Yuheng Jia, Hui Liu 0032, Junhui Hou
KDD4
2024 Noisy Label Removal for Partial Multi-Label Learning
abstract
This paper addresses the problem of partial multi-label learning (PML), a challenging weakly supervised learning framework, where each sample is associated with a candidate label set comprising both ground-true labels and noisy labels. We theoretically reveal that an increased number of noisy labels in the candidate label set leads to an enlarged generalization error bound, consequently degrading the classification performance. Accordingly, the key to solving PML lies in accurately removing the noisy labels within the candidate label set. To achieve this objective, we leverage prior knowledge about the noisy labels in PML, which suggests that they only exist within the candidate label set and possess binary values. Specifically, we propose a constrained regression model to learn a PML classifier and select the noisy labels. The constraints of the model strictly enforce the location and value of the noisy labels. Simultaneously, the supervision information provided by the candidate label set is unreliable due to the presence of noisy labels. In contrast, the non-candidate labels of a sample precisely indicate the classes to which the sample does not belong. To aid in the selection of noisy labels, we construct a competitive classifier based on the non-candidate labels. The PML classifier and the competitive classifier form a competitive relationship, encouraging mutual learning. We formulate the proposed model as a discrete optimization problem to effectively remove the noisy labels, and we solve it using an alternative algorithm. Extensive experiments conducted on 6 real-world partial multi-label data sets and 7 synthetic data sets, employing various evaluation metrics, demonstrate that our method significantly outperforms state-of-the-art PML methods. The code implementation is publicly available at https://github.com/Yangfc-ML/NLR.
Fuchao Yang, Yuheng Jia, Hui Liu 0032, Yongqiang Dong, Junhui Hou
KDD5
2024 RainyScape: Unsupervised Rainy Scene Reconstruction using Decoupled Neural Rendering
Xianqiang Lyu, Hui Liu 0032, Junhui Hou
ACM Multimedia3
2024 Rethinking the One-shot Object Detection: Cross-Domain Object Search
abstract
One-shot object detection (OSOD) uses a query patch to identify the same category of object in a target image. As the OSOD setting, the target images are required to contain the object category of the query patch, and the image styles (domains) of the query patch and target images are always similar. However, in practical application, the above requirements are not commonly satisfied. Therefore, we propose a new problem namely Cross-Domain Object Search (CDOS), where the object categories of the query patch and target image are decoupled, and the image styles between them may also be significantly different. For this problem, we develop a new method, which incorporates both foreground-background contrastive learning heads and a domain-generalized feature augmentation technique. This makes our method effectively handle the object category gap and domain distribution gap, between the query patch and target image in the training and testing datasets. We further build a new benchmark for the proposed CDOS problem, on which our method shows significant performance improvements over the comparison methods.
Shuqi Zheng, Rui-Ze Han, Yuzhong Feng, Junhui Hou, Linqi Song, Wei Feng 0005
ACM Multimedia5
2024 PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference
abstract
In this paper, we make the first attempt to align diffusion models for image inpainting with human aesthetic standards via a reinforcement learning framework, significantly improving the quality and visual appeal of inpainted images. Specifically, instead of directly measuring the divergence with paired images, we train a reward model with the dataset we construct, consisting of nearly 51,000 images annotated with human preferences. Then, we adopt a reinforcement learning process to fine-tune the distribution of a pre-trained diffusion model for image inpainting in the direction of higher reward. Moreover, we theoretically deduce the upper bound on the error of the reward model, which illustrates the potential confidence of reward estimation throughout the reinforcement alignment process, thereby facilitating accurate regularization. Extensive experiments on inpainting comparison and downstream tasks, such as image extension and 3D reconstruction, demonstrate the effectiveness of our approach, showing significant improvements in the alignment of inpainted images with human preference compared with state-of-the-art methods. This research not only advances the field of image inpainting but also provides a framework for incorporating human preference into the iterative refinement of generative models based on modeling reward accuracy, with broad implications for the design of visually driven AI applications. Our code and dataset are publicly available at \url{https://prefpaint.github.io}.
Kendong Liu, Chuanhao Li 0002, Hui Liu 0032, Huanqiang Zeng, Junhui Hou
NeurIPS6
2024 E-Motion: Future Motion Simulation via Event Sequence Diffusion
abstract
Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict future motion with a level of detail and precision previously unachievable. Inspired by that, we propose to integrate the strong learning capacity of the video diffusion model with the rich motion information of an event camera as a motion simulation framework. Specifically, we initially employ pre-trained stable video diffusion models to adapt the event sequence dataset. This process facilitates the transfer of extensive knowledge from RGB videos to an event-centric domain. Moreover, we introduce an alignment mechanism that utilizes reinforcement learning techniques to enhance the reverse generation trajectory of the diffusion model, ensuring improved performance and accuracy. Through extensive testing and validation, we demonstrate the effectiveness of our method in various complex scenarios, showcasing its potential to revolutionize motion flow prediction in computer vision applications such as autonomous vehicle guidance, robotic navigation, and interactive media. Our findings suggest a promising direction for future research in enhancing the interpretative power and predictive accuracy of computer vision systems. The source code is publicly available at https://github.com/p4r4mount/E-Motion.
Junhui Hou, Guangming Shi, Jinjian Wu
NeurIPS3
2024 Fine-grained Image-to-LiDAR Contrastive Distillation with Visual Foundation Models
abstract
Contrastive image-to-LiDAR knowledge transfer, commonly used for learning 3D representations with synchronized images and point clouds, often faces a self-conflict dilemma. This issue arises as contrastive losses unintentionally dissociate features of unmatched points and pixels that share semantic labels, compromising the integrity of learned representations. To overcome this, we harness Visual Foundation Models (VFMs), which have revolutionized the acquisition of pixel-level semantics, to enhance 3D representation learning. Specifically, we utilize off-the-shelf VFMs to generate semantic labels for weakly-supervised pixel-to-point contrastive distillation. Additionally, we employ von Mises-Fisher distributions to structure the feature space, ensuring semantic embeddings within the same class remain consistent across varying inputs. Furthermore, we adapt sampling probabilities of points to address imbalances in spatial distribution and category frequency, promoting comprehensive and balanced learning. Extensive experiments demonstrate that our approach mitigates the challenges posed by traditional methods and consistently surpasses existing image-to-LiDAR contrastive distillation methods in downstream tasks. We have included the code in supplementary materials.
Yifan Zhang 0036, Junhui Hou
NeurIPS2
2024 Flatten Anything: Unsupervised Neural Surface Parameterization
abstract
Surface parameterization plays an essential role in numerous computer graphics and geometry processing applications. Traditional parameterization approaches are designed for high-quality meshes laboriously created by specialized 3D modelers, thus unable to meet the processing demand for the current explosion of ordinary 3D data. Moreover, their working mechanisms are typically restricted to certain simple topologies, thus relying on cumbersome manual efforts (e.g., surface cutting, part segmentation) for pre-processing. In this paper, we introduce the Flatten Anything Model (FAM), an unsupervised neural architecture to achieve global free-boundary surface parameterization via learning point-wise mappings between 3D points on the target geometric surface and adaptively-deformed UV coordinates within the 2D parameter domain. To mimic the actual physical procedures, we ingeniously construct geometrically-interpretable sub-networks with specific functionalities of surface cutting, UV deforming, unwrapping, and wrapping, which are assembled into a bi-directional cycle mapping framework. Compared with previous methods, our FAM directly operates on discrete surface points without utilizing connectivity information, thus significantly reducing the strict requirements for mesh quality and even applicable to unstructured point cloud data. More importantly, our FAM is fully-automated without the need for pre-cutting and can deal with highly-complex topologies, since its learning process adaptively finds reasonable cutting seams and UV boundaries. Extensive experiments demonstrate the universality, superiority, and inspiring potential of our proposed neural surface parameterization paradigm. Our code is available at https://github.com/keeganhk/FlattenAnything.
Qijian Zhang, Junhui Hou, Wenping Wang 0001, Ying He 0001
NeurIPS2
2024 Probabilistic-Based Feature Embedding of 4-D Light Fields for Compressive Imaging and Denoising
Xianqiang Lyu, Junhui Hou
Int. J. Comput. Vis.2
2024 A Comprehensive Study of the Robustness for LiDAR-Based 3D Object Detectors Against Adversarial Attacks
Yifan Zhang 0036, Junhui Hou, Yixuan Yuan
Int. J. Comput. Vis.2
2024 Unsupervised video-based action recognition using two-stream generative adversarial network
Wei Lin 0021, Huanqiang Zeng, Jianqing Zhu, Chih-Hsien Hsia, Junhui Hou, Kai-Kuang Ma
Neural Comput. Appl.5
2024 Deep Diversity-Enhanced Feature Representation of Hyperspectral Images
abstract
In this paper, we study the problem of efficiently and effectively embedding the high-dimensional spatio-spectral information of hyperspectral (HS) images, guided by feature diversity. Specifically, based on the theoretical formulation that feature diversity is correlated with the rank of the unfolded kernel matrix, we rectify 3D convolution by modifying its topology to enhance the rank upper-bound. This modification yields a rank-enhanced spatial-spectral symmetrical convolution set (ReS$^{3}$-ConvSet), which not only learns diverse and powerful feature representations but also saves network parameters. Additionally, we also propose a novel diversity-aware regularization (DA-Reg) term that directly acts on the feature maps to maximize independence among elements. To demonstrate the superiority of the proposed ReS$^{3}$-ConvSet and DA-Reg, we apply them to various HS image processing and analysis tasks, including denoising, spatial super-resolution, and classification. Extensive experiments show that the proposed approaches outperform state-of-the-art methods both quantitatively and qualitatively to a significant extent. The code is publicly available athttps://github.com/jinnh/ReSSS-ConvSet.
Jinhui Hou, Junhui Hou, Hui Liu 0032, Huanqiang Zeng, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Dynamic 3D Point Cloud Sequences as 2D Videos
abstract
Dynamic 3D point cloud sequences serve as one of the most common and practical representation modalities of dynamic real-world environments. However, their unstructured nature in both spatial and temporal domains poses significant challenges to effective and efficient processing. Existing deep point cloud sequence modeling approaches imitate the mature 2D video learning mechanisms by developing complex spatio-temporal point neighbor grouping and feature aggregation schemes, often resulting in methods lacking effectiveness, efficiency, and expressive power. In this paper, we propose a novel generic representation called Structured Point Cloud Videos (SPCVs). Intuitively, by leveraging the fact that 3D geometric shapes are essentially 2D manifolds, SPCV re-organizes a point cloud sequence as a 2D video with spatial smoothness and temporal consistency, where the pixel values correspond to the 3D coordinates of points. The structured nature of our SPCV representation allows for the seamless adaptation of well-established 2D image/video techniques, enabling efficient and effective processing and analysis of 3D point cloud sequences. To achieve such re-organization, we design a self-supervised learning pipeline that is geometrically regularized and driven by self-reconstructive and deformation field learning objectives. Additionally, we construct SPCV-based frameworks for both low-level and high-level 3D point cloud sequence processing and analysis tasks, including action recognition, temporal interpolation, and compression. Extensive experiments demonstrate the versatility and superiority of the proposed SPCV, which has the potential to offer new possibilities for deep learning on unstructured 3D point cloud sequences.
Yiming Zeng 0002, Junhui Hou, Qijian Zhang, Wenping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Spatial-Temporal Graph Enhanced DETR Towards Multi-Frame 3D Object Detection
abstract
The Detection Transformer (DETR) has revolutionized the design of CNN-based object detection systems, showcasing impressive performance. However, its potential in the domain of multi-frame 3D object detection remains largely unexplored. In this paper, we present STEMD, a novel end-to-end framework that enhances the DETR-like paradigm for multi-frame 3D object detection by addressing three key aspects specifically tailored for this task. First, to model the inter-object spatial interaction and complex temporal dependencies, we introduce the spatial-temporal graph attention network, which represents queries as nodes in a graph and enables effective modeling of object interactions within a social context. To solve the problem of missing hard cases in the proposed output of the encoder in the current frame, we incorporate the output of the previous frame to initialize the query input of the decoder. Finally, it poses a challenge for the network to distinguish between the positive query and other highly similar queries that are not the best match. And similar queries are insufficiently suppressed and turn into redundant prediction boxes. To address this issue, our proposed IoU regularization term encourages similar queries to be distinct during the refinement. Through extensive experiments, we demonstrate the effectiveness of our approach in handling challenging scenarios, while incurring only a minor additional computational overhead.
Yifan Zhang 0036, Junhui Hou, Dapeng Oliver Wu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Superpixel Graph Contrastive Clustering With Semantic-Invariant Augmentations for Hyperspectral Images
abstract
Hyperspectral images (HSI) clustering is an important but challenging task. The state-of-the-art (SOTA) methods usually rely on superpixels, however, they do not fully utilize the spatial and spectral information in HSI 3-D structure, and their optimization targets are not clustering-oriented. In this work, we first use 3-D and 2-D hybrid convolutional neural networks to extract the high-order spatial and spectral features of HSI through pre-training, and then design a superpixel graph contrastive clustering (SPGCC) model to learn discriminative superpixel representations. Reasonable augmented views are crucial for contrastive clustering, and conventional contrastive learning may hurt the cluster structure since different samples are pushed away in the embedding space even if they belong to the same class. In SPGCC, we design two semantic-invariant data augmentations for HSI superpixels: pixel sampling augmentation and model weight augmentation. Then sample-level alignment and clustering-center-level contrast are performed for better intra-class similarity and inter-class dissimilarity of superpixel embeddings. We perform clustering and network optimization alternatively. Experimental results on several HSI datasets verify the advantages of the proposed SPGCC compared to SOTA methods. Our code is available athttps://github.com/jhqi/spgcc.
Jianhan Qi, Yuheng Jia, Hui Liu 0032, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.4
2024 Enhancing Low-Light Light Field Images With a Deep Compensation Unfolding Network
abstract
This paper presents a novel and interpretable end-to-end learning framework, called the deep compensation unfolding network (DCUNet), for restoring light field (LF) images captured under low-light conditions. DCUNet is designed with a multi-stage architecture that mimics the optimization process of solving an inverse imaging problem in a data-driven fashion. The framework uses the intermediate enhanced result to estimate the illumination map, which is then employed in the unfolding process to produce a new enhanced result. Additionally, DCUNet includes a content-associated deep compensation module at each optimization stage to suppress noise and illumination map estimation errors. To properly mine and leverage the unique characteristics of LF images, this paper proposes a pseudo-explicit feature interaction module that comprehensively exploits redundant information in LF images. The experimental results on both simulated and real datasets demonstrate the superiority of our DCUNet over state-of-the-art methods, both qualitatively and quantitatively. Moreover, DCUNet preserves the essential geometric structure of enhanced LF images much better. The code is publicly available at https://github.com/lyuxianqiang/LFLL-DCU.
Xianqiang Lyu, Junhui Hou
IEEE Trans. Image Process.2
2024 PointVST: Self-Supervised Pre-Training for 3D Point Clouds via View-Specific Point-to-Image Translation
abstract
The past few years have witnessed the great success and prevalence of self-supervised representation learning within the language and 2D vision communities. However, such advancements have not been fully migrated to the field of 3D point cloud learning. Different from existing pre-training paradigms designed for deep point cloud feature extractors that fall into the scope of generative modeling or contrastive learning, this paper proposes a translative pre-training framework, namely PointVST, driven by a novel self-supervised pretext task of cross-modal translation from 3D point clouds to their corresponding diverse forms of 2D rendered images. More specifically, we begin with deducing view-conditioned point-wise embeddings through the insertion of the viewpoint indicator, and then adaptively aggregate a view-specific global codeword, which can be further fed into subsequent 2D convolutional translation heads for image generation. Extensive experimental evaluations on various downstream task scenarios demonstrate that our PointVST shows consistent and prominent performance superiority over current state-of-the-art approaches as well as satisfactory domain transfer capability.
Qijian Zhang, Junhui Hou
IEEE Trans. Vis. Comput. Graph.2
2024 Light Field Depth Estimation via Stitched Epipolar Plane Images
abstract
Depth estimation is a fundamental problem in light field processing. Epipolar-plane image (EPI)-based methods often encounter challenges such as low accuracy in slope computation due to discretization errors and limited angular resolution. Besides, existing methods perform well in most regions but struggle to produce sharp edges in occluded regions and resolve ambiguities in texture-less regions. To address these issues, we propose the concept of stitched-EPI (SEPI) to enhance slope computation. SEPI achieves this by shifting and concatenating lines from different EPIs that correspond to the same 3D point. Moreover, we introduce the half-SEPI algorithm, which focuses exclusively on the non-occluded portion of lines to handle occlusion. Additionally, we present a depth propagation strategy aimed at improving depth estimation in texture-less regions. This strategy involves determining the depth of such regions by progressing from the edges towards the interior, prioritizing accurate regions over coarse regions. Through extensive experimental evaluations and ablation studies, we validate the effectiveness of our proposed method. The results demonstrate its superior ability to generate more accurate and robust depth maps across all regions compared to state-of-the-art methods.
Langqing Shi, Xiaoyang Liu 0013, Jing Jin 0006, Junhui Hou
IEEE Trans. Vis. Comput. Graph.6
2023 PointCA: Evaluating the Robustness of 3D Point Cloud Completion Models against Adversarial Examples
abstract
Point cloud completion, as the upstream procedure of 3D recognition and segmentation, has become an essential part of many tasks such as navigation and scene understanding. While various point cloud completion models have demonstrated their powerful capabilities, their robustness against adversarial attacks, which have been proven to be fatally malicious towards deep neural networks, remains unknown. In addition, existing attack approaches towards point cloud classifiers cannot be applied to the completion models due to different output forms and attack purposes. In order to evaluate the robustness of the completion models, we propose PointCA, the first adversarial attack against 3D point cloud completion models. PointCA can generate adversarial point clouds that maintain high similarity with the original ones, while being completed as another object with totally different semantic information. Specifically, we minimize the representation discrepancy between the adversarial example and the target point set to jointly explore the adversarial point clouds in the geometry space and the feature space. Furthermore, to launch a stealthier attack, we innovatively employ the neighbourhood density information to tailor the perturbation constraint, leading to geometry-aware and distribution-adaptive modifications for each point. Extensive experiments against different premier point cloud completion networks show that PointCA can cause the performance degradation from 77.9% to 16.7%, with the structure chamfer distance kept below 0.01. We conclude that existing completion models are severely vulnerable to adversarial examples, and state-of-the-art defenses for point cloud classification will be partially invalid when applied to incomplete and uneven point cloud data.
Shengshan Hu, Wei Liu 0004, Junhui Hou, Leo Yu Zhang, Hai Jin 0001, Lichao Sun 0001
AAAI4
2023 GeoUDF: Surface Reconstruction from 3D Point Clouds via Geometry-guided Distance Representation
abstract
We present a learning-based method, namely GeoUDF, to tackle the long-standing and challenging problem of reconstructing a discrete surface from a sparse point cloud. To be specific, we propose a geometry-guided learning method for UDF and its gradient estimation that explicitly formulates the unsigned distance of a query point as the learnable affine averaging of its distances to the tangent planes of neighboring points on the surface. Besides, we model the local geometric structure of the input point clouds by explicitly learning a quadratic polynomial for each point. This not only facilitates upsampling the input sparse point cloud but also naturally induces unoriented normal, which further augments UDF estimation. Finally, to extract triangle meshes from the predicted UDF we propose a customized edge-based marching cube module. We conduct extensive experiments and ablation studies to demonstrate the significant advantages of our method over state-of-the-art methods in terms of reconstruction accuracy, efficiency, and generality. The source code is publicly available at https://github.com/rsy6318/GeoUDF.
Junhui Hou, Xiaodong Chen 0009, Ying He 0001, Wenping Wang 0001
ICCV2
2023 Downstream-agnostic Adversarial Examples
abstract
Self-supervised learning usually uses a large amount of unlabeled data to pre-train an encoder which can be used as a general-purpose feature extractor, such that downstream users only need to perform fine-tuning operations to enjoy the benefit of "large model". Despite this promising prospect, the security of pre-trained encoder has not been thoroughly investigated yet, especially when the pre-trained encoder is publicly available for commercial use.In this paper, we propose AdvEncoder, the first framework for generating downstream-agnostic universal adversarial examples based on the pre-trained encoder. AdvEncoder aims to construct a universal adversarial perturbation or patch for a set of natural images that can fool all the downstream tasks inheriting the victim pre-trained encoder. Unlike traditional adversarial example works, the pre-trained encoder only outputs feature vectors rather than classification labels. Therefore, we first exploit the high frequency component information of the image to guide the generation of adversarial examples. Then we design a generative attack framework to construct adversarial perturbations/patches by learning the distribution of the attack surrogate dataset to improve their attack success rates and transferability. Our results show that an attacker can successfully attack downstream tasks without knowing either the pre-training dataset or the downstream dataset. We also tailor four defenses for pre-trained encoders, the results of which further prove the attack ability of AdvEncoder. Our codes are available at: https://github.com/CGCL-codes/AdvEncoder.
Ziqi Zhou 0001, Shengshan Hu, Ruizhi Zhao, Qian Wang 0002, Leo Yu Zhang, Junhui Hou, Hai Jin 0001
ICCV6
2023 Cross-modal Orthogonal High-rank Augmentation for RGB-Event Transformer-trackers
abstract
This paper addresses the problem of cross-modal object tracking from RGB videos and event data. Rather than constructing a complex cross-modal fusion network, we explore the great potential of a pre-trained vision Transformer (ViT). Particularly, we delicately investigate plug-and-play training augmentations that encourage the ViT to bridge the vast distribution gap between the two modalities, enabling comprehensive cross-modal information interaction and thus enhancing its ability. Specifically, we propose a mask modeling strategy that randomly masks a specific modality of some tokens to enforce the interaction between tokens from different modalities interacting proactively. To mitigate network oscillations resulting from the masking strategy and further amplify its positive effect, we then theoretically propose an orthogonal high-rank loss to regularize the attention matrix. Extensive experiments demonstrate that our plug-and-play training augmentation techniques can significantly boost state-of-the-art one-stream and two-stream trackers to a large extent in terms of both tracking precision and success rate. Our new perspective and findings will potentially bring insights to the field of leveraging powerful pre-trained ViTs to model cross-modal data. The code is publicly available at https://github.com/ZHU-Zhiyu/High-Rank_RGB-Event_Tracker.
Junhui Hou, Dapeng Oliver Wu
ICCV2
2023 CAS-Net: Cascade Attention-Based Sampling Neural Network for Point Cloud Simplification
abstract
Point cloud sampling can reduce storage requirements and computation costs for various vision tasks. Traditional sampling methods, such as farthest point sampling, are not geared towards downstream tasks and may fail on such tasks. In this paper, we propose a cascade attention-based sampling network (CAS-Net), which is end-to-end trainable. Specifically, we propose an attention-based sampling module (ASM) to capture the semantic features and preserve the geometry of the original point cloud. Experimental results on the ModelNet40 dataset show that CAS-Net outperforms state-of-the-art methods in a sampling-based point cloud classification task, while preserving the geometric structure of the sampled point cloud.
Chen Chen 0063, Hui Yuan 0001, Hao Liu 0044, Junhui Hou, Raouf Hamzaoui
ICME4
2023 PointCRT: Detecting Backdoor in 3D Point Cloud via Corruption Robustness
abstract
Backdoor attacks for point clouds have elicited mounting interest with the proliferation of deep learning. The point cloud classifiers can be vulnerable to malicious actors who seek to manipulate or fool the model with specific backdoor triggers. Detecting and rejecting backdoor samples during the inference stage can effectively alleviate backdoor attacks. Recently, some black-box test-time backdoor sample detection methods have been proposed in the 2D image domain, without any underlying assumptions about the backdoor triggers. However, upon examination, we have found that these detection techniques are not effective for 3D point clouds. As a result, there is a pressing need to bridge the gap for the development of a universal approach that is specifically designed for 3D point clouds.
Shengshan Hu, Wei Liu 0304, Yechao Zhang, Xiaogeng Liu, Xianlong Wang 0001, Leo Yu Zhang, Junhui Hou
ACM Multimedia8
2023 Global Structure-Aware Diffusion Process for Low-light Image Enhancement
abstract
This paper studies a diffusion-based framework to address the low-light image enhancement problem. To harness the capabilities of diffusion models, we delve into this intricate process and advocate for the regularization of its inherent ODE-trajectory. To be specific, inspired by the recent research that low curvature ODE-trajectory results in a stable and effective diffusion process, we formulate a curvature regularization term anchored in the intrinsic non-local structures of image data, i.e., global structure-aware regularization, which gradually facilitates the preservation of complicated details and the augmentation of contrast during the diffusion process. This incorporation mitigates the adverse effects of noise and artifacts resulting from the diffusion process, leading to a more precise and flexible enhancement. To additionally promote learning in challenging regions, we introduce an uncertainty-guided regularization technique, which wisely relaxes constraints on the most extreme regions of the image. Experimental evaluations reveal that the proposed diffusion-based framework, complemented by rank-informed regularization, attains distinguished performance in low-light enhancement. The outcomes indicate substantial advancements in image quality, noise suppression, and contrast amplification in comparison with state-of-the-art methods. We believe this innovative approach will stimulate further exploration and advancement in low-light image processing, with potential implications for other applications of diffusion models. The code is publicly available at https://github.com/jinnh/GSAD.
Jinhui Hou, Junhui Hou, Hui Liu 0032, Huanqiang Zeng, Hui Yuan 0001
NeurIPS3
2023 NeuroGF: A Neural Representation for Fast Geodesic Distance and Path Queries
abstract
Geodesics play a critical role in many geometry processing applications. Traditional algorithms for computing geodesics on 3D mesh models are often inefficient and slow, which make them impractical for scenarios requiring extensive querying of arbitrary point-to-point geodesics. Recently, deep implicit functions have gained popularity for 3D geometry representation, yet there is still no research on neural implicit representation of geodesics. To bridge this gap, we make the first attempt to represent geodesics using implicit learning frameworks. Specifically, we propose neural geodesic field (NeuroGF), which can be learned to encode all-pairs geodesics of a given 3D mesh model, enabling to efficiently and accurately answer queries of arbitrary point-to-point geodesic distances and paths. Evaluations on common 3D object models and real-captured scene-level meshes demonstrate our exceptional performances in terms of representation accuracy and querying efficiency. Besides, NeuroGF also provides a convenient way of jointly encoding both 3D geometry and geodesics in a unified representation. Moreover, the working mode of per-model overfitting is further extended to generalizable learning frameworks that can work on various input formats such as unstructured point clouds, which also show satisfactory performances for unseen shapes and categories. Our code and data are available at https://github.com/keeganhk/NeuroGF.
Qijian Zhang, Junhui Hou, Yohanes Yudhi Adikusuma, Wenping Wang 0001, Ying He 0001
NeurIPS2
2023 Unleash the Potential of Image Branch for Cross-modal 3D Object Detection
abstract
To achieve reliable and precise scene understanding, autonomous vehicles typically incorporate multiple sensing modalities to capitalize on their complementary attributes. However, existing cross-modal 3D detectors do not fully utilize the image domain information to address the bottleneck issues of the LiDAR-based detectors. This paper presents a new cross-modal 3D object detector, namely UPIDet, which aims to unleash the potential of the image branch from two aspects. First, UPIDet introduces a new 2D auxiliary task called normalized local coordinate map estimation. This approach enables the learning of local spatial-aware features from the image modality to supplement sparse point clouds. Second, we discover that the representational capability of the point cloud backbone can be enhanced through the gradients backpropagated from the training objectives of the image branch, utilizing a succinct and effective point-to-pixel module. Extensive experiments and ablation studies validate the effectiveness of our method. Notably, we achieved the top rank in the highly competitive cyclist class of the KITTI benchmark at the time of submission. The source code is available at https://github.com/Eaphan/UPIDet.
Yifan Zhang 0036, Qijian Zhang, Junhui Hou, Yixuan Yuan, Guoliang Xing
NeurIPS3
2023 GLENet: Boosting 3D Object Detectors with Generative Label Uncertainty Estimation
Yifan Zhang 0036, Qijian Zhang, Junhui Hou, Yixuan Yuan
Int. J. Comput. Vis.4
2023 Screen content video quality assessment based on spatiotemporal sparse feature
Huanqiang Zeng, Hailiang Huang 0002, Shan Cheng, Junhui Hou
J. Vis. Commun. Image Represent.6
2023 Content-Aware Warping for View Synthesis
abstract
Existing image-based rendering methods usually adopt depth-based image warping operation to synthesize novel views. In this paper, we reason the essential limitations of the traditional warping operation to be the limited neighborhood and only distance-based interpolation weights. To this end, we propose content-aware warping, which adaptively learns the interpolation weights for pixels of a relatively large neighborhood from their contextual information via a lightweight neural network. Based on this learnable warping module, we propose a new end-to-end learning-based framework for novel view synthesis from a set of input source views, in which two additional modules, namely confidence-based blending and feature-assistant spatial refinement, are naturally proposed to handle the occlusion issue and capture the spatial correlation among pixels of the synthesized view, respectively. Besides, we also propose a weight-smoothness loss term to regularize the network. Experimental results on light field datasets with wide baselines and multi-view datasets show that the proposed method significantly outperforms state-of-the-art methods both quantitatively and visually. The source code is publicly available at https://github.com/MantangGuo/CW4VS.
Mantang Guo, Junhui Hou, Jing Jin 0006, Hui Liu 0032, Huanqiang Zeng, Jiwen Lu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Light Field Reconstruction via Deep Adaptive Fusion of Hybrid Lenses
abstract
This paper explores the problem of reconstructing high-resolution light field (LF) images from hybrid lenses, including a high-resolution camera surrounded by multiple low-resolution cameras. The performance of existing methods is still limited, as they produce either blurry results on plain textured areas or distortions around depth discontinuous boundaries. To tackle this challenge, we propose a novel end-to-end learning-based approach, which can comprehensively utilize the specific characteristics of the input from two complementary and parallel perspectives. Specifically, one module regresses a spatially consistent intermediate estimation by learning a deep multidimensional and cross-domain feature representation, while the other module warps another intermediate estimation, which maintains the high-frequency textures, by propagating the information of the high-resolution view. We finally leverage the advantages of the two intermediate estimations adaptively via the learned confidence maps, leading to the final high-resolution LF image with satisfactory results on both plain textured areas and depth discontinuous boundaries. Besides, to promote the effectiveness of our method trained with simulated hybrid data on real hybrid data captured by a hybrid LF imaging system, we carefully design the network architecture and the training strategy. Extensive experiments on both real and simulated hybrid data demonstrate the significant superiority of our approach over state-of-the-art ones. To the best of our knowledge, this is the first end-to-end deep learning method for LF reconstruction from a real hybrid input. We believe our framework could potentially decrease the cost of high-resolution LF data acquisition and benefit LF data storage and transmission. The code will be publicly available at https://github.com/jingjin25/LFhybridSR-Fusion.
Jing Jin 0006, Mantang Guo, Junhui Hou, Hui Liu 0032, Hongkai Xiong
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Flattening-Net: Deep Regular 2D Representation for 3D Point Cloud Analysis
abstract
Point clouds are characterized by irregularity and unstructuredness, which pose challenges in efficient data exploitation and discriminative feature extraction. In this paper, we present an unsupervised deep neural architecture called Flattening-Net to represent irregular 3D point clouds of arbitrary geometry and topology as a completely regular 2D point geometry image (PGI) structure, in which coordinates of spatial points are captured in colors of image pixels. Intuitively, Flattening-Net implicitly approximates a locally smooth 3D-to-2D surface flattening process while effectively preserving neighborhood consistency. As a generic representation modality, PGI inherently encodes the intrinsic property of the underlying manifold structure and facilitates surface-style point feature aggregation. To demonstrate its potential, we construct a unified learning framework directly operating on PGIs to achieve diverse types of high-level and low-level downstream applications driven by specific task networks, including classification, segmentation, reconstruction, and upsampling. Extensive experiments demonstrate that our methods perform favorably against the current state-of-the-art competitors.
Qijian Zhang, Junhui Hou, Yiming Zeng 0002, Juyong Zhang, Ying He 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Semi-supervised adaptive kernel concept factorization
Wenhui Wu 0001, Junhui Hou, Shiqi Wang 0001, Sam Kwong, Yu Zhou 0027
Pattern Recognit.2
2023 ECSNet: Spatio-Temporal Feature Learning for Event Camera
abstract
The neuromorphic event cameras can efficiently sense the latent geometric structures and motion clues of a scene by generating asynchronous and sparse event signals. Due to the irregular layout of the event signals, how to leverage their plentiful spatio-temporal information for recognition tasks remains a significant challenge. Existing methods tend to treat events as dense image-like or point-serie representations. However, they either suffer from severe destruction on the sparsity of event data or fail to encode robust spatial cues. To fully exploit their inherent sparsity with reconciling the spatio-temporal information, we introduce a compact event representation, namely 2D-1T event cloud sequence (2D-1T ECS). We couple this representation with a novel light-weight spatio-temporal learning framework (ECSNet) that accommodates both object classification and action recognition tasks. The core of our framework is a hierarchical spatial relation module. Equipped with specially designed surface-event-based sampling unit and local event normalization unit to enhance the inter-event relation encoding, this module learns robust geometric features from the 2D event clouds. And we propose a motion attention module for efficiently capturing long-term temporal context evolving with the 1T cloud sequence. Empirically, the experiments show that our framework achieves par or even better state-of-the-art performance. Importantly, our approach cooperates well with the sparsity of event data without any sophisticated operations, hence leading to low computational costs and prominent inference speeds.
Zhiwen Chen 0002, Jinjian Wu, Junhui Hou, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.3
2023 Semi-Supervised Subspace Clustering via Tensor Low-Rank Representation
abstract
In this letter, we propose a novel semi-supervised subspace clustering method, which is able to simultaneously augment the initial supervisory information and construct a discriminative affinity matrix. By representing the limited amount of supervisory information as a pairwise constraint matrix, we observe that the ideal affinity matrix for clustering shares the same low-rank structure as the ideal pairwise constraint matrix. Thus, we stack the two matrices into a 3-D tensor, where a global low-rank constraint is imposed to promote the affinity matrix construction and augment the initial pairwise constraints synchronously. Besides, we use the local geometry structure of input samples to complement the global low-rank prior to achieve better affinity matrix learning. The proposed model is formulated as a Laplacian graph regularized convex low-rank tensor representation problem, which is further solved with an alternative iterative algorithm. In addition, we propose to refine the affinity matrix with the augmented pairwise constraints. Comprehensive experimental results on eight commonly-used benchmark datasets demonstrate the superiority of our method over state-of-the-art methods. The code is publicly available athttps://github.com/GuanxingLu/Subspace-Clustering.
Yuheng Jia, Guanxing Lu, Hui Liu 0032, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.4
2023 Deep Attention-Guided Graph Clustering With Dual Self-Supervision
abstract
Existing deep embedding clustering methods fail to sufficiently utilize the available off-the-shelf information from feature embeddings and cluster assignments, limiting their performance. To this end, we propose a novel method, namely deep attention-guided graph clustering with dual self-supervision (DAGC). Specifically, DAGC first utilizes a heterogeneity-wise fusion module to adaptively integrate the features of the auto-encoder and the graph convolutional network in each layer and then uses a scale-wise fusion module to dynamically concatenate the multi-scale features in different layers. Such modules are capable of learning an informative feature embedding via an attention-based mechanism. In addition, we design a distribution-wise fusion module that leverages cluster assignments to acquire clustering results directly. To better explore the off-the-shelf information from the cluster assignments, we develop a dual self-supervision solution consisting of a soft self-supervision strategy with a Kullback-Leibler divergence loss and a hard self-supervision strategy with a pseudo supervision loss. Extensive experiments on nine benchmark datasets validate that our method consistently outperforms state-of-the-art methods. Especially, our method improves the ARI by more than 10.29% over the best baseline. The code will be publicly available athttps://github.com/ZhihaoPENG-CityU/DAGC.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.4
2023 Task-Oriented Compact Representation of 3D Point Clouds via A Matrix Optimization-Driven Network
abstract
This paper explores the task-oriented compact representation of 3D point clouds, which should maintain the performance of subsequent applications applied to such compact point clouds as much as possible. Designing from the perspective of matrix optimization, we propose MOPS-Net, a novel deep learning-based method that is distinguishable from existing approaches due to its interpretability and flexibility. The matrix optimization problem is challenging due to the discrete and combinatorial nature of the sampling matrix. Therefore, we tackle the challenges by relaxing the binary constraint of the sampling matrix and formulating a constrained and differentiable optimization problem. We then design a deep neural network to mimic the matrix optimization by exploring both the local and global structures of the input data. MOPS-Net can be end-to-end trained with a task network and is permutation-invariant, making it robust to the input. We also extend MOPS-Net such that a single network after one-time training is capable of handling arbitrary downsampling ratios. Extensive experimental results show that MOPS-Net can achieve favorable performance against state-of-the-art deep learning-based methods over various tasks, including classification, reconstruction, and registration. Besides, we validate the robustness of MOPS-Net on noisy data.
Junhui Hou, Qijian Zhang, Yiming Zeng 0002, Sam Kwong, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 CorrI2P: Deep Image-to-Point Cloud Registration via Dense Correspondence
abstract
Motivated by the intuition that the critical step of localizing a 2D image in the corresponding 3D point cloud is establishing 2D-3D correspondence between them, we propose the first feature-based dense correspondence framework for addressing the challenging problem of 2D image-to-3D point cloud registration, dubbed CorrI2P. CorrI2P is mainly composed of three modules, i.e., feature embedding, symmetric overlapping region detection, and pose estimation through the established correspondence. Specifically, given a pair of a 2D image and a 3D point cloud, we first transform them into high-dimensional feature spaces and feed the resulting features into a symmetric overlapping region detector to determine the region where the image and point cloud overlap. Then we use the features of the overlapping regions to establish dense 2D-3D correspondence, on which EPnP within RANSAC is performed to estimate the camera pose, i.e., translation and rotation matrices. Experimental results on KITTI and NuScenes datasets show that our CorrI2P outperforms state-of-the-art image-to-point cloud registration methods significantly. The code will be publicly available athttps://github.com/rsy6318/CorrI2P.
Yiming Zeng 0002, Junhui Hou, Xiaodong Chen 0009
IEEE Trans. Circuits Syst. Video Technol.3
2023 Plausible Proxy Mining With Credibility for Unsupervised Person Re-Identification
abstract
One effective way to address unsupervised person re-identification is to use a clustering-based contrastive learning approach. Existing state-of-the-art methods adopt clustering algorithms (e.g., DBSCAN) and camera ID information to divide all person images into several camera-aware proxies. Then, for each person image, the extracted feature representation is pulled closer to the centroids of its pseudo-positive proxies (the proxies that share the same pseudo-identity label with this image) and pushed away from the centroids of other pseudo-negative proxies (the proxies that share the different pseudo-identity label with this image). However, the quality of the proxy centroid is significantly affected by the proxy impurity issue and thus deteriorates the learned feature representations. On the premise that we cannot introduce superior supervision signals by thoroughly solving the proxy impurity issue, for a person image, identifying its plausible proxies: the pseudo-negative proxies which potentially include its wrongly-clustered instances (the instances with the same ground-truth identity with this image), and further fixing the resulted incorrect supervision signals become an urgent and challenging problem. This paper proposes a simple yet effective approach to address this problem. With a given image, our method can effectively locate its plausible proxies. Then we introduce credibility to measure how much we should treat the centroid of each mined plausible proxy as a positive supervision signal rather than entirely negative. Extensive experiments on three widely-used person re-ID datasets validate the effectiveness of our proposed approach. Codes will be available at:https://github.com/Dingyuan-Zheng/PPCL.
Dingyuan Zheng, Jimin Xiao, Mingjie Sun, Huihui Bai 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.5
2023 t-Linear Tensor Subspace Learning for Robust Feature Extraction of Hyperspectral Images
abstract
Subspace learning has been widely applied for feature extraction of hyperspectral images (HSIs) and achieved great success. However, the current methods still leave two problems that need to be further investigated. First, those methods mainly focus on finding one or multiple projection matrices for mapping the high-dimensional data into a low-dimensional subspace, which can only capture the information from each direction of high-order hyperspectral data separately. Second, the performance of feature extraction is barely satisfactory when the hyperspectral data is severely corrupted by noise. To address these issues, this article presents a t-linear tensor subspace learning (tLTSL) model for robust feature extraction of HSIs based on t-product projection. In the model, t-product projection is a new defined tensor transformation way similar to linear transformation in vector space, which can maximally capture the intrinsic structure of tensor data. The integrated tensor low-rank and sparse decomposition can effectively remove the noise corruption and the learned t-product projection can directly transform the high-order hyperspectral data into a subspace with information from all modes comprehensively considered. Moreover, a proposition related to tensor rank is proofed for interpreting the meaning of the tLTSL model. Extensive experiments are conducted on two different kinds of noise (i.e., simulated and real noise) corrupted HSI data, which validate the effectiveness of tLTSL.
Yangjun Deng, Heng-Chao Li 0001, Siqiao Tan, Junhui Hou, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2023 EGRC-Net: Embedding-Induced Graph Refinement Clustering Network
abstract
Existing graph clustering networks heavily rely on a predefined yet fixed graph, which can lead to failures when the initial graph fails to accurately capture the data topology structure of the embedding space. In order to address this issue, we propose a novel clustering network called Embedding-Induced Graph Refinement Clustering Network (EGRC-Net), which effectively utilizes the learned embedding to adaptively refine the initial graph and enhance the clustering performance. To begin, we leverage both semantic and topological information by employing a vanilla auto-encoder and a graph convolution network, respectively, to learn a latent feature representation. Subsequently, we utilize the local geometric structure within the feature embedding space to construct an adjacency matrix for the graph. This adjacency matrix is dynamically fused with the initial one using our proposed fusion architecture. To train the network in an unsupervised manner, we minimize the Jeffreys divergence between multiple derived distributions. Additionally, we introduce an improved approximate personalized propagation of neural predictions to replace the standard graph convolution network, enabling EGRC-Net to scale effectively. Through extensive experiments conducted on nine widely-used benchmark datasets, we demonstrate that our proposed methods consistently outperform several state-of-the-art approaches. Notably, EGRC-Net achieves an improvement of more than 11.99% in Adjusted Rand Index (ARI) over the best baseline on the DBLP dataset. Furthermore, our scalable approach exhibits a 10.73% gain in ARI while reducing memory usage by 33.73% and decreasing running time by 19.71%. The code for EGRC-Net will be made publicly available at https://github.com/ZhihaoPENG-CityU/EGRC-Net.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
IEEE Trans. Image Process.4
2023 GQE-Net: A Graph-Based Quality Enhancement Network for Point Cloud Color Attribute
abstract
In recent years, point clouds have become increasingly popular for representing three-dimensional (3D) visual objects and scenes. To efficiently store and transmit point clouds, compression methods have been developed, but they often result in a degradation of quality. To reduce color distortion in point clouds, we propose a graph-based quality enhancement network (GQE-Net) that uses geometry information as an auxiliary input and graph convolution blocks to extract local features efficiently. Specifically, we use a parallel-serial graph attention module with a multi-head graph attention mechanism to focus on important points or features and help them fuse together. Additionally, we design a feature refinement module that takes into account the normals and geometry distance between points. To work within the limitations of GPU memory capacity, the distorted point cloud is divided into overlap-allowed 3D patches, which are sent to GQE-Net for quality enhancement. To account for differences in data distribution among different color components, three models are trained for the three color components. Experimental results show that our method achieves state-of-the-art performance. For example, when implementing GQE-Net on a recent test model of the geometry-based point cloud compression (G-PCC) standard, 0.43 dB, 0.25 dB and 0.36 dB Bjφntegaard delta (BD)-peak-signal-to-noise ratio (PSNR), corresponding to 14.0%, 9.3% and 14.5% BD-rate savings were achieved on dense point clouds for the Y, Cb, and Cr components, respectively. The source code of our method is available at https://github.com/xjr998/GQE-Net.
Jinrui Xing, Hui Yuan 0001, Raouf Hamzaoui, Hao Liu 0044, Junhui Hou
IEEE Trans. Image Process.5
2023 Learning a Locally Unified 3D Point Cloud for View Synthesis
abstract
In this paper, we explore the problem of 3D point cloud representation-based view synthesis from a set of sparse source views. To tackle this challenging problem, we propose a new deep learning-based view synthesis paradigm that learns a locally unified 3D point cloud from source views. Specifically, we first construct sub-point clouds by projecting source views to 3D space based on their depth maps. Then, we learn the locally unified 3D point cloud by adaptively fusing points at a local neighborhood defined on the union of the sub-point clouds. Besides, we also propose a 3D geometry-guided image restoration module to fill the holes and recover high-frequency details of the rendered novel views. Experimental results on three benchmark datasets demonstrate that our method can improve the average PSNR by more than 4 dB while preserving more accurate visual details, compared with state-of-the-art view synthesis methods. The code will be publicly available at https://github.com/mengyou2/PCVS.
Meng You, Mantang Guo, Xianqiang Lyu, Hui Liu 0032, Junhui Hou
IEEE Trans. Image Process.5
2023 A Two-Level Rectification Attention Network for Scene Text Recognition
abstract
Scene text recognition is a challenging task in the computer vision field due to the diversity of text styles and the complexity of the image backgrounds. In recent decades, numerous text rectification and recognition methods have been proposed to solve these problems. However, most of these methods rectify texts at the geometry level or pixel level. The former is limited by geometric constraints, and the latter is prone to blurring the text. In this paper, we propose a two-level rectification attention network (TRAN) to rectify and recognize texts. This network consists of two parts: a two-level rectification network (TORN) and an attention-based recognition network (ABRN). Specifically, the TORN first rectifies texts at the geometry level and then performs a pixel-level adjustment, which not only eliminates the geometric constraints but also renders clear texts. The ABRN’s role is to recognize text in the rectified images. To improve the feature extraction ability of our model, we design a new channel-wise and kernel-wise attention unit, which enables the network to handle significant variations of character size and channel interdependencies. Furthermore, we propose a skip training strategy to make our model converge smoothly. We conduct experiments on various benchmarks, including regular and irregular datasets. The experimental results show that our method achieves a state-of-the-art performance.
Lintai Wu, Yong Xu 0001, Junhui Hou, C. L. Philip Chen, Cheng-Lin Liu 0001
IEEE Trans. Multim.3
2023 Light Field Compression With Graph Learning and Dictionary-Guided Sparse Coding
abstract
Light field (LF) data are widely used in the immersive representations of the 3D world. To record the light rays along with different directions, an LF requires much larger storage space and transmission bandwidth than a conventional 2D image with similar spatial dimension. In this paper, we propose a novel framework for light field image compression that leverages graph learning and dictionary learning to remove structural redundancies between different views. Specifically, to significantly reduce the bit-rates, only a few key views are sampled and encoded, whereas the remaining non-key views are reconstructed via the graph adjacency matrix learned from the angular patch. Furthermore, dictionary-guided sparse coding is developed to compress the graph adjacency matrices and reduce the coding overheads. To our best knowledge, this paper is the first to achieve compact representation of cross-view structural information via adaptive learning on graphs. Experimental results demonstrate that the proposed framework achieves better performance than the standardized HEVC-based codec.
Wenrui Dai, Yong Li 0033, Junhui Hou, Junni Zou, Hongkai Xiong
IEEE Trans. Multim.5
2023 Exploiting Manifold Feature Representation for Efficient Classification of 3D Point Clouds
abstract
In this paper, we propose an efficient point cloud classification method via manifold learning based feature representation. Different from conventional methods, we use manifold learning algorithms to embed point cloud features for better considering the geometric continuity on the surface. Then, the nature of point cloud can be acquired in low dimensional space, and after being concatenated with features in the original three-dimensional (3D) space, both the capability of feature representation and the classification network performance can be improved. We explore three traditional manifold algorithms (i.e., Isomap, Locally-Linear Embedding, and Laplacian eigenmaps) in detail, and finally, we select the Locally-Linear Embedding (LLE) algorithm due to its low complexity and locality consistency preservation. Furthermore, we propose a neural network based manifold learning (NNML) method to implement manifold learning based non-linear projection. Experiments demonstrate that the proposed two manifold learning methods can obtain better performances than the state-of-the-art methods, and the obtained mean class accuracy (mA) and overall accuracy (oA) can reach 91.4% and 94.4%, respectively. Moreover, because of the improved feature learning capability, the proposed NNML method can also have better classification accuracy on models with prominent geometric shapes. To further demonstrate the advantages of PointManifold, we extend it as a plug and play method for point cloud classification task, which can be directly used with existing methods and gain a significant improvement.
Dinghao Yang, Wei Gao 0003, Ge Li 0002, Hui Yuan 0001, Junhui Hou, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.5
2023 PU-Flow: A Point Cloud Upsampling Network With Normalizing Flows
abstract
Point cloud upsampling aims to generate dense point clouds from given sparse ones, which is a challenging task due to the irregular and unordered nature of point sets. To address this issue, we present a novel deep learning-based model, called PU-Flow, which incorporates normalizing flows and weight prediction techniques to produce dense points uniformly distributed on the underlying surface. Specifically, we exploit the invertible characteristics of normalizing flows to transform points between euclidean and latent spaces and formulate the upsampling process as ensemble of neighbouring points in a latent space, where the ensemble weights are adaptively learned from local geometric context. Extensive experiments show that our method is competitive and, in most test cases, it outperforms state-of-the-art methods in terms of reconstruction quality, proximity-to-surface accuracy, and computation efficiency. The source code will be publicly available at https://github.com/unknownue/puflow.
Aihua Mao, Zihui Du, Junhui Hou, Yaqi Duan, Yong-Jin Liu 0001, Ying He 0001
IEEE Trans. Vis. Comput. Graph.3
2022 WarpingGAN: Warping Multiple Uniform Priors for Adversarial 3D Point Cloud Generation
abstract
We propose WarpingGAN, an effective and efficient 3D point cloud generation network. Unlike existing methods that generate point clouds by directly learning the mapping functions between latent codes and 3D shapes, Warping-GAN learns a unified local-warping function to warp multiple identical pre-defined priors (i.e., sets of points uniformly distributed on regular 3D grids) into 3D shapes driven by local structure-aware semantics. In addition, we also in-geniously utilize the principle of the discriminator and tai-lor a stitching loss to eliminate the gaps between different partitions of a generated shape corresponding to different priors for boosting quality. Owing to the novel gen-erating mechanism, WarpingGAN, a single lightweight network after one-time training, is capable of efficiently gen-erating uniformly distributed 3D point clouds with various resolutions. Extensive experimental results demonstrate the superiority of our WarpingGAN over state-of-the-art methods in terms of quantitative metrics, visual quality, and efficiency. The source code is publicly available at https://github.com/yztang4/WarpingGAN.git.
Yingzhi Tang, Qijian Zhang, Yiming Zeng 0002, Junhui Hou, Xuefei Zhe
CVPR5
2022 IDEA-Net: Dynamic 3D Point Cloud Interpolation via Deep Embedding Alignment
abstract
This paper investigates the problem of temporally interpolating dynamic 3D point clouds with large non-rigid deformation. We formulate the problem as estimation of point-wise trajectories (i.e., smooth curves) and further reason that temporal irregularity and under-sampling are two major challenges. To tackle the challenges, we propose IDEA-Net, an end-to-end deep learning framework, which disentangles the problem under the assistance of the explicitly learned temporal consistency. Specifically, we propose a temporal consistency learning module to align two consecutive point cloud frames point-wisely, based on which we can employ linear interpolation to obtain coarse trajectories/in-between frames. To compensate the high-order nonlinear components of trajectories, we apply aligned feature embeddings that encode local geometry properties to regress point-wise increments, which are combined with the coarse estimations. We demonstrate the effectiveness of our method on various point cloud sequences and observe large improvement over state-of-the-art methods both quantitatively and visually. Our framework can bring benefits to 3D motion data acquisition. The source code is publicly available at https://github.com/ZENGYIMING-EAMON/IDEANet.git.
Yiming Zeng 0002, Qijian Zhang, Junhui Hou, Yixuan Yuan, Ying He 0001
CVPR4
2022 AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular Data
abstract
Dynamic Vision Sensor (DVS) is a compelling neuromorphic camera compared to conventional camera, but it suffers from fiercer noise. Due to the nature of irregular format and asynchronous readout, DVS data is always transformed into a regular tensor (e.g., 3D voxel or image) for deep learning method, which corrupts its own asynchronous properties. To maintain asynchronous, we establish an innovative asynchronous event denoise neural network, named AEDNet, which directly consumes the correlation of the irregular signal in spatial-temporal range without destroying its original structural property. Based on the property of continuation in temporal domain and discreteness in spatial domain, we decompose the DVS signal into two parts, i.e., temporal correlation and spatial affinity, and separately process these two parts. Our spatial feature embedding unit is a unique feature extraction module that extracts feature from event-level, which perfectly maintains its spatial-temporal correlation. To test effectiveness, we build a novel dataset named DVSCLEAN containing both simulated and real-world data. The experimental results of AEDNet achieve SOTA.
Huachen Fang, Jinjian Wu, Leida Li, Junhui Hou, Weisheng Dong, Guangming Shi
ACM Multimedia4
2022 Learning Graph-embedded Key-event Back-tracing for Object Tracking in Event Clouds
abstract
Event data-based object tracking is attracting attention increasingly. Unfortunately, the unusual data structure caused by the unique sensing mechanism poses great challenges in designing downstream algorithms. To tackle such challenges, existing methods usually re-organize raw event data (or event clouds) with the event frame/image representation to adapt to mature RGB data-based tracking paradigms, which compromises the high temporal resolution and sparse characteristics. By contrast, we advocate developing new designs/techniques tailored to the special data structure to realize object tracking. To this end, we make the first attempt to construct a new end-to-end learning-based paradigm that directly consumes event clouds. Specifically, to process a non-uniformly distributed large-scale event cloud efficiently, we propose a simple yet effective density-insensitive downsampling strategy to sample a subset called key-events. Then, we employ a graph-based network to embed the irregular spatio-temporal information of key-events into a high-dimensional feature space, and the resulting embeddings are utilized to predict their target likelihoods via semantic-driven Siamese-matching. Besides, we also propose motion-aware target likelihood prediction, which learns the motion flow to back-trace the potential initial positions of key-events and measures them with the previous proposal. Finally, we obtain the bounding box by adaptively fusing the two intermediate ones separately regressed from the weighted embeddings of key-events by the two types of predicted target likelihoods. Extensive experiments on both synthetic and real event datasets demonstrate the superiority of the proposed framework over state-of-the-art methods in terms of both the tracking accuracy and speed. The code is publicly available at https://github.com/ZHU-Zhiyu/Event-tracking.
Junhui Hou, Xianqiang Lyu
NeurIPS2
2022 Learning hyperspectral images from RGB images via a coarse-to-fine CNN
Shaohui Mei, Yunhao Geng, Junhui Hou, Qian Du 0001
Sci. China Inf. Sci.3
2022 RegGeoNet: Learning Regular Representations for Large-Scale 3D Point Clouds
Qijian Zhang, Junhui Hou, Antoni B. Chan, Juyong Zhang, Ying He 0001
Int. J. Comput. Vis.2
2022 Deep Spatial-Angular Regularization for Light Field Imaging, Denoising, and Super-Resolution
abstract
Coded aperture is a promising approach for capturing the 4-D light field (LF), in which the 4-D data are compressively modulated into 2-D coded measurements that are further decoded by reconstruction algorithms. The bottleneck lies in the reconstruction algorithms, resulting in rather limited reconstruction quality. To tackle this challenge, we propose a novel learning-based framework for the reconstruction of high-quality LFs from acquisitions via learned coded apertures. The proposed method incorporates the measurement observation into the deep learning framework elegantly to avoid relying entirely on data-driven priors for LF reconstruction. Specifically, we first formulate the compressive LF reconstruction as an inverse problem with an implicit regularization term. Then, we construct the regularization term with a deep efficient spatial-angular separable convolutional sub-network in the form of local and global residual learning to comprehensively explore the signal distribution free from the limited representation ability and inefficiency of deterministic mathematical modeling. Furthermore, we extend this pipeline to LF denoising and spatial super-resolution, which could be considered as variants of coded aperture imaging equipped with different degradation matrices. Extensive experimental results demonstrate that the proposed methods outperform state-of-the-art approaches to a significant extent both quantitatively and qualitatively, i.e., the reconstructed LFs not only achieve much higher PSNR/SSIM but also preserve the LF parallax structure better on both real and synthetic LF benchmarks. The code will be publicly available at https://github.com/MantangGuo/DRLF.
Mantang Guo, Junhui Hou, Jing Jin 0006, Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Deep Coarse-to-Fine Dense Light Field Reconstruction With Flexible Sampling and Geometry-Aware Fusion
abstract
A densely-sampled light field (LF) is highly desirable in various applications, such as 3-D reconstruction, post-capture refocusing and virtual reality. However, it is costly to acquire such data. Although many computational methods have been proposed to reconstruct a densely-sampled LF from a sparsely-sampled one, they still suffer from either low reconstruction quality, low computational efficiency, or the restriction on the regularity of the sampling pattern. To this end, we propose a novel learning-based method, which accepts sparsely-sampled LFs with irregular structures, and produces densely-sampled LFs with arbitrary angular resolution accurately and efficiently. We also propose a simple yet effective method for optimizing the sampling pattern. Our proposed method, an end-to-end trainable network, reconstructs a densely-sampled LF in a coarse-to-fine manner. Specifically, the coarse sub-aperture image (SAI) synthesis module first explores the scene geometry from an unstructured sparsely-sampled LF and leverages it to independently synthesize novel SAIs, in which a confidence-based blending strategy is proposed to fuse the information from different input SAIs, giving an intermediate densely-sampled LF. Then, the efficient LF refinement module learns the angular relationship within the intermediate result to recover the LF parallax structure. Comprehensive experimental evaluations demonstrate the superiority of our method on both real-world and synthetic LF images when compared with state-of-the-art methods. In addition, we illustrate the benefits and advantages of the proposed approach when applied in various LF-based applications, including image-based rendering and depth estimation enhancement. The code is available at https://github.com/jingjin25/LFASR-FS-GAF.
Jing Jin 0006, Junhui Hou, Jie Chen 0026, Huanqiang Zeng, Sam Kwong, Jingyi Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Point Cloud Quality Assessment via 3D Edge Similarity Measurement
abstract
In this letter, a new full-reference metric is presented to assess the perceptual quality of the point clouds (PCs). The human visual system (HVS) always shows a high sensitivity to the three-dimensional (3D) edge features inherent in the PCs. With this motivation, the three-dimensional edge similarity-based model (TDESM) is proposed, which makes the first attempt to apply 3D Difference of Gaussian (3D-DOG) on point cloud quality assessment (PCQA). Specifically, the 3D edge features are captured by convolving the dual-scale 3D-DOG filters with both reference and distorted PCs. The quality scores of distorted PCs are generated by combining the 3D edge similarity measured from different scales. The experiments are conducted on four publicly available PCQA datasets, i.e., Torlig2018, M-PCCD, ICIP2020, and SJTU-PCQA. Compared with multiple state-of-the-art PCQA metrics, our proposed approach is able to be higher consistent with the subjective perception on the PCs.
Zian Lu, Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Kai-Kuang Ma
IEEE Signal Process. Lett.4
2022 Self-Supervised Symmetric Nonnegative Matrix Factorization
abstract
Symmetric nonnegative matrix factorization (SNMF) has demonstrated to be a powerful method for data clustering. However, SNMF is mathematically formulated as a non-convex optimization problem, making it sensitive to the initialization of variables. Inspired by ensemble clustering that aims to seek a better clustering result from a set of clustering results, we propose self-supervised SNMF (S3NMF), which is capable of boosting clustering performance progressively by taking advantage of the sensitivity to initialization characteristic of SNMF, without relying on any additional information. Specifically, we first perform SNMF repeatedly with a random positive matrix for initialization each time, leading to multiple decomposed matrices. Then, we rank the quality of the resulting matrices with adaptively learned weights, from which a new similarity matrix that is expected to be more discriminative is reconstructed for SNMF again. These two steps are iterated until the stopping criterion/maximum number of iterations is achieved. We mathematically formulate S3NMF as a constrained optimization problem, and provide an alternative optimization algorithm to solve it with the theoretical convergence guaranteed. Extensive experimental results on 10 commonly used benchmark datasets demonstrate the significant advantage of our S3NMF over 14 state-of-the-art methods in terms of 5 quantitative metrics. The source code is publicly available athttps://github.com/jyh-learning/SSSNMF.
Yuheng Jia, Hui Liu 0032, Junhui Hou, Sam Kwong, Qingfu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Finding Stars From Fireworks: Improving Non-Cooperative Iris Tracking
abstract
We revisit the problem of iris tracking with RGB cameras, aiming to obtain iris contours from captured images of eyes. We find the reason that limits the performance of the state-of-the-art method in more general non-cooperative environments, which prohibits a wider adoption of this useful technique in practice. We believe that because the iris boundary could be inherently unclear and blocked, as its pixels occupy only an extremely limited percentage of those on the entire image of the eye, similar to the stars hidden in fireworks, we should not treat the boundary pixels as one class to conduct end-to-end recognition directly. Thus, we propose to learn features from iris and sclera regions first, and then leverage entropy to sketch the thin and sharp iris boundary pixels, where we can trace more precise parameterized iris contours. In this work, we also collect a new dataset by smartphone with 22 K images of eyes from video clips. We annotate a subset of 2 K images, so that label propagation can be applied to further enhance the system performance. Extensive experiments over both public and our own datasets show that our method outperforms the state-of-the-art method. The results also indicate that our method can improve the coarsely labeled data to enhance the iris contour’s accuracy and support the downstream application better than the prior method.
Chengdong Lin, Zhenjiang Li 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.4
2022 Global-Local Balanced Low-Rank Approximation of Hyperspectral Images for Classification
abstract
This paper explores the problem of recovering the discriminative representation of a hyperspectral remote sensing image (HRSI), which suffers from spectral variations, to boost its classification accuracy. To tackle this challenge, we propose a new method, namely local-global balanced low-rank approximation (GLB-LRA), which can increase the similarity between pixels belonging to an identical category while promoting the discriminability between pixels of different categories. Specifically, by taking advantage of the particular structural spatial information of HRSIs, we exploit the low-rankness of an HRSI robustly in both spatial and spectral domains from the perspective of local and global balance. We mathematically formulate GLB-LRA as an explicit optimization problem and propose an iterative algorithm to solve it efficiently. Experimental results over three commonly-used benchmark datasets demonstrate the significant superiority of our method over state-of-the-art methods.
Hui Liu 0032, Yuheng Jia, Junhui Hou, Qingfu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Learning Low-Rank Graph With Enhanced Supervision
abstract
In this paper, we propose a new semi-supervised graph construction method, which is capable of adaptively learning the similarity relationship between data samples by fully exploiting the potential of pairwise constraints, a kind of weakly supervisory information. Specifically, to adaptively learn the similarity relationship, we linearly approximate each sample with others under the regularization of the low-rankness of the matrix formed by the approximation coefficient vectors of all the samples. In the meanwhile, by taking advantage of the underlying local geometric structure of data samples that is empirically obtained, we enhance the dissimilarity information of the available pairwise constraints via propagation. We seamlessly combine the two adversarial learning processes to achieve mutual guidance. We cast our method as a constrained optimization problem and provide an efficient alternating iterative algorithm to solve it. Experimental results on five commonly-used benchmark datasets demonstrate that our method produces much higher classification accuracy than state-of-the-art methods, while running faster.
Hui Liu 0032, Yuheng Jia, Junhui Hou, Qingfu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 A Hybrid Compression Framework for Color Attributes of Static 3D Point Clouds
abstract
The emergence of 3D point clouds (3DPCs) is promoting the rapid development of immersive communication, autonomous driving, and so on. Due to the huge data volume, the compression of 3DPCs is becoming more and more attractive. We propose a novel and efficient color attribute compression method for static 3DPCs. First, a 3DPC is partitioned into several sub-point clouds by color distribution analysis. Each sub-point cloud is then decomposed into a lot of 3D blocks by an improved k-d tree-based decomposition algorithm. Afterwards, a novel virtual adaptive sampling-based sparse representation strategy is proposed for each 3D block to remove the redundancy among points, in which the bases of the graph transform (GT) and the discrete cosine transform (DCT) are used as candidates of the complete dictionary. Experimental results over 10 common 3DPCs demonstrate that the proposed method can achieve superior or comparable coding performance when compared with the current state-of-the-art methods.
Hao Liu 0044, Hui Yuan 0001, Qi Liu 0029, Junhui Hou, Huanqiang Zeng, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2022 Maximum Entropy Subspace Clustering Network
abstract
Deep subspace clustering networks have attracted much attention in subspace clustering, in which an auto-encoder non-linearly maps the input data into a latent space, and a fully connected layer named self-expressiveness module is introduced to learn the affinity matrix via a typical regularization term (e.g., sparse or low-rank). However, the adopted regularization terms ignore the connectivity within each subspace, limiting their clustering performance. In addition, the adopted framework suffers from the coupling issue between the auto-encoder module and the self-expressiveness module, making the network training non-trivial. To tackle these two issues, we propose a novel deep subspace clustering method named Maximum Entropy Subspace Clustering Network (MESC-Net). Specifically, MESC-Net maximizes the entropy of the affinity matrix to promote the connectivity within each subspace, in which its elements corresponding to the same subspace are uniformly and densely distributed. Meanwhile, we design a novel framework to explicitly decouple the auto-encoder module and the self-expressiveness module. Besides, we also theoretically prove that the learned affinity matrix satisfies the block-diagonal property under the assumption of independent subspaces. Extensive quantitative and qualitative results on commonly used benchmark datasets validate MESC-Net significantly outperforms state-of-the-art methods. The code is publicly available athttps://github.com/ZhihaoPENG-CityU/MESC.
Zhihao Peng 0002, Yuheng Jia, Hui Liu 0032, Junhui Hou, Qingfu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Semisupervised Affinity Matrix Learning via Dual-Channel Information Recovery
abstract
This article explores the problem of semisupervised affinity matrix learning, that is, learning an affinity matrix of data samples under the supervision of a small number of pairwise constraints (PCs). By observing that both the matrix encoding PCs, called pairwise constraint matrix (PCM) and the empirically constructed affinity matrix (EAM), express the similarity between samples, we assume that both of them are generated from a latent affinity matrix (LAM) that can depict the ideal pairwise relation between samples. Specifically, the PCM can be thought of as a partial observation of the LAM, while the EAM is a fully observed one but corrupted with noise/outliers. To this end, we innovatively cast the semisupervised affinity matrix learning as the recovery of the LAM guided by the PCM and EAM, which is technically formulated as a convex optimization problem. We also provide an efficient algorithm for solving the resulting model numerically. Extensive experiments on benchmark datasets demonstrate the significant superiority of our method over state-of-the-art ones when used for constrained clustering and dimensionality reduction. The code is publicly available at https://github.com/jyh-learning/LAM.
Yuheng Jia, Hui Liu 0032, Junhui Hou, Sam Kwong, Qingfu Zhang 0001
IEEE Trans. Cybern.3
2022 Attention-Guided Progressive Neural Texture Fusion for High Dynamic Range Image Restoration
abstract
High Dynamic Range (HDR) imaging via multi-exposure fusion is an important task for most modern imaging platforms. In spite of recent developments in both hardware and algorithm innovations, challenges remain over content association ambiguities caused by saturation, motion, and various artifacts introduced during multi-exposure fusion such as ghosting, noise, and blur. In this work, we propose an Attention-guided Progressive Neural Texture Fusion (APNT-Fusion) HDR restoration model which aims to address these issues within one framework. An efficient two-stream structure is proposed which separately focuses on texture feature transfer over saturated regions and multi-exposure tonal and texture feature fusion. A neural feature transfer mechanism is proposed which establishes spatial correspondence between different exposures based on multi-scale VGG features in the masked saturated HDR domain for discriminative contextual clues over the ambiguous image areas. A progressive texture blending module is designed to blend the encoded two-stream features in a multi-scale and progressive manner. In addition, we introduce several novel attention mechanisms, i.e., the motion attention module detects and suppresses the content discrepancies among the reference images; the saturation attention module facilitates differentiating the misalignment caused by saturation from those caused by motion; and the scale attention module ensures texture blending consistency between different coder/decoder scales. We carry out comprehensive qualitative and quantitative evaluations and ablation studies, which validate that these novel modules work coherently under the same framework and outperform state-of-the-art methods.
Jie Chen 0026, Zaifeng Yang, Tsz Nam Chan, Hui Li 0029, Junhui Hou, Lap-Pui Chau
IEEE Trans. Image Process.5
2022 Deep Posterior Distribution-Based Embedding for Hyperspectral Image Super-Resolution
abstract
In this paper, we investigate the problem of hyperspectral (HS) image spatial super-resolution via deep learning. Particularly, we focus on how to embed the high-dimensional spatial-spectral information of HS images efficiently and effectively. Specifically, in contrast to existing methods adopting empirically-designed network modules, we formulate HS embedding as an approximation of the posterior distribution of a set of carefully-defined HS embedding events, including layer-wise spatial-spectral feature extraction and network-level feature aggregation. Then, we incorporate the proposed feature embedding scheme into a source-consistent super-resolution framework that is physically-interpretable, producing PDE-Net, in which high-resolution (HR) HS images are iteratively refined from the residuals between input low-resolution (LR) HS images and pseudo-LR-HS images degenerated from reconstructed HR-HS images via probability-inspired HS embedding. Extensive experiments over three common benchmark datasets demonstrate that PDE-Net achieves superior performance over state-of-the-art methods. Besides, the probabilistic characteristic of this kind of networks can provide the epistemic uncertainty of the network outputs, which may bring additional benefits when used for other HS image-based applications. The code will be publicly available at https://github.com/jinnh/PDE-Net.
Jinhui Hou, Junhui Hou, Huanqiang Zeng, Jinjian Wu, Jiantao Zhou 0001
IEEE Trans. Image Process.3
2022 A Spatial and Geometry Feature-Based Quality Assessment Model for the Light Field Images
abstract
This paper proposes a new full-reference image quality assessment (IQA) model for performing perceptual quality evaluation on light field (LF) images, called the spatial and geometry feature-based model (SGFM). Considering that the LF image describe both spatial and geometry information of the scene, the spatial features are extracted over the sub-aperture images (SAIs) by using contourlet transform and then exploited to reflect the spatial quality degradation of the LF images, while the geometry features are extracted across the adjacent SAIs based on 3D-Gabor filter and then explored to describe the viewing consistency loss of the LF images. These schemes are motivated and designed based on the fact that the human eyes are more interested in the scale, direction, contour from the spatial perspective and viewing angle variations from the geometry perspective. These operations are applied to the reference and distorted LF images independently. The degree of similarity can be computed based on the above-measured quantities for jointly arriving at the final IQA score of the distorted LF image. Experimental results on three commonly-used LF IQA datasets show that the proposed SGFM is more in line with the quality assessment of the LF images perceived by the human visual system (HVS), compared with multiple classical and state-of-the-art IQA models.
Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Image Process.3
2022 Occlusion-Aware Unsupervised Learning of Depth From 4-D Light Fields
abstract
Depth estimation is a fundamental issue in 4-D light field processing and analysis. Although recent supervised learning-based light field depth estimation methods have significantly improved the accuracy and efficiency of traditional optimization-based ones, these methods rely on the training over light field data with ground-truth depth maps which are challenging to obtain or even unavailable for real-world light field data. Besides, due to the inevitable gap (or domain difference) between real-world and synthetic data, they may suffer from serious performance degradation when generalizing the models trained with synthetic data to real-world data. By contrast, we propose an unsupervised learning-based method, which does not require ground-truth depth as supervision during training. Specifically, based on the basic knowledge of the unique geometry structure of light field data, we present an occlusion-aware strategy to improve the accuracy on occlusion areas, in which we explore the angular coherence among subsets of the light field views to estimate initial depth maps, and utilize a constrained unsupervised loss to learn their corresponding reliability for final depth prediction. Additionally, we adopt a multi-scale network with a weighted smoothness loss to handle the textureless areas. Experimental results on synthetic data show that our method can significantly shrink the performance gap between the previous unsupervised method and supervised ones, and produce depth maps with comparable accuracy to traditional methods with obviously reduced computational cost. Moreover, experiments on real-world datasets show that our method can avoid the domain shift problem presented in supervised methods, demonstrating the great potential of our method. The code will be publicly available at https://github.com/jingjin25/LFDE-OccUnNet.
Jing Jin 0006, Junhui Hou
IEEE Trans. Image Process.2
2022 PUFA-GAN: A Frequency-Aware Generative Adversarial Network for 3D Point Cloud Upsampling
abstract
We propose a generative adversarial network for point cloud upsampling, which can not only make the upsampled points evenly distributed on the underlying surface but also efficiently generate clean high frequency regions. The generator of our network includes a dynamic graph hierarchical residual aggregation unit and a hierarchical residual aggregation unit for point feature extraction and upsampling, respectively. The former extracts multiscale point-wise descriptive features, while the latter captures rich feature details with hierarchical residuals. To generate neat edges, our discriminator uses a graph filter to extract and retain high frequency points. The generated high resolution point cloud and corresponding high frequency points help the discriminator learn the global and high frequency properties of the point cloud. We also propose an identity distribution loss function to make sure that the upsampled points remain on the underlying surface of the input low resolution point cloud. To assess the regularity of the upsampled points in high frequency regions, we introduce two evaluation metrics. Objective and subjective results demonstrate that the visual quality of the upsampled point clouds generated by our method is better than that of the state-of-the-art methods.
Hao Liu 0044, Hui Yuan 0001, Junhui Hou, Raouf Hamzaoui, Wei Gao 0003
IEEE Trans. Image Process.3
2022 Adaptive Attribute and Structure Subspace Clustering Network
abstract
Deep self-expressiveness-based subspace clustering methods have demonstrated effectiveness. However, existing works only consider the attribute information to conduct the self-expressiveness, limiting the clustering performance. In this paper, we propose a novel adaptive attribute and structure subspace clustering network (AASSC-Net) to simultaneously consider the attribute and structure information in an adaptive graph fusion manner. Specifically, we first exploit an auto-encoder to represent input data samples with latent features for the construction of an attribute matrix. We also construct a mixed signed and symmetric structure matrix to capture the local geometric structure underlying data samples. Then, we perform self-expressiveness on the constructed attribute and structure matrices to learn their affinity graphs separately. Finally, we design a novel attention-based fusion module to adaptively leverage these two affinity graphs to construct a more discriminative affinity graph. Extensive experimental results on commonly used benchmark datasets demonstrate that our AASSC-Net significantly outperforms state-of-the-art methods. In addition, we conduct comprehensive ablation studies to discuss the effectiveness of the designed modules. The code is publicly available at https://github.com/ZhihaoPENG-CityU/AASSC-Net.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
IEEE Trans. Image Process.4
2022 Screen Content Video Quality Assessment Model Using Hybrid Spatiotemporal Features
abstract
In this paper, a full-reference video quality assessment (VQA) model is designed for the perceptual quality assessment of the screen content videos (SCVs), called the hybrid spatiotemporal feature-based model (HSFM). The SCVs are of hybrid structure including screen and natural scenes, which are perceived by the human visual system (HVS) with different visual effects. With this consideration, the three dimensional Laplacian of Gaussian (3D-LOG) filter and three dimensional Natural Scene Statistics (3D-NSS) are exploited to extract the screen and natural spatiotemporal features, based on the reference and distorted SCV sequences separately. The similarities of these extracted features are then computed independently, followed by generating the distorted screen and natural quality scores for screen and natural scenes. After that, an adaptive screen and natural quality fusion scheme through the local video activity is developed to combine them for arriving at the final VQA score of the distorted SCV under evaluation. The experimental results on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases have shown that the proposed HSFM is more in line with the perceptual quality assessment of the SCVs perceived by the HVS, compared with a variety of classic and latest IQA/VQA models.
Huanqiang Zeng, Hailiang Huang 0002, Junhui Hou, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Trans. Image Process.3
2021 Clustering Ensemble Meets Low-rank Tensor Approximation
abstract
This paper explores the problem of clustering ensemble, which aims to combine multiple base clusterings to produce better performance than that of the individual one. The existing clustering ensemble methods generally construct a co-association matrix, which indicates the pairwise similarity between samples, as the weighted linear combination of the connective matrices from different base clusterings, and the resulting co-association matrix is then adopted as the input of an off-the-shelf clustering algorithm, e.g., spectral clustering. However, the co-association matrix may be dominated by poor base clusterings, resulting in inferior performance. In this paper, we propose a novel low-rank tensor approximation based method to solve the problem from a global perspective. Specifically, by inspecting whether two samples are clustered to an identical cluster under different base clusterings, we derive a coherent-link matrix, which contains limited but highly reliable relationships between samples. We then stack the coherent-link matrix and the co-association matrix to form a three-dimensional tensor, the low-rankness property of which is further explored to propagate the information of the coherent-link matrix to the co-association matrix, producing a refined co-association matrix. We formulate the proposed method as a convex constrained optimization problem and solve it efficiently. Experimental results over 7 benchmark data sets show that the proposed model achieves a breakthrough in clustering performance, compared with 12 state-of-the-art methods. To the best of our knowledge, this is the first work to explore the potential of low-rank tensor on clustering ensemble, which is fundamentally different from previous approaches. Last but not least, our method only contains one parameter, which can be easily tuned.
Yuheng Jia, Hui Liu 0032, Junhui Hou, Qingfu Zhang 0001
AAAI3
2021 Recurrent Multi-View Alignment Network for Unsupervised Surface Registration
abstract
Learning non-rigid registration in an end-to-end manner is challenging due to the inherent high degrees of freedom and the lack of labeled training data. In this paper, we resolve these two challenges simultaneously. First, we propose to represent the non-rigid transformation with a point-wise combination of several rigid transformations. This representation not only makes the solution space well-constrained but also enables our method to be solved iteratively with a recurrent framework, which greatly reduces the difficulty of learning. Second, we introduce a differentiable loss function that measures the 3D shape similarity on the projected multi-view 2D depth images so that our full framework can be trained end-to-end without ground truth supervision. Extensive experiments on several different datasets demonstrate that our proposed method outperforms the previous state-of-the-art by a large margin.
Wanquan Feng, Juyong Zhang, Hongrui Cai, Haofei Xu, Junhui Hou, Hujun Bao
CVPR5
2021 CorrNet3D: Unsupervised End-to-End Learning of Dense Correspondence for 3D Point Clouds
abstract
Motivated by the intuition that one can transform two aligned point clouds to each other more easily and meaningfully than a misaligned pair, we propose CorrNet3D – the first unsupervised and end-to-end deep learning-based framework – to drive the learning of dense correspondence between 3D shapes by means of deformation-like reconstruction to overcome the need for annotated data. Specifically, CorrNet3D consists of a deep feature embedding module and two novel modules called correspondence indicator and symmetric deformer. Feeding a pair of raw point clouds, our model first learns the pointwise features and passes them into the indicator to generate a learnable correspondence matrix used to permute the input pair. The symmetric deformer, with an additional regularized loss, transforms the two permuted point clouds to each other to drive the unsupervised learning of the correspondence. The extensive experiments on both synthetic and real-world datasets of rigid and non-rigid 3D shapes show our CorrNet3D outperforms state-of-the-art methods to a large extent, including those taking meshes as input. CorrNet3D is a flexible framework in that it can be easily adapted to supervised learning if annotated data are available. The source code and pre-trained model will be available at https://github.com/ZENGYIMINGEAMON/CorrNet3D.git.
Yiming Zeng 0002, Junhui Hou, Hui Yuan 0001, Ying He 0001
CVPR4
2021 Learning Dynamic Interpolation for Extremely Sparse Light Fields with Wide Baselines
abstract
In this paper, we tackle the problem of dense light field (LF) reconstruction from sparsely-sampled ones with wide baselines and propose a learnable model, namely dynamic interpolation, to replace the commonly-used geometry warping operation. Specifically, with the estimated geometric relation between input views, we first construct a lightweight neural network to dynamically learn weights for interpolating neighbouring pixels from input views to synthesize each pixel of novel views independently. In contrast to the fixed and content-independent weights employed in the geometry warping operation, the learned interpolation weights implicitly incorporate the correspondences between the source and novel views and adapt to different image content information. Then, we recover the spatial correlation between the independently synthesized pixels of each novel view by referring to that of input views using a geometry-based spatial refinement module. We also constrain the angular correlation between the novel views through a disparity-oriented LF structure loss. Experimental results on LF datasets with wide baselines show that the reconstructed LFs achieve much higher PSNR/SSIM and preserve the LF parallax structure better than state-of-the-art methods. The source code is publicly available at https://github.com/MantangGuo/DI4SLF.
Mantang Guo, Jing Jin 0006, Hui Liu 0032, Junhui Hou
ICCV4
2021 Semantic-embedded Unsupervised Spectral Reconstruction from Single RGB Images in the Wild
abstract
This paper investigates the problem of reconstructing hyperspectral (HS) images from single RGB images captured by commercial cameras, without using paired HS and RGB images during training. To tackle this challenge, we propose a new lightweight and end-to-end learning-based framework. Specifically, on the basis of the intrinsic imaging degradation model of RGB images from HS images, we progressively spread the differences between input RGB images and re-projected RGB images from recovered HS images via effective unsupervised camera spectral response function estimation. To enable the learning without paired ground-truth HS images as supervision, we adopt the adversarial learning manner and boost it with a simple yet effective ℒ1gradient clipping scheme. Besides, we embed the semantic information of input RGB images to locally regularize the unsupervised learning, which is expected to promote pixels with identical semantics to have consistent spectral signatures. In addition to conducting quantitative experiments over two widely-used datasets for HS image reconstruction from synthetic RGB images, we also evaluate our method by applying recovered HS images from real RGB images to HS-based visual tracking. Extensive results show that our method significantly outperforms state-of-the-art unsupervised methods and even exceeds the latest supervised method under some settings. The source code is public available at https://github.com/zbzhzhy/Unsupervised-Spectral-Reconstruction.
Hui Liu 0032, Junhui Hou, Huanqiang Zeng, Qingfu Zhang 0001
ICCV3
2021 DRLFNet: A Dense-Connection Residual Learning Neural Network for Light Field Super Resolution
Congrui Fu, Junhui Hou, Hui Yuan 0001
ICIG (3)4
2021 Learning Spatial-angular Fusion for Compressive Light Field Imaging in a Cycle-consistent Framework
abstract
This paper investigates the 4-D light field (LF) reconstruction from 2-D measurements captured by the coded aperture camera. To tackle such an ill-posed inverse problem, we propose a cycle-consistent reconstruction network (CR-Net). To be specific, based on the intrinsic linear imaging model of the coded aperture, CR-Net reconstructs an LF through progressively eliminating the residuals between the projected measurements from the reconstructed LF and input measurements. Moreover, to address the crucial issue of extracting representative features from high-dimensional LF data efficiently and effectively, we formulate the problem in a probability space and propose to approximate a posterior distribution of a set of carefully-defined LF processing events, including both layer-wise spatial-angular feature extraction and network-level feature aggregation. Through droppath from a densely-connected template network, we derive an adaptively learned spatial-angular fusion strategy, which is sharply contrasted with existing manners that combine spatial and angular features empirically. Extensive experiments on both simulated measurements and measurements by a real coded aperture camera demonstrate the significant advantage of our method over state-of-the-art ones, i.e., our method improves the reconstruction quality by 4.5 dB.
Xianqiang Lyu, Mantang Guo, Jing Jin 0006, Junhui Hou, Huanqiang Zeng
ACM Multimedia5
2021 Attention-driven Graph Clustering Network
abstract
The combination of the traditional convolutional network (i.e., an auto-encoder) and the graph convolutional network has attracted much attention in clustering, in which the auto-encoder extracts the node attribute feature and the graph convolutional network captures the topological graph feature. However, the existing works (i) lack a flexible combination mechanism to adaptively fuse those two kinds of features for learning the discriminative representation and (ii) overlook the multi-scale information embedded at different layers for subsequent cluster assignment, leading to inferior clustering results. To this end, we propose a novel deep clustering method named Attention-driven Graph Clustering Network (AGCN). Specifically, AGCN exploits a heterogeneity-wise fusion module to dynamically fuse the node attribute feature and the topological graph feature. Moreover, AGCN develops a scale-wise fusion module to adaptively aggregate the multi-scale features embedded at different layers. Based on a unified optimization framework, AGCN can jointly perform feature learning and cluster assignment in an unsupervised fashion. Compared with the existing deep clustering methods, our method is more flexible and effective since it comprehensively considers the numerous and discriminative information embedded in the network and directly produces the clustering results. Extensive quantitative and qualitative results on commonly used benchmark datasets validate that our AGCN consistently outperforms state-of-the-art methods.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
ACM Multimedia4
2021 Categorical Matrix Completion With Active Learning for High-Throughput Screening
abstract
The recent advances in wet-lab automation enable high-throughput experiments to be conducted seamlessly. In particular, the exhaustive enumeration of all possible conditions is always involved in high-throughput screening. Nonetheless, such a screening strategy is hardly believed to be optimal and cost-effective. By incorporating artificial intelligence, we design an open-source model based on categorical matrix completion and active machine learning to guide high throughput screening experiments. Specifically, we narrow our scope to the high-throughput screening for chemical compound effects on diverse protein sub-cellular locations. In the proposed model, we believe that exploration is more important than the exploitation in the long-run of high-throughput screening experiment, Therefore, we design several innovations to circumvent the existing limitations. In particular, categorical matrix completion is designed to accurately impute the missing experiments while margin sampling is also implemented for uncertainty estimation. The model is systematically tested on both simulated and real data. The simulation results reflect that our model can be robust to diverse scenarios, while the real data results demonstrate the wet-lab applicability of our model for high-throughput screening experiments. Lastly, we attribute the model success to its exploration ability by revealing the related matrix ranks and distinct experiment coverage comparisons.
Junhui Hou, Ka-Chun Wong
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Multi-View Spectral Clustering Tailored Tensor Low-Rank Representation
abstract
This paper explores the problem of multi-view spectral clustering (MVSC) based on tensor low-rank modeling. Unlike the existing methods that all adopt an off-the-shelf tensor low-rank norm without considering the special characteristics of the tensor in MVSC, we design a novel structured tensor low-rank norm tailored to MVSC. Specifically, we explicitly impose a symmetric low-rank constraint and a structured sparse low-rank constraint on the frontal and horizontal slices of the tensor to characterize the intra-view and inter-view relationships, respectively. Moreover, the two constraints could be jointly optimized to achieve mutual refinement. On basis of the novel tensor low-rank norm, we formulate MVSC as a convex low-rank tensor recovery problem, which is then efficiently solved with an augmented Lagrange multiplier-based method iteratively. Extensive experimental results on seven commonly used benchmark datasets show that the proposed method outperforms state-of-the-art methods to a significant extent. Impressively, our method is able to produce perfect clustering. In addition, the parameters of our method can be easily tuned, and the proposed model is robust to different datasets, demonstrating its potential in practice. The code is available athttps://github.com/jyh-learning/MVSC-TLRR.
Yuheng Jia, Hui Liu 0032, Junhui Hou, Sam Kwong, Qingfu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 PQA-Net: Deep No Reference Point Cloud Quality Assessment via Multi-View Projection
abstract
Recently, 3D point cloud is becoming popular due to its capability to represent the real world for advanced content modality in modern communication systems. In view of its wide applications, especially for immersive communication towards human perception, quality metrics for point clouds are essential. Existing point cloud quality evaluations rely on a full or certain portion of the original point cloud, which severely limits their applications. To overcome this problem, we propose a novel deep learning-based no reference point cloud quality assessment method, namely PQA-Net. Specifically, the PQA-Net consists of a multi-view-based joint feature extraction and fusion (MVFEF) module, a distortion type identification (DTI) module, and a quality vector prediction (QVP) module. The DTI and QVP modules share the feature generated from the MVFEF module. By using the distortion type labels, the DTI and the MVFEF modules are first pre-trained to initialize the network parameters, based on which the whole network is then jointly trained to finally evaluate the point cloud quality. Experimental results on the Waterloo Point Cloud dataset show that PQA-Net achieves better or equivalent performance comparing with the state-of-the-art quality assessment methods. The code of the proposed model will be made publicly available to facilitate reproducible researchhttps://github.com/qdushl/PQA-Net.
Qi Liu 0029, Hui Yuan 0001, Honglei Su, Hao Liu 0044, Yu Wang 0106, Huan Yang 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.7
2021 A Light Field Image Quality Assessment Model Based on Symmetry and Depth Features
abstract
This paper presents a new full-reference image quality assessment (IQA) method for conducting the perceptual quality evaluation of the light field (LF) images, called the symmetry and depth feature-based model (SDFM). Specifically, the radial symmetry transform is first employed on the luminance components of the reference and distorted LF images to extract their symmetry features for capturing the spatial quality of each view of an LF image. Second, the depth feature extraction scheme is designed to explore the geometry information inherited in an LF image for modeling its LF structural consistency across views. The similarity measurements are subsequently conducted on the comparison of their symmetry and depth features separately, which are further combined to achieve the quality score for the distorted LF image. Note that the proposed SDFM that explores the symmetry and depth features is conformable to the human vision system, which identifies the objects by sensing their structures and geometries. Extensive simulation results on the dense light fields dataset have clearly shown that the proposed SDFM outperforms multiple classical and recently developed IQA algorithms on quality evaluation of the LF images.
Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Circuits Syst. Video Technol.3
2021 Subjective Quality Database and Objective Study of Compressed Point Clouds With 6DoF Head-Mounted Display
abstract
In this paper, we focus on subjective and objective Point Cloud Quality Assessment (PCQA) in an immersive environment and study the effect of geometry and texture attributes in compression distortion. Using a Head-Mounted Display (HMD) with six degrees of freedom, we establish a subjective PCQA database, named SIAT Point Cloud Quality Database (SIAT-PCQD). Our database consists of 340 distorted point clouds compressed by the MPEG point cloud encoder with the combination of 20 sequences and 17 pairs of geometry and texture quantization parameters. The impact of distorted geometry and texture attributes is further discussed in this paper. Then, we propose two projection-based objective quality evaluation methods, i.e., a weighted view projection based model and a patch projection based model. Our subjective database and findings can be used in point cloud processing, transmission, and coding, especially for virtual reality applications. The subjective datasethttps://dx.doi.org/10.21227/ad8d-7r28http://codec.siat.ac.cn/video_download_siat-pcqd.htmlhasbeen released in the public repository.
Xinju Wu, Yun Zhang 0002, Chunling Fan, Junhui Hou, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2021 Semisupervised Adaptive Symmetric Non-Negative Matrix Factorization
abstract
As a variant of non-negative matrix factorization (NMF), symmetric NMF (SymNMF) can generate the clustering result without additional post-processing, by decomposing a similarity matrix into the product of a clustering indicator matrix and its transpose. However, the similarity matrix in the traditional SymNMF methods is usually predefined, resulting in limited clustering performance. Considering that the quality of the similarity graph is crucial to the final clustering performance, we propose a new semisupervised model, which is able to simultaneously learn the similarity matrix with supervisory information and generate the clustering results, such that the mutual enhancement effect of the two tasks can produce better clustering performance. Our model fully utilizes the supervisory information in the form of pairwise constraints to propagate it for obtaining an informative similarity matrix. The proposed model is finally formulated as a non-negativity-constrained optimization problem. Also, we propose an iterative method to solve it with the convergence theoretically proven. Extensive experiments validate the superiority of the proposed model when compared with nine state-of-the-art NMF models.
Yuheng Jia, Hui Liu 0032, Junhui Hou, Sam Kwong
IEEE Trans. Cybern.3
2021 ASIF-Net: Attention Steered Interweave Fusion Network for RGB-D Salient Object Detection
abstract
Salient object detection from RGB-D images is an important yet challenging vision task, which aims at detecting the most distinctive objects in a scene by combining color information and depth constraints. Unlike prior fusion manners, we propose an attention steered interweave fusion network (ASIF-Net) to detect salient objects, which progressively integrates cross-modal and cross-level complementarity from the RGB image and corresponding depth map via steering of an attention mechanism. Specifically, the complementary features from RGB-D images are jointly extracted and hierarchically fused in a dense and interweaved manner. Such a manner breaks down the barriers of inconsistency existing in the cross-modal data and also sufficiently captures the complementarity. Meanwhile, an attention mechanism is introduced to locate the potential salient regions in an attention-weighted fashion, which advances in highlighting the salient objects and suppressing the cluttered background regions. Instead of focusing only on pixelwise saliency, we also ensure that the detected salient objects have the objectness characteristics (e.g., complete structure and sharp boundary) by incorporating the adversarial learning that provides a global semantic constraint for RGB-D salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably against 17 state-of-the-art saliency detectors on four publicly available RGB-D salient object detection datasets. The code and results of our method are available at https://github.com/Li-Chongyi/ASIF-Net.
Chongyi Li, Runmin Cong, Sam Kwong, Junhui Hou, Huazhu Fu, Guopu Zhu, Dingwen Zhang, Qingming Huang
IEEE Trans. Cybern.4
2021 Underwater Image Enhancement via Medium Transmission-Guided Multi-Color Space Embedding
abstract
Underwater images suffer from color casts and low contrast due to wavelength- and distance-dependent attenuation and scattering. To solve these two degradation issues, we present an underwater image enhancement network via medium transmission-guided multi-color space embedding, called Ucolor. Concretely, we first propose a multi-color space encoder network, which enriches the diversity of feature representations by incorporating the characteristics of different color spaces into a unified structure. Coupled with an attention mechanism, the most discriminative features extracted from multiple color spaces are adaptively integrated and highlighted. Inspired by underwater imaging physical models, we design a medium transmission (indicating the percentage of the scene radiance reaching the camera)-guided decoder network to enhance the response of network towards quality-degraded regions. As a result, our network can effectively improve the visual quality of underwater images by exploiting multiple color spaces embedding and the advantages of both physical model-based and learning-based methods. Extensive experiments demonstrate that our Ucolor achieves superior performance against state-of-the-art methods in terms of both visual quality and quantitative metrics. The code is publicly available at: https://li-chongyi.github.io/Proj_Ucolor.html.
Chongyi Li, Saeed Anwar, Junhui Hou, Runmin Cong, Chunle Guo, Wenqi Ren
IEEE Trans. Image Process.3
2021 Reduced Reference Perceptual Quality Model With Application to Rate Control for Video-Based Point Cloud Compression
abstract
In rate-distortion optimization, the encoder settings are determined by maximizing a reconstruction quality measure subject to a constraint on the bitrate. One of the main challenges of this approach is to define a quality measure that can be computed with low computational cost and which correlates well with the perceptual quality. While several quality measures that fulfil these two criteria have been developed for images and videos, no such one exists for point clouds. We address this limitation for the video-based point cloud compression (V-PCC) standard by proposing a linear perceptual quality model whose variables are the V-PCC geometry and color quantization step sizes and whose coefficients can easily be computed from two features extracted from the original point cloud. Subjective quality tests with 400 compressed point clouds show that the proposed model correlates well with the mean opinion score, outperforming state-of-the-art full reference objective measures in terms of Spearman rank-order and Pearson linear correlation coefficient. Moreover, we show that for the same target bitrate, rate-distortion optimization based on the proposed model offers higher perceptual quality than rate-distortion optimization based on exhaustive search with a point-to-point objective quality metric. Our datasets are publicly available at https://github.com/qdushl/Waterloo-Point-Cloud-Database-2.0.
Qi Liu 0029, Hui Yuan 0001, Raouf Hamzaoui, Honglei Su, Junhui Hou, Huan Yang 0001
IEEE Trans. Image Process.5
2021 Deep Magnification-Flexible Upsampling Over 3D Point Clouds
abstract
This paper addresses the problem of generating dense point clouds from given sparse point clouds to model the underlying geometric structures of objects/scenes. To tackle this challenging issue, we propose a novel end-to-end learning-based framework. Specifically, by taking advantage of the linear approximation theorem, we first formulate the problem explicitly, which boils down to determining the interpolation weights and high-order approximation errors. Then, we design a lightweight neural network to adaptively learn unified and sorted interpolation weights as well as the high-order refinements, by analyzing the local geometry of the input point cloud. The proposed method can be interpreted by the explicit formulation, and thus is more memory-efficient than existing ones. In sharp contrast to the existing methods that work only for a pre-defined and fixed upsampling factor, the proposed framework only requires a single neural network with one-time training to handle various upsampling factors within a typical range, which is highly desired in real-world applications. In addition, we propose a simple yet effective training strategy to drive such a flexible ability. In addition, our method can handle non-uniformly distributed and noisy data well. Extensive experiments on both synthetic and real-world data demonstrate the superiority of the proposed method over state-of-the-art methods both quantitatively and qualitatively. The code will be publicly available at https://github.com/ninaqy/Flexible-PU.
Junhui Hou, Sam Kwong, Ying He 0001
IEEE Trans. Image Process.2
2021 A Self-Training Approach for Point-Supervised Object Detection and Counting in Crowds
abstract
In this article, we propose a novel self-training approach named Crowd-SDNet that enables a typical object detector trained only with point-level annotations (i.e., objects are labeled with points) to estimate both the center points and sizes of crowded objects. Specifically, during training, we utilize the available point annotations to supervise the estimation of the center points of objects directly. Based on a locally-uniform distribution assumption, we initialize pseudo object sizes from the point-level supervisory information, which are then leveraged to guide the regression of object sizes via a crowdedness-aware loss. Meanwhile, we propose a confidence and order-aware refinement scheme to continuously refine the initial pseudo object sizes such that the ability of the detector is increasingly boosted to detect and count objects in crowds simultaneously. Moreover, to address extremely crowded scenes, we propose an effective decoding method to improve the detector's representation ability. Experimental results on the WiderFace benchmark show that our approach significantly outperforms state-of-the-art point-supervised methods under both detection and counting tasks, i.e., our method improves the average precision by more than 10% and reduces the counting error by 31.2%. Besides, our method obtains the best results on the crowd counting and localization datasets (i.e., ShanghaiTech and NWPU-Crowd) and vehicle counting datasets (i.e., CARPK and PUCPR+) compared with state-of-the-art counting-by-detection methods. The code will be publicly available at https://github.com/WangyiNTU/Point-supervised-crowd-detection.
Yi Wang 0068, Junhui Hou, Xinyu Hou, Lap-Pui Chau
IEEE Trans. Image Process.2
2021 Superpixel-Guided Discriminative Low-Rank Representation of Hyperspectral Images for Classification
Shujun Yang, Junhui Hou, Yuheng Jia, Shaohui Mei, Qian Du 0001
IEEE Trans. Image Process.2
2021 Hyperspectral Image Super-Resolution via Deep Progressive Zero-Centric Residual Learning
abstract
This paper explores the problem of hyperspectral image (HSI) super-resolution that merges a low resolution HSI (LR-HSI) and a high resolution multispectral image (HR-MSI). The cross-modality distribution of the spatial and spectral information makes the problem challenging. Inspired by the classic wavelet decomposition-based image fusion, we propose a novel lightweight deep neural network-based framework, namely progressive zero-centric residual network (PZRes-Net), to address this problem efficiently and effectively. Specifically, PZRes-Net learns a high resolution and zero-centric residual image, which contains high-frequency spatial details of the scene across all spectral bands, from both inputs in a progressive fashion along the spectral dimension. And the resulting residual image is then superimposed onto the up-sampled LR-HSI in a mean-value invariant manner, leading to a coarse HR-HSI, which is further refined by exploring the coherence across all spectral bands simultaneously. To learn the residual image efficiently and effectively, we employ spectral-spatial separable convolution with dense connections. In addition, we propose zero-mean normalization implemented on the feature maps of each layer to realize the zero-mean characteristic of the residual image. Extensive experiments over both real and synthetic benchmark datasets demonstrate that our PZRes-Net outperforms state-of-the-art methods to a significant extent in terms of both 4 quantitative metrics and visual quality, e.g., our PZRes-Net improves the PSNR more than 3dB, while saving 2.3× parameters and consuming 15× less FLOPs. The code is publicly available at https://github.com/zbzhzhy/PZRes-Net.
Junhui Hou, Jie Chen 0026, Huanqiang Zeng, Jiantao Zhou 0001
IEEE Trans. Image Process.2
2021 Model-Based Joint Bit Allocation Between Geometry and Color for Video-Based 3D Point Cloud Compression
abstract
In video-based 3D point cloud compression, the quality of the reconstructed 3D point cloud depends on both the geometry, and color distortions. Finding an optimal allocation of the total bitrate between the geometry coder, and the color coder is a challenging task due to the large number of possible solutions. To solve this bit allocation problem, we first propose analytical distortion, and rate models for the geometry, and color information. Using these models, we formulate the joint bit allocation problem as a constrained convex optimization problem, and solve it with an interior point method. Experimental results show that the rate-distortion performance of the proposed solution is close to that obtained with exhaustive search but at only 0.66$\%$of its time complexity.
Qi Liu 0029, Hui Yuan 0001, Junhui Hou, Raouf Hamzaoui, Honglei Su
IEEE Trans. Multim.3
2021 Patch Based Video Summarization With Block Sparse Representation
abstract
In recent years, sparse representation has been successfully utilized for video summarization (VS). However, most of the sparse representation based VS methods characterize each video frame with global features. As a result, some important local details could be neglected by global features, which may compromise the performance of summarization. In this paper, we propose to partition each video frame into a number of patches and characterize each patch with global features. Instead of concatenating the features of each patch and utilizing conventional sparse representation, we formulate the VS problem with such video frame representation as block sparse representation by considering each video frame as a block containing a number of patches. By taking the reconstruction constraint into account, we devise a simultaneous version of block-based OMP (Orthogonal Matching Pursuit) algorithm, namely SBOMP, to solve the proposed model. The proposed model is further extended to a neighborhood based model which considers temporally adjacent frames as a super block. This is one of the first sparse representation based VS methods taking both spatial and temporal contexts into account with blocks. Experimental results on two widely used VS datasets have demonstrated that our proposed methods present clear superiority over existing sparse representation based VS methods and are highly comparable to some deep learning ones requiring supervision information for extra model training.
Shaohui Mei, Mingyang Ma 0004, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.4
2021 Constrained Clustering With Dissimilarity Propagation-Guided Graph-Laplacian PCA
abstract
In this article, we propose a novel model for constrained clustering, namely, the dissimilarity propagation-guided graph-Laplacian principal component analysis (DP-GLPCA). By fully utilizing a limited number of weakly supervisory information in the form of pairwise constraints, the proposed DP-GLPCA is capable of capturing both the local and global structures of input samples to exploit their characteristics for excellent clustering. More specifically, we first formulate a convex semisupervised low-dimensional embedding model by incorporating a new dissimilarity regularizer into GLPCA (i.e., an unsupervised dimensionality reduction model), in which both the similarity and dissimilarity between low-dimensional representations are enforced with the constraints to improve their discriminability. An efficient iterative algorithm based on the inexact augmented Lagrange multiplier is designed to solve it with the global convergence guaranteed. Furthermore, we innovatively propose to propagate the cannot-link constraints (i.e., dissimilarity) to refine the dissimilarity regularizer to be more informative. The resulting DP model is iteratively solved, and we also prove that it can converge to a Karush-Kuhn-Tucker point. Extensive experimental results over nine commonly used benchmark data sets show that the proposed DP-GLPCA can produce much higher clustering accuracy than state-of-the-art constrained clustering methods. Besides, the effectiveness and advantage of the proposed DP model are experimentally verified. To the best of our knowledge, it is the first time to investigate DP, which is contrast to existing pairwise constraint propagation that propagates similarity. The code is publicly available at https://github.com/jyh-learning/DP-GLPCA.
Yuheng Jia, Junhui Hou, Sam Kwong
IEEE Trans. Neural Networks Learn. Syst.2
2021 Joint Optimization for Pairwise Constraint Propagation
abstract
Constrained spectral clustering (SC) based on pairwise constraint propagation has attracted much attention due to the good performance. All the existing methods could be generally cast as the following two steps, i.e., a small number of pairwise constraints are first propagated to the whole data under the guidance of a predefined affinity matrix, and the affinity matrix is then refined in accordance with the resulting propagation and finally adopted for SC. Such a stepwise manner, however, overlooks the fact that the two steps indeed depend on each other, i.e., the two steps form a "chicken-egg" problem, leading to suboptimal performance. To this end, we propose a joint PCP model for constrained SC by simultaneously learning a propagation matrix and an affinity matrix. Especially, it is formulated as a bounded symmetric graph regularized low-rank matrix completion problem. We also show that the optimized affinity matrix by our model exhibits an ideal appearance under some conditions. Extensive experimental results in terms of constrained SC, semisupervised classification, and propagation behavior validate the superior performance of our model compared with state-of-the-art methods.
Yuheng Jia, Wenhui Wu 0001, Ran Wang 0001, Junhui Hou, Sam Kwong
IEEE Trans. Neural Networks Learn. Syst.4
2021 Convolutional Neural Networks With Dynamic Regularization
abstract
Regularization is commonly used for alleviating overfitting in machine learning. For convolutional neural networks (CNNs), regularization methods, such as DropBlock and Shake-Shake, have illustrated the improvement in the generalization performance. However, these methods lack a self-adaptive ability throughout training. That is, the regularization strength is fixed to a predefined schedule, and manual adjustments are required to adapt to various network architectures. In this article, we propose a dynamic regularization method for CNNs. Specifically, we model the regularization strength as a function of the training loss. According to the change of the training loss, our method can dynamically adjust the regularization strength in the training procedure, thereby balancing the underfitting and overfitting of CNNs. With dynamic regularization, a large-scale model is automatically regularized by the strong perturbation, and vice versa. Experimental results show that the proposed method can improve the generalization capability on off-the-shelf network architectures and outperform state-of-the-art regularization methods.
Yi Wang 0068, Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau
IEEE Trans. Neural Networks Learn. Syst.3
2020 Learning Light Field Angular Super-Resolution via a Geometry-Aware Network
abstract
The acquisition of light field images with high angular resolution is costly. Although many methods have been proposed to improve the angular resolution of a sparsely-sampled light field, they always focus on the light field with a small baseline, which is captured by a consumer light field camera. By making full use of the intrinsic geometry information of light fields, in this paper we propose an end-to-end learning-based approach aiming at angularly super-resolving a sparsely-sampled light field with a large baseline. Our model consists of two learnable modules and a physically-based module. Specifically, it includes a depth estimation module for explicitly modeling the scene geometry, a physically-based warping for novel views synthesis, and a light field blending module specifically designed for light field reconstruction. Moreover, we introduce a novel loss function to promote the preservation of the light field parallax structure. Experimental results over various light field datasets including large baseline light field images demonstrate the significant superiority of our method when compared with state-of-the-art ones, i.e., our method improves the PSNR of the second best method up to 2 dB in average, while saves the execution time 48×. In addition, our method preserves the light field parallax structure better.
Jing Jin 0006, Junhui Hou, Hui Yuan 0001, Sam Kwong
AAAI2
2020 Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement
abstract
The paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network. Our method trains a lightweight deep network, DCE-Net, to estimate pixel-wise and high-order curves for dynamic range adjustment of a given image. The curve estimation is specially designed, considering pixel value range, monotonicity, and differentiability. Zero-DCE is appealing in its relaxed assumption on reference images, i.e., it does not require any paired or unpaired data during training. This is achieved through a set of carefully formulated non-reference loss functions, which implicitly measure the enhancement quality and drive the learning of the network. Our method is efficient as image enhancement can be achieved by an intuitive and simple nonlinear curve mapping. Despite its simplicity, we show that it generalizes well to diverse lighting conditions. Extensive experiments on various benchmarks demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively. Furthermore, the potential benefits of our Zero-DCE to face detection in the dark are discussed.
Chunle Guo, Chongyi Li, Jichang Guo, Chen Change Loy, Junhui Hou, Sam Kwong, Runmin Cong
CVPR5
2020 Light Field Spatial Super-Resolution via Deep Combinatorial Geometry Embedding and Structural Consistency Regularization
abstract
Light field (LF) images acquired by hand-held devices usually suffer from low spatial resolution as the limited sampling resources have to be shared with the angular dimension. LF spatial super-resolution (SR) thus becomes an indispensable part of the LF camera processing pipeline. The high-dimensionality characteristic and complex geometrical structure of LF images makes the problem more challenging than traditional single-image SR. The performance of existing methods are still limited as they fail to thoroughly explore the coherence among LF views and are insufficient in accurately preserving the parallax structure of the scene. In this paper, we propose a novel learning-based LF spatial SR framework, in which each view of an LF image is first individually super-resolved by exploring the complementary information among views with combinatorial geometry embedding. For accurate preservation of the parallax structure among the reconstructed views, a regularization network trained over a structure-aware loss function is subsequently appended to enforce correct parallax relationships over the intermediate estimation. Our proposed approach is evaluated over datasets with a large number of testing images including both synthetic and real-world scenes. Experimental results demonstrate the advantage of our approach over state-of-the-art methods, i.e., our method not only improves the average PSNR by more than 1.0 dB but also preserves more accurate parallax details, at a lower computation cost.
Jing Jin 0006, Junhui Hou, Jie Chen 0026, Sam Kwong
CVPR2
2020 Deep Spatial-Angular Regularization for Compressive Light Field Reconstruction over Coded Apertures
Mantang Guo, Junhui Hou, Jing Jin 0006, Jie Chen 0026, Lap-Pui Chau
ECCV (2)2
2020 PUGeo-Net: A Geometry-Centric Network for 3D Point Cloud Upsampling
Junhui Hou, Sam Kwong, Ying He 0001
ECCV (19)2
2020 Surface Consistent Light Field Extrapolation Over Stratified Disparity And Spatial Granularities
abstract
The light field captures both the spatial and angular configurations of the scene, which facilitates a wide range of imaging possibilities. In this work, we propose an LF view extrapolation algorithm which renders high quality novel LF views far outside the range of given angular baselines. A stratified synthesis strategy is adopted which projects the scene content based on stratified disparity layers and across varying scales of spatial granularities. Such a stratified methodology proves to help preserve scene structures over large angular shifts, and provide informative clues for inferring the contents of occluded regions. A generative-adversarial network model is further adopted for parallax correction and occlusion completion conditioned on surface consistent feature. Experiments show that our proposed model can provide more reliable novel view extrapolation quality at large baseline extension ratios compared with state-of-the-art LF synthesis algorithms.
Jie Chen 0026, Lap-Pui Chau, Junhui Hou
ICME3
2020 Accurate Light Field Depth Estimation via an Occlusion-Aware Network
abstract
Depth estimation is a fundamental problem for light field based applications. Although recent learning-based methods have proven to be effective for light field depth estimation, they still have troubles when handling occlusion regions. In this paper, by leveraging the explicitly learned occlusion map, we propose an occlusion-aware network, which is capable of estimating accurate depth maps with sharp edges. Our main idea is to separate the depth estimation on non-occlusion and occlusion regions, as they contain different properties with respect to the light field structure, i.e., obeying and violating the angular photo consistency constraint. To this end, three modules are involved in our network: the occlusion region detection network (ORDNet), the coarse depth estimation network (CDENet), and the refined depth estimation network (RDENet). Specifically, ORDNet predicts the occlusion map as a mask, while under the guidance of the resulting occlusion map, CDENet and REDNet focus on the depth estimation on non-occlusion and occlusion areas, respectively. Experimental results show that our method achieves better performance on 4D light field benchmark, especially in occlusion regions, when compared with current state-of-the-art light-field depth estimation algorithms.
Chunle Guo, Jing Jin 0006, Junhui Hou, Jie Chen 0026
ICME3
2020 Lossy Geometry Compression Of 3d Point Cloud Data Via An Adaptive Octree-Guided Network
abstract
In this paper, we propose a deep learning based framework for point cloud geometry lossy compression via hybrid representation of point cloud. First, the input raw 3D point cloud data is adaptively decomposed into non-overlapping local patches through adaptive Octree decomposition and clustering. Second, a framework of point cloud auto-encoder network with quantization layer is proposed for learning compact latent feature representation from each patch. Specifically, the proposed point cloud auto-encoder networks with different input size are trained for achieving optimal rate-distortion (RD) performance. Final, bitstream specifications of proposed compression systems with additional signaled meta-data and header information are designed to support parallel decoding and successive reconstruction. Experimental results shows that our proposed method can achieve 40.20% bitrate saving in average than the existing standard Geometry based Point Cloud Compression (G-PCC) codec.
Xuanzheng Wen, Xu Wang 0006, Junhui Hou, Lin Ma 0002, Yu Zhou 0027, Jianmin Jiang
ICME3
2020 Light Field Super-resolution via Attention-Guided Fusion of Hybrid Lenses
abstract
This paper explores the problem of reconstructing high-resolution light field (LF) images from hybrid lenses, including a high-resolution camera surrounded by multiple low-resolution cameras. To tackle this challenge, we propose a novel end-to-end learning-based approach, which can comprehensively utilize the specific characteristics of the input from two complementary and parallel perspectives. Specifically, one module regresses a spatially consistent intermediate estimation by learning a deep multidimensional and cross-domain feature representation; the other one constructs another intermediate estimation, which maintains the high-frequency textures, by propagating the information of the high-resolution view. We finally leverage the advantages of the two intermediate estimations via the learned attention maps, leading to the final high-resolution LF image. Extensive experiments demonstrate the significant superiority of our approach over state-of-the-art ones. That is, our method not only improves the PSNR by more than 2 dB, but also preserves the LF structure much better. To the best of our knowledge, this is the first end-to-end deep learning method for reconstructing a high-resolution LF image with a hybrid input. We believe our framework could potentially decrease the cost of high-resolution LF data acquisition and also be beneficial to LF data storage and transmission. The code is available at https://github.com/jingjin25/LFhybridSR-Fusion.
Jing Jin 0006, Junhui Hou, Jie Chen 0026, Sam Kwong, Jingyi Yu 0001
ACM Multimedia2
2020 CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object Detection
abstract
Co-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue is how to effectively capture co-saliency cues by modeling and exploiting inter-image relationships. In this paper, we present an end-to-end collaborative aggregation-and-distribution network (CoADNet) to capture both salient and repetitive visual patterns from multiple images. First, we integrate saliency priors into the backbone features to suppress the redundant background information through an online intra-saliency guidance structure. After that, we design a two-stage aggregate-and-distribute architecture to explore group-wise semantic interactions and produce the co-saliency features. In the first stage, we propose a group-attentional semantic aggregation module that models inter-image relationships to generate the group-wise semantic representations. In the second stage, we propose a gated group distribution module that adaptively distributes the learned group semantics to different individuals in a dynamic gating mechanism. Finally, we develop a group consistency preserving decoder tailored for the CoSOD task, which maintains group constraints during feature decoding to predict more consistent full-resolution co-saliency maps. The proposed CoADNet is evaluated on four prevailing CoSOD benchmark datasets, which demonstrates the remarkable performance improvement over ten state-of-the-art competitors.
Qijian Zhang, Runmin Cong, Junhui Hou, Chongyi Li, Yao Zhao 0001
NeurIPS3
2020 Video summarization via block sparse dictionary selection
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing4
2020 Hyperspectral Image Classification via Sparse Representation With Incremental Dictionaries
abstract
In this letter, we propose a new sparse representation (SR)-based method for hyperspectral image (HSI) classification, namely SR with incremental dictionaries (SRID). Our SRID boosts existing SR-based HSI classification methods significantly, especially when used for the task with extremely limited training samples. Specifically, by exploiting unlabeled pixels with spatial information and multiple-feature-based SR classifiers, we select and add some of them to dictionaries in an iterative manner, such that the representation abilities of the dictionaries are progressively augmented, and likewise more discriminative representations. In addition, to deal with large-scale data sets, we use a certainty sampling strategy to control the sizes of the dictionaries, such that the computational complexity is well balanced. Experiments over two benchmark data sets show that our proposed method achieves higher classification accuracy than the state-of-the-art methods, i.e., the overall classification accuracy can improve more than 4%.
Shujun Yang, Junhui Hou, Yuheng Jia, Shaohui Mei, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2020 Single image-based head pose estimation with spherical parametrization and 3D morphing
Hui Yuan 0001, Junhui Hou, Jimin Xiao
Pattern Recognit.3
2020 3D Point Cloud Attribute Compression via Graph Prediction
abstract
3D point clouds associated with attributes are considered as a promising data representation for immersive communication. The large amount of data, however, poses great challenges to the subsequent transmission and storage processes. In this letter, we propose a new compression scheme for the color attribute of static voxelized 3D point clouds. Specifically, we first partition the colors of a 3D point cloud into clusters by applying k-d tree to the geometry information, which are then successively encoded. To eliminate the redundancy, we propose a novel prediction module, namely graph prediction, in which a small number of representative points selected from previously encoded clusters are used to predict the points to be encoded by exploring the underlying graph structure constructed from the geometry information. Furthermore, the prediction residuals are transformed with the graph transform, and the resulting transform coefficients are finally uniformly quantified and entropy encoded. Experimental results show that the proposed compression scheme is able to achieve better rate-distortion performance at a lower computational cost when compared with state-of-the-art methods.
Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001
IEEE Signal Process. Lett.2
2020 Non-Negative Transfer Learning With Consistent Inter-Domain Distribution
abstract
In this letter, we propose a novel transfer learning approach, which simultaneously exploits the intra-domain differentiation and inter-domain correlation to comprehensively solve the drawbacks many existing transfer learning methods suffer from, i.e., they either are unable to handle the negative samples or have strict assumptions on the distribution. Specifically, the sample selection strategy is introduced to handle negative samples by using the local geometry structure and the label information of source samples. Furthermore, the pseudo target label is imposed to slack the assumption on the inter-domain distribution for considering the inter-domain correlation. Then, an efficient alternating iterative algorithm is proposed to solve the formulated optimization problem with multiple constraints. The extensive experiments conducted on eleven real-world datasets show the superiority of our method over state-of-the-art approaches, i.e., our method achieves 11.23% improvement on the MNIST dataset.
Zhihao Peng 0002, Yuheng Jia, Junhui Hou
IEEE Signal Process. Lett.3
2020 Correlation Filter Tracking via Distractor-Aware Learning and Multi-Anchor Detection
abstract
Correlation filter has demonstrated the power in object tracking, benefiting from its superior speed and competitive performance. However, existing correlation filter based trackers (CFTs) are fragile for some inherent defects caused by the boundary effect. To address this issue, we propose a novel correlation filter based tracking framework by integrating three highly collaborative components, including a fast target proposal module, a distractor-aware filter, and a correlation filter based refiner. Specifically, the target proposal aims at determining some target-like regions in contexts efficiently, which provides target-like patches to learn a distractor-aware filter and detect. Multi-region strategy enlarges space fields for learning and prediction. The filter learned from both target and distractors enhances its ability to identify background. Therefore, our method is capable of evaluating multiple candidates in wider context with less risk of drifting to distractors, namely multi-anchor detection. Besides, the proposed Proposal-Detect-Refine hierarchical searching process progressively achieves data alignment between testing and training samples, which benefits for reliable model prediction. A refiner is used to fine-tune positions after multi-anchor detection for lessening error accumulation and preventing model from drifting. Comprehensive experiments on five challenging datasets, i.e. OTB2013, OTB2015, VOT2017, VOT19, and TC128, demonstrate that the proposed method achieves superior performance against the state-of-the-art methods.
Guochun Chen, Gengzheng Pan, Yongxin Zhou 0003, Wenxiong Kang, Junhui Hou, Feiqi Deng
IEEE Trans. Circuits Syst. Video Technol.5
2020 Going From RGB to RGBD Saliency: A Depth-Guided Transformation Model
abstract
Depth information has been demonstrated to be useful for saliency detection. However, the existing methods for RGBD saliency detection mainly focus on designing straightforward and comprehensive models, while ignoring the transferable ability of the existing RGB saliency detection models. In this article, we propose a novel depth-guided transformation model (DTM) going from RGB saliency to RGBD saliency. The proposed model includes three components, that is: 1) multilevel RGBD saliency initialization; 2) depth-guided saliency refinement; and 3) saliency optimization with depth constraints. The explicit depth feature is first utilized in the multilevel RGBD saliency model to initialize the RGBD saliency by combining the global compactness saliency cue and local geodesic saliency cue. The depth-guided saliency refinement is used to further highlight the salient objects and suppress the background regions by introducing the prior depth domain knowledge and prior refined depth shape. Benefiting from the consistency of the entire object in the depth map, we formulate an optimization model to attain more consistent and accurate saliency results via an energy function, which integrates the unary data term, color smooth term, and depth consistency term. Experiments on three public RGBD saliency detection benchmarks demonstrate the effectiveness and performance improvement of the proposed DTM from RGB to RGBD saliency.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Junhui Hou, Qingming Huang, Sam Kwong
IEEE Trans. Cybern.4
2020 Screen Content Video Quality Assessment: Subjective and Objective Study
abstract
In this paper, we make the first attempt to study the subjective and objective quality assessment for the screen content videos (SCVs). For that, we construct the first large-scale video quality assessment (VQA) database specifically for the SCVs, called the screen content video database (SCVD). This SCVD provides 16 reference SCVs, 800 distorted SCVs, and their corresponding subjective scores, and it is made publicly available for research usage. The distorted SCVs are generated from each reference SCV with 10 distortion types and 5 degradation levels for each distortion type. Each distorted SCV is rated by at least 32 subjects in the subjective test. Furthermore, we propose the first full-reference VQA model for the SCVs, called the spatiotemporal Gabor feature tensor-based model (SGFTM), to objectively evaluate the perceptual quality of the distorted SCVs. This is motivated by the observation that 3D-Gabor filter can well stimulate the visual functions of the human visual system (HVS) on perceiving videos, being more sensitive to the edge and motion information that are often-encountered in the SCVs. Specifically, the proposed SGFTM exploits 3D-Gabor filter to individually extract the spatiotemporal Gabor feature tensors from the reference and distorted SCVs, followed by measuring their similarities and later combining them together through the developed spatiotemporal feature tensor pooling strategy to obtain the final SGFTM score. Experimental results on SCVD have shown that the proposed SGFTM yields a high consistency on the subjective perception of SCV quality and consistently outperforms multiple classical and state-of-the-art image/video quality assessment models.
Shan Cheng, Huanqiang Zeng, Jing Chen 0001, Junhui Hou, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Image Process.4
2020 3D Point Cloud Attribute Compression Using Geometry-Guided Sparse Representation
abstract
3D point clouds associated with attributes are considered as a promising paradigm for immersive communication. However, the corresponding compression schemes for this media are still in the infant stage. Moreover, in contrast to conventional image/video compression, it is a more challenging task to compress 3D point cloud data, arising from the irregular structure. In this paper, we propose a novel and effective compression scheme for the attributes of voxelized 3D point clouds. In the first stage, an input voxelized 3D point cloud is divided into blocks of equal size. Then, to deal with the irregular structure of 3D point clouds, a geometry-guided sparse representation (GSR) is proposed to eliminate the redundancy within each block, which is formulated as an ℓ0-norm regularized optimization problem. Also, an inter-block prediction scheme is applied to remove the redundancy between blocks. Finally, by quantitatively analyzing the characteristics of the resulting transform coefficients by GSR, an effective entropy coding strategy that is tailored to our GSR is developed to generate the bitstream. Experimental results over various benchmark datasets show that the proposed compression scheme is able to achieve better rate-distortion performance and visual quality, compared with state-of-the-art methods.
Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001, Kai-Kuang Ma
IEEE Trans. Image Process.2
2020 An Underwater Image Enhancement Benchmark Dataset and Beyond
abstract
Underwater image enhancement has been attracting much attention due to its significance in marine engineering and aquatic robotics. Numerous underwater image enhancement algorithms have been proposed in the last few years. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real-world images. It is thus unclear how these algorithms would perform on images acquired in the wild and how we could gauge the progress in the field. To bridge this gap, we present the first comprehensive perceptual study and analysis of underwater image enhancement using large-scale real-world images. In this paper, we construct an Underwater Image Enhancement Benchmark (UIEB) including 950 real-world underwater images, 890 of which have the corresponding reference images. We treat the rest 60 underwater images which cannot obtain satisfactory reference images as challenging data. Using this dataset, we conduct a comprehensive study of the state-of-the-art underwater image enhancement algorithms qualitatively and quantitatively. In addition, we propose an underwater image enhancement network (called Water-Net) trained on this benchmark as a baseline, which indicates the generalization of the proposed UIEB for training Convolutional Neural Networks (CNNs). The benchmark evaluations and the proposed Water-Net demonstrate the performance and limitations of state-of-the-art algorithms, which shed light on future research in underwater image enhancement. The dataset and code are available at.
Chongyi Li, Chunle Guo, Wenqi Ren, Runmin Cong, Junhui Hou, Sam Kwong, Dacheng Tao
IEEE Trans. Image Process.5
2020 Light Field Image Quality Assessment via the Light Field Coherence
abstract
In this paper, a novel full-referenceimage quality assessment(IQA) method for evaluating the quality of the distortedlight field(LF) image against its reference LF image is proposed, called thelog-Gabor feature-basedlight field coherence (LGF-LFC). Based on the fact that to compare two LF images, it essentially boils down to measure howcoherentof these two LF images, we attempt to measure the degree of their LFcoherence(LFC). To pursue this goal, the salient features from the reference and distorted LF images under comparison need to be extracted. By considering that the Gabor feature has the ability to well characterize thehuman visual system(HVS) perception, and the special characteristics of the LF images, themulti-scale andsingle-scale Gabor feature extraction schemes are developed to extract the multi-scale log-Gabor features from thesub-aperture images(SAIs) and the single-scale log-Gabor feature from theepi-polar images(EPIs), respectively. Note that the former can reflect the image details (via the SAIs), while the latter indicates the viewing consistency (via the EPI’s depth information). The similarity measurements are subsequently conducted on the comparison of their SAIs and that of their EPIs separately, followed by combining them together for arriving at the final score. Extensive simulation results have clearly demonstrated that the proposed LGF-LFC is more consistent with the perception of the HVS on the quality evaluation of the LF images than multiple classical and state-of-the-art IQA methods.
Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma
IEEE Trans. Image Process.3
2020 Semi-Supervised Non-Negative Matrix Factorization With Dissimilarity and Similarity Regularization
abstract
In this article, we propose a semi-supervised non-negative matrix factorization (NMF) model by means of elegantly modeling the label information. The proposed model is capable of generating discriminable low-dimensional representations to improve clustering performance. Specifically, a pair of complementary regularizers, i.e., similarity and dissimilarity regularizers, is incorporated into the conventional NMF to guide the factorization. And, they impose restrictions on both the similarity and dissimilarity of the low-dimensional representations of data samples with labels as well as a small number of unlabeled ones. The proposed model is formulated as a well-posed constrained optimization problem and further solved with an efficient alternating iterative algorithm. Moreover, we theoretically prove that the proposed algorithm can converge to a limiting point that meets the Karush-Kuhn-Tucker conditions. Extensive experiments as well as comprehensive analysis demonstrate that the proposed model outperforms the state-of-the-art NMF methods to a large extent over five benchmark data sets, i.e., the clustering accuracy increases to 82.2% from 57.0%.
Yuheng Jia, Sam Kwong, Junhui Hou, Wenhui Wu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2020 Pairwise Constraint Propagation With Dual Adversarial Manifold Regularization
abstract
Pairwise constraints (PCs) composed of must-links (MLs) and cannot-links (CLs) are widely used in many semisupervised tasks. Due to the limited number of PCs, pairwise constraint propagation (PCP) has been proposed to augment them. However, the existing PCP algorithms only adopt a single matrix to contain all the information, which overlooks the differences between the two types of links such that the discriminability of the propagated PCs is compromised. To this end, this article proposes a novel PCP model via dual adversarial manifold regularization to fully explore the potential of the limited initial PCs. Specifically, we propagate MLs and CLs with two separated variables, called similarity and dissimilarity matrices, under the guidance of the graph structure constructed from data samples. At the same time, the adversarial relationship between the two matrices is taken into consideration. The proposed model is formulated as a nonnegative constrained minimization problem, which can be efficiently solved with convergence theoretically guaranteed. We conduct extensive experiments to evaluate the proposed model, including propagation effectiveness and applications on constrained clustering and metric learning, all of which validate the superior performance of our model to state-of-the-art PCP models.
Yuheng Jia, Hui Liu 0032, Junhui Hou, Sam Kwong
IEEE Trans. Neural Networks Learn. Syst.3
2019 An Articulated Structure-aware Network for 3D Human Pose Estimation
abstract
In this paper, we propose a new end-to-end articulated structure-aware network to regress 3D joint coordinates from the given 2D joint detections. The proposed method is capable of dealing with hard joints well that usually fail existing methods. Specifically, our framework cascades a refinement network with a basic network for two types of joints, and employs a attention module to simulate a camera projection model. In addition, we propose to use a random enhancement module to intensify the constraints between joints. Experimental results on the Human3.6M and HumanEva databases demonstrate the effectiveness and flexibility of the proposed network, and errors of hard joints and bone lengths are significantly reduced, compared with state-of-the-art approaches.
Zhenhua Tang 0001, Xiaoyan Zhang 0002, Junhui Hou
ACML3
2019 Object Counting in Video Surveillance Using Multi-scale Density Map Regression
abstract
In this paper, we present an effective convolutional neural network (CNN) for object counting in video surveillance, namely multi-scale density map regressor (MSDMR). In contrast to existing CNN-based methods that achieve high accuracy by means of empirically increasing the model capacity with more complex structures/layers, we focus on a compact CNN. Specifically, the MSDMR is mainly designed with the supervision of multi-scale outputs, in which two CNN stacks estimate coarse- and fine-scale density maps, respectively. The integral of the fine density map provides the count of objects. The two stacks are connected in a cascaded manner and jointly trained such that the overall model can learn discriminative and complementary features to produce expressive performance. Experimental results show that the proposed MSDMR can achieve higher accuracy compared with state-of-the-art methods on the surveillance datasets.
Yi Wang 0068, Junhui Hou, Lap-Pui Chau
ICASSP2
2019 Imbalance-aware Pairwise Constraint Propagation
abstract
Pairwise constraint propagation (PCP) aims to propagate a limited number of initial pairwise constraints (PCs, including must-link and cannot-link constraints) from the constrained data samples to the unconstrained ones to boost subsequent PC-based applications. The existing PCP approaches always suffer from the imbalance characteristic of PCs, which limits their performance significantly. To this end, we propose a novel imbalance-aware PCP method, by comprehensively and theoretically exploring the intrinsic structures of the underlying PCs. Specifically, different from the existing methods that adopt a single representation, we propose to use two separate carriers to represent the two types of links. And the propagation is driven by the structure embedded in data samples and the regularization of the local, global, and complementary structures of the two carries. Our method is elegantly cast as a well-posed constrained optimization model, which can be efficiently solved. Experimental results demonstrate that the proposed PCP method is capable of generating more high-fidelity PCs than the recent PCP algorithms. In addition, the augmented PCs by our method produce higher accuracy than state-of-the-art semi-supervised clustering methods when applied to constrained clustering. To the best of our knowledge, this is the first PCP method taking the imbalance property of PCs into account.
Hui Liu 0032, Yuheng Jia, Junhui Hou, Qingfu Zhang 0001
ACM Multimedia3
2019 Correlation filter tracker with siamese: A robust and real-time object tracking framework
Gengzheng Pan, Guochun Chen, Wenxiong Kang, Junhui Hou
Neurocomputing4
2019 ImmerTai: Immersive Motion Learning in VR Environments
Xiaoming Chen 0006, Zhibo Chen 0001, Tianyu He, Junhui Hou, Sen Liu 0001, Ying He 0001
J. Vis. Commun. Image Represent.5
2019 3D human pose estimation via human structure-aware fully connected network
Xiaoyan Zhang 0002, Zhenhua Tang 0001, Junhui Hou, Yanbin Hao
Pattern Recognit. Lett.3
2019 Permuted Sparse Representation for 3D Point Clouds
abstract
The irregular structure of a 3D point cloud, which is composed of the 3D coordinates of irregularly sampled points, poses great challenges to its sparse representation. In this letter, by taking advantage of the permutation-invariant characteristic, we propose a novel method for sparsely representing 3D point clouds, namely permuted sparse representation (PSR). Specifically, we permute the points of a 3D point cloud for increasing its regularity to adapt to a predefined transform, e.g., discrete cosine/wavelet transform. More precisely, the permutation is directly driven by optimizing the objective of sparse representation. Our PSR is elegantly and explicitly formulated as a constrained optimization problem, and an efficient algorithm is proposed to solve it iteratively with the convergence guaranteed. Experimental results demonstrate the advantage of our PSR over the existing ones, i.e., with the same approximation error, the number of non-zero coefficients by our method is only 30% of that of the existing method.
Junhui Hou
IEEE Signal Process. Lett.1
2019 Light Field Image Compression Based on Bi-Level View Compensation With Rate-Distortion Optimization
abstract
Compared with conventional color images, light field images (LFIs) contain richer scene information, which allows a wide range of interesting applications. However, such additional information is obtained at the cost of generating substantially more data, which poses challenges to both data storage and transmission. In this paper, we propose a new hybrid framework for effective compression of LFIs. The proposed framework takes the particular characteristics of LFIs into account so that the inter- and intra-view correlations of LFIs can be more efficiently exploited to produce better compression performance. Specifically, the proposed scheme partitions sub-aperture images (SAIs) of an LFI into two groups, namely, key SAIs and non-key SAIs. Bi-level view compensation is proposed to exploit the inter-view correlation: first, based on the group of selected key SAIs, learning-based angular super-resolution is performed to compensate non-key SAIs in pixel-wise, during which heterogeneous inter-view correlation between the non-key SAIs is efficiently removed; second, the two groups of SAIs are respectively reorganized as pseudo-sequences, and block-wise motion compensation is carried out with a standard video encoder, during which the homogeneous inter-view correlation is subsequently exploited. The video encoder also helps to remove the intra-view correlation of the SAIs and finally generates the encoded bitstream. Moreover, the bits allocated to each group are optimally determined via model-based rate distortion optimization. Extensive experimental evaluations and comparisons demonstrate the advantage of the proposed framework over existing methods in terms of rate-distortion performance.
Junhui Hou, Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.1
2019 Nested Network With Two-Stream Pyramid for Salient Object Detection in Optical Remote Sensing Images
abstract
Arising from the various object types and scales, diverse imaging orientations, and cluttered backgrounds in optical remote sensing image (RSI), it is difficult to directly extend the success of salient object detection for nature scene image to the optical RSI. In this paper, we propose an end-to-end deep network called LV-Net based on the shape of network architecture, which detects salient objects from optical RSIs in a purely data-driven fashion. The proposed LV-Net consists of two key modules, i.e., a two-stream pyramid module (L-shaped module) and an encoder-decoder module with nested connections (V-shaped module). Specifically, the L-shaped module extracts a set of complementary information hierarchically by using a two-stream pyramid structure, which is beneficial to perceiving the diverse scales and local details of salient objects. The V-shaped module gradually integrates encoder detail features with decoder semantic features through nested connections, which aims at suppressing the cluttered backgrounds and highlighting the salient objects. In addition, we construct the first publicly available optical RSI data set for salient object detection, including 800 images with varying spatial resolutions, diverse saliency types, and pixel-wise ground truth. Experiments on this benchmark data set demonstrate that the proposed method outperforms the state-of-the-art salient object detection methods both qualitatively and quantitatively.
Chongyi Li, Runmin Cong, Junhui Hou, Sanyi Zhang, Sam Kwong
IEEE Trans. Geosci. Remote. Sens.3
2019 Simultaneous Dimensionality Reduction and Classification via Dual Embedding Regularized Nonnegative Matrix Factorization
abstract
Nonnegative matrix factorization (NMF) is a well-known paradigm for data representation. Traditional NMF-based classification methods first perform NMF or one of its variants on input data samples to obtain their low-dimensional representations, which are successively classified by means of a typical classifier [e.g., k -nearest neighbors (KNN) and support vector machine (SVM)]. Such a stepwise manner may overlook the dependency between the two processes, resulting in the compromise of the classification accuracy. In this paper, we elegantly unify the two processes by formulating a novel constrained optimization model, namely dual embedding regularized NMF (DENMF), which is semi-supervised. Our DENMF solution simultaneously finds the low-dimensional representations and assignment matrix via joint optimization for better classification. Specifically, input data samples are projected onto a couple of low-dimensional spaces (i.e., feature and label spaces), and locally linear embedding is employed to preserve the identical local geometric structure in different spaces. Moreover, we propose an alternating iteration algorithm to solve the resulting DENMF, whose convergence is theoretically proven. Experimental results over five benchmark datasets demonstrate that DENMF can achieve higher classification accuracy than state-of-the-art algorithms.
Wenhui Wu 0001, Sam Kwong, Junhui Hou, Yuheng Jia, Horace Ho-Shing Ip
IEEE Trans. Image Process.3
2019 Light Field Spatial Super-Resolution Using Deep Efficient Spatial-Angular Separable Convolution
abstract
Light field (LF) photography is an emerging paradigm for capturing more immersive representations of the real-world. However, arising from the inherent trade-off between the angular and spatial dimensions, the spatial resolution of LF images captured by commercial micro-lens based LF cameras are significantly constrained. In this paper, we propose effective and efficient end-to-end convolutional neural network models for spatially super-resolving LF images. Specifically, the proposed models have an hourglass shape, which allows feature extraction to be performed at the low resolution level to save both computational and memory costs. To fully make use of the four-dimensional (4-D) structure information of LF data in both spatial and angular domains, we propose to use 4-D convolution to characterize the relationship among pixels. Moreover, as an approximation of 4-D convolution, we also propose to use spatialangular separable (SAS) convolutions for more computationallyand memory-efficient extraction of spatial-angular joint features. Extensive experimental results on 57 test LF images with various challenging natural scenes show significant advantages from the proposed models over state-of-the-art methods. That is, an average PSNR gain of more than 3.0 dB and better visual quality are achieved, and our methods preserve the LF structure of the super-resolved LF images better, which is highly desirable for subsequent applications. In addition, the SAS convolutionbased model can achieve 3× speed up with only negligible reconstruction quality decrease when compared with the 4-D convolution-based one. The source code of our method is online available at https://github.com/spatialsr/DeepLightFieldSSR.
Henry Wing Fung Yeung, Junhui Hou, Xiaoming Chen 0006, Jie Chen 0026, Zhibo Chen 0001, Vera Chung
IEEE Trans. Image Process.2
2018 Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework
abstract
Rain removal is important for improving the robustness of outdoor vision based systems. Current rain removal methods show limitations either for complex dynamic scenes shot from fast moving cameras, or under torrential rain fall with opaque occlusions. We propose a novel derain algorithm, which applies superpixel (SP) segmentation to decompose the scene into depth consistent units. Alignment of scene contents are done at the SP level, which proves to be robust towards rain occlusion and fast camera motion. Two alignment output tensors, i.e., optimal temporal match tensor and sorted spatial-temporal match tensor, provide informative clues for rain streak location and occluded background contents to generate an intermediate derain output. These tensors will be subsequently prepared as input features for a convolutional neural network to restore high frequency details to the intermediate output for compensation of mis-alignment blur. Extensive evaluations show that up to 5dB reconstruction PSNR advantage is achieved over state-of-the-art methods. Visual inspection shows that much cleaner rain removal is achieved especially for highly dynamic scenes with heavy and opaque rainfall from a fast moving camera.
Jie Chen 0026, Cheen-Hau Tan, Junhui Hou, Lap-Pui Chau
CVPR3
2018 Fast Light Field Reconstruction with Deep Coarse-to-Fine Modeling of Spatial-Angular Clues
Henry Wing Fung Yeung, Junhui Hou, Jie Chen 0026, Vera Chung, Xiaoming Chen 0006
ECCV (6)2
2018 Convex Constrained Clustering with Graph-Laplacian Pca
abstract
In this paper, we propose a new algorithm for constrained clustering, in which a new regularizer elegantly incorporates a small amount of weakly supervisory information in the form of pair-wise constraints to regularize the similarity between the low-dimensional representations of a set of data samples. By exploring both the local and global structures of the data samples with the guidance of the supervisory information, the proposed algorithm is capable of learning the low-dimensional representations with strong separability. Technically, the proposed algorithm is formulated and relaxed as a convex optimization model, which is further efficiently solved with the global convergence guaranteed. Experimental results on multiple benchmark data sets show that our proposed model can produce higher clustering accuracy than state-of-the-art algorithms.
Yuheng Jia, Sam Kwong, Junhui Hou, Wenhui Wu 0001
ICME3
2018 Hyperspectral Classification Via Spatial Context Exploration with Multi-Scale CNN
abstract
Spatial context has shown to be very useful in hyperspectral image processing. Existing convolutional neural network (CNN)-based methods for hyperspectral classification explore spatial context by single-scale convolution kernels in 2D or 3D shapes. However, such single-scale convolution may not be capable to explore the complex spatial context in a hyperspectral image. In this paper, we propose a multi-scale CNN, MS-CNN to explore the spatial context in different extents, in which adaptive spatial neighborhood convolution kernels are used to simultaneously extract multiple spectral-spatial features from spatial context of pixels. These features obtained by different spatial kernels are then concatenated and fused for further feature extraction and classification. Experimental results show that the proposed adaptive spatial neighborhood convolution are more effective to explore spatial context than traditional single-scale spatial convolution and the performance of the proposed MS-CNN outperforms several state-of-art CNNs for classification of hyperspectral images.
Zhongqi Tian, Jingyu Ji, Shaohui Mei, Junhui Hou, Shuai Wan, Qian Du 0001
IGARSS4
2018 A multi-scale contrast-based image quality assessment model for multi-exposure image fusion
Huanqiang Zeng, Jing Chen 0001, Jianqing Zhu, Junhui Hou
Signal Process.6
2018 Light Field Denoising via Anisotropic Parallax Analysis in a CNN Framework
abstract
Light field (LF) cameras provide perspective information of scenes by taking directional measurements of the focusing light rays. The raw outputs are usually dark with additive camera noise, which impedes subsequent processing and applications. We propose a novel LF denoising framework based on anisotropic parallax analysis (APA). Two convolutional neural networks are jointly designed for the task: first, the structural parallax synthesis network predicts the parallax details for the entire LF based on a set of anisotropic parallax features. These novel features can efficiently capture the high-frequency perspective components of a LF from noisy observations. Second, the view-dependent detail compensation network restores non-Lambertian variation to each LF view by involving view-specific spatial energies. Extensive experiments show that the proposed APA LF denoiser provides a much better denoising performance than state-of-the-art methods in terms of visual quality and in preservation of parallax details.
Jie Chen 0026, Junhui Hou, Lap-Pui Chau
IEEE Signal Process. Lett.2
2018 Semi-Supervised Spectral Clustering With Structured Sparsity Regularization
abstract
Spectral clustering (SC) is one of the most widely used clustering methods. In this letter, we extend the traditional SC with a semi-supervised manner. Specifically, with the guidance of small amount of supervisory information, we build a matrix with anti-block-diagonal appearance, which is further utilized to regularize the product of the low-dimensional embedding and its transpose. Technically, we formulate the proposed model as a constrained optimization problem. Then, we relax it as a convex problem, which can be efficiently solved with the global convergence guaranteed via the inexact augmented Lagrangian multiplier method. Experimental results over four real-world datasets demonstrate that higher accuracy and normalized mutual information are achieved when compared with state-of-the-art methods.
Yuheng Jia, Sam Kwong, Junhui Hou
IEEE Signal Process. Lett.3
2018 Simultaneous Spatial and Spectral Low-Rank Representation of Hyperspectral Images for Classification
abstract
Arising from various environmental and atmos- pheric conditions and sensor interference, spectral variations are inevitable during hyperspectral remote sensing, which degrade the subsequent hyperspectral image analysis significantly. In this paper, we propose simultaneous spatial and spectral low-rank representation (S3LRR) that can effectively suppress the within-class spectral variations for classification purposes. The S3LRR recovers an intrinsic component with the same dimension as the original image, in which both spatial and spectral low-rank priors are adopted to regularize the intrinsic component simultaneously and compensate to each other, together with robust modeling of spectral variations. Compared with existing methods that explore only the spectral low-rank prior, the novel spatial low-rank prior (i.e., low-rank prior in band-wise) can take the spatial structure information of hyperspectral images into account, which has demonstrated to be very useful. Technically, we formulate S3LRR as a constrained convex optimization problem, and solve it using the efficient inexact augmented Lagrangian multiplier method. The resulting intrinsic component is less interfered by within-class spectral variations, and more discriminatory to offer higher classification accuracy. Comprehensive experiments on benchmark data sets demonstrate that the proposed S3LRR improves classification accuracy significantly, which outperforms state-of-the-art methods.
Shaohui Mei, Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 Light Field Compression With Disparity-Guided Sparse Coding Based on Structural Key Views
abstract
Recent imaging technologies are rapidly evolving for sampling richer and more immersive representations of the 3D world. One of the emerging technologies is light field (LF) cameras based on micro-lens arrays. To record the directional information of the light rays, a much larger storage space and transmission bandwidth are required by an LF image as compared with a conventional 2D image of similar spatial dimension. Hence, the compression of LF data becomes a vital part of its application. In this paper, we propose an LF codec with disparity guided Sparse Coding over a learned perspective-shifted LF dictionary based on selected Structural Key Views (SC-SKV). The sparse coding is based on a limited number of optimally selected SKVs; yet the entire LF can be recovered from the coding coefficients. By keeping the approximation identical between encoder and decoder, only the residuals of the non-key views, disparity map, and the SKVs need to be compressed into the bit stream. An optimized SKV selection method is proposed such that most LF spatial information can be preserved. To achieve optimum dictionary efficiency, the LF is divided into several coding regions, over which the reconstruction works individually. Experiments and comparisons have been carried out over benchmark LF data set, which show that the proposed SC-SKV codec produces convincing compression results in terms of both rate-distortion performance and visual quality compared with Joint Exploration Model: with 37.9% BD-rate reduction and 1.17-dB BD-PSNR improvement achieved on average, especially with up to 6-dB improvement for low bit rate scenarios.
Jie Chen 0026, Junhui Hou, Lap-Pui Chau
IEEE Trans. Image Process.2
2018 Accurate Light Field Depth Estimation With Superpixel Regularization Over Partially Occluded Regions
abstract
Depth estimation is a fundamental problem for light field photography applications. Numerous methods have been proposed in recent years, which either focus on crafting cost terms for more robust matching, or on analyzing the geometry of scene structures embedded in the epipolar-plane images. Significant improvements have been made in terms of overall depth estimation error; however, current state-of-the-art methods still show limitations in handling intricate occluding structures and complex scenes with multiple occlusions. To address these challenging issues, we propose a very effective depth estimation framework which focuses on regularizing the initial label confidence map and edge strength weights. Specifically, we first detect partially occluded boundary regions (POBR) via superpixel-based regularization. Series of shrinkage/reinforcement operations are then applied on the label confidence map and edge strength weights over the POBR. We show that after weight manipulations, even a low-complexity weighted least squares model can produce much better depth estimation than the state-of-the-art methods in terms of average disparity error rate, occlusion boundary precision-recall rate, and the preservation of intricate visual features.
Jie Chen 0026, Junhui Hou, Yun Ni, Lap-Pui Chau
IEEE Trans. Image Process.2
2018 A Gabor Feature-Based Quality Assessment Model for the Screen Content Images
abstract
In this paper, an accurate and efficient full-reference image quality assessment (IQA) model using the extracted Gabor features, called Gabor feature-based model (GFM), is proposed for conducting objective evaluation of screen content images (SCIs). It is well-known that the Gabor filters are highly consistent with the response of the human visual system (HVS), and the HVS is highly sensitive to the edge information. Based on these facts, the imaginary part of the Gabor filter that has odd symmetry and yields edge detection is exploited to the luminance of the reference and distorted SCI for extracting their Gabor features, respectively. The local similarities of the extracted Gabor features and two chrominance components, recorded in the LMN color space, are then measured independently. Finally, the Gabor-feature pooling strategy is employed to combine these measurements and generate the final evaluation score. Experimental simulation results obtained from two large SCI databases have shown that the proposed GFM model not only yields a higher consistency with the human perception on the assessment of SCIs but also requires a lower computational complexity, compared with that of classical and state-of-the-art IQA models. The source code for the proposed GFM will be available at http://smartviplab.org/pubilcations/GFM.html.
Zhangkai Ni, Huanqiang Zeng, Lin Ma 0002, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma
IEEE Trans. Image Process.4
2018 Non-Cooperative Game Theory Based Rate Adaptation for Dynamic Video Streaming over HTTP
abstract
Dynamic Adaptive Streaming over HTTP (DASH) has demonstrated to be an emerging and promising multimedia streaming technique, owing to its capability of dealing with the variability of networks. Rate adaptation mechanism, a challenging and open issue, plays an important role in DASH based systems since it affects Quality of Experience (QoE) of users, network utilization, etc. In this paper, based on non-cooperative game theory, we propose a novel algorithm to optimally allocate the limited export bandwidth of the server to multi-users to maximize their QoE with fairness guaranteed. The proposed algorithm is proxy-free. Specifically, a novel user QoE model is derived by taking a variety of factors into account, like the received video quality, the reference buffer length, and user accumulated buffer lengths, etc. Then, the bandwidth competing problem is formulated as a non-cooperation game with the existence of Nash Equilibrium that is theoretically proven. Finally, a distributed iterative algorithm with stability analysis is proposed to find the Nash Equilibrium. Compared with state-of-the-art methods, extensive experimental results in terms of both simulated and realistic networking scenarios demonstrate that the proposed algorithm can produce higher QoE, and the actual buffer lengths of all users keep nearly optimal states, i.e., moving around the reference buffer all the time. Besides, the proposed algorithm produces no playback interruption.
Hui Yuan 0001, Huayong Fu, Junhui Hou, Sam Kwong
IEEE Trans. Mob. Comput.4
2018 Pairwise Constraint Propagation-Induced Symmetric Nonnegative Matrix Factorization
abstract
As a variant of nonnegative matrix factorization (NMF), symmetric NMF (SNMF) has shown to be effective for capturing the cluster structure embedded in the graph representation. In contrast to the existing SNMF-based clustering methods that empirically construct the similarity matrix and rigidly introduce the supervisory information to the assignment matrix, in this paper, we propose a novel SNMF-based semisupervised clustering method, namely, pairwise constraint propagation-induced SNMF (PCPSNMF). By formulating a single-constrained optimization problem, PCPSNMF is capable of learning the similarity and assignment matrices adaptively and simultaneously, in which a small amount of supervisory information in the form of pairwise constraints is introduced in a flexible way to guide the construction of the similarity matrix, and the two matrices communicate with each other to achieve mutual refinement until convergence. In addition, we propose an efficient alternating iterative algorithm to solve the optimization problem, whose convergence is theoretically proven. Experimental results over several benchmark image data sets demonstrate that PCPSNMF is less sensitive to initialization and produces higher clustering performance, compared with the state-of-the-art methods.
Wenhui Wu 0001, Yuheng Jia, Sam Kwong, Junhui Hou
IEEE Trans. Neural Networks Learn. Syst.4
2017 Sparse representation for colors of 3D point cloud via virtual adaptive sampling
abstract
Sparse signal representation has proven to be an extremely powerful tool in a wide range of engineering applications. However, most of the existing techniques are designed for regular data (such as audio signals and images/videos) that uniformly lies in regular Euclidian spaces. This paper aims at extending sparse representation for irregular data (such as colors of 3D point clouds) that is defined on irregular domains embedded in Euclidean spaces. Dealing with the irregular structure of such data via a virtual adaptive sampling process, we formulate sparse representation as an ℓ0-norm regularized optimization problem. Experimental results show that the proposed algorithm outperforms the state-of-the-art algorithm to a large extent: with the same number of nonzero coefficients, we improve the reconstruction quality up to 5 dB; conversely, fixing the reconstruction quality, our method uses only 55% coefficients. Using compressive sensing theory, we provide an intuitive explanation on how and why our algorithm works well in practice.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Philip A. Chou
ICASSP1
2017 Hyperspectral image super-resolution via convolutional neural network
abstract
Due to the tradeoff between spatial and spectral resolution in remote sensing imaging, hyperspectral images are often acquired with a relative low spatial resolution, which limits their applications in many areas. Inspired by recent achievements in convolutional neural network (CNN) based super resolution (SR), a novel CNN based framework is constructed for SR of hyperspectral images by considering both spatial context and spectral correlation. As a result, the spectral distortion incurred by directly applying traditional SR algorithms to hyperspectral images is alleviated. Experimental results on several benchmark hyperspectral datasets have demonstrated that higher quality of reconstruction and spectral fidelity can be achieved, compared to band-wise manner based algorithms.
Shaohui Mei, Xin Yuan 0002, Jingyu Ji, Shuai Wan, Junhui Hou, Qian Du 0001
ICIP5
2017 Nonlinear kernel sparse dictionary selection for video summarization
abstract
Sparse dictionary selection (SDS) has demonstrated to be an effective solution for keyframe based video summarization (VS), which generally assumes a linear relation among similar video frames. However, such a linear assumption is not always true for videos. In this paper, the nonlinearity among frames is taken into consideration and a nonlinear SDS model is formulated for VS, in which the nonlinearity is transformed to linearity by projecting a video to a high dimensional feature space induced by a kernel function. Moreover, a kernel simultaneous orthogonal matching pursuit (KSOMP) is proposed to solve the problem. In order to achieve an intuitive and flexible configuration of the VS process, an adaptive criterion is devised to produce video summaries with different lengths for different video content. Experimental results on benchmark video datasets demonstrate that the proposed algorithm outperforms several state-of-the-art VS algorithms.
Mingyang Ma 0004, Shaohui Mei, Junhui Hou, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICME3
2017 Learning sensor-specific features for hyperspectral images via 3-dimensional convolutional autoencoder
abstract
Deep learning techniques have brought in revolutionary achievements for feature learning of images. In this paper, a novel structure of 3-Dimensional Convolutional AutoEncoder (3D-CAE) is proposed for hyperspectral spatial-spectral feature learning, in which the spatial context is considered by constructing a 3-Dimensional input using pixels in a spatial neighborhood. All the parameters involved in the 3D-CAE are trained without the need of labeled training samples such that feature learning is conducted in an unsupervised fashion. Such unsupervised spatial-spectral feature extraction is also extended to different images from the same sensor to learn sensor-specific features. As a result, spatial-spectral features of hyperspectral images are extracted for a specific sensor under an unsupervised manner. Experimental results on several benchmark hyperspectral datasets have demonstrated that our proposed 3D-CAE are very effective in extracting sensor-specific spatial-spectral features and outperform several state-of-the-art deep learning neural networks in classification application.
Jingyu Ji, Shaohui Mei, Junhui Hou, Xu Li 0010, Qian Du 0001
IGARSS3
2017 Fusing different levels of deep features by deep stacked neural network for hyperspectral images
abstract
Deep learning techniques have been demonstrated to be a powerful tool to learn features of images automatically. In this paper, a novel deep learning structure, i.e., deep stacked neural network (DSNN), is constructed to extract different levels of deep features of hyperspectral images. Specifically, convolutional neural network (CNN) is used as basic units in the proposed DSNN for feature extraction of hyperspectral images. Then, different levels of deep features are concatenated to form a novel fused feature for classification with a typical classifier, e.g., SVM. Experimental results on two benchmark hyperspectral datasets show that the fusion of features extracted in DSNN can produce higher classification accuracy than state-of-the-art deep learning based methods, indicating its effectiveness in feature learning.
Shaohui Mei, Yanfu Chen, Jingyu Ji, Junhui Hou, Qian Du 0001
IGARSS4
2017 Immersive and collaborative Taichi motion learning in various VR environments
abstract
Learning “motion” online or from video tutorials is usually inefficient since it is difficult to deliver “motion” information in traditional ways and in the ordinary PC platform. This paper presents ImmerTai, a system that can efficiently teach motion, in particular Chinese Taichi motion, in various immersive environments. ImmerTai captures the Taichi expert's motion and delivers to students the captured motion in multi-modal forms in immersive CAVE, HMD as well as ordinary PC environments. The students' motions are captured too for quality assessment and utilized to form a virtual collaborative learning atmosphere. We built up a Taichi motion dataset with 150 fundamental Taichi motions captured from 30 students, on which we evaluated the learning effectiveness and user experience of ImmerTai. The results show that ImmerTai can enhance the learning efficiency by up to 17.4% and the learning quality by up to 32.3%.
Tianyu He, Xiaoming Chen 0006, Zhibo Chen 0001, Sen Liu 0001, Junhui Hou, Ying He 0001
VR6
2017 Sparse Low-Rank Matrix Approximation for Data Compression
abstract
Low-rank matrix approximation (LRMA) is a powerful technique for signal processing and pattern analysis. However, its potential for data compression has not yet been fully investigated. In this paper, we propose sparse LRMA (SLRMA), an effective computational tool for data compression. SLRMA extends conventional LRMA by exploring both the intra and inter coherence of data samples simultaneously. With the aid of prescribed orthogonal transforms (e.g., discrete cosine/wavelet transform and graph transform), SLRMA decomposes a matrix into a product of two smaller matrices, where one matrix is made up of extremely sparse and orthogonal column vectors and the other consists of the transform coefficients. Technically, we formulate SLRMA as a constrained optimization problem, i.e., minimizing the approximation error in the least-squares sense regularized by the $\ell _{0}$ -norm and orthogonality, and solve it using the inexact augmented Lagrangian multiplier method. Through extensive tests on real-world data, such as 2D image sets and 3D dynamic meshes, we observe that: 1) SLRMA empirically converges well; 2) SLRMA can produce approximation error comparable to LRMA but in a much sparse form; and 3) SLRMA-based compression schemes significantly outperform the state of the art in terms of rate-distortion performance.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Learning Sensor-Specific Spatial-Spectral Features of Hyperspectral Images via Convolutional Neural Networks
abstract
Convolutional neural network (CNN) is well known for its capability of feature learning and has made revolutionary achievements in many applications, such as scene recognition and target detection. In this paper, its capability of feature learning in hyperspectral images is explored by constructing a five-layer CNN for classification (C-CNN). The proposed C-CNN is constructed by including recent advances in deep learning area, such as batch normalization, dropout, and parametric rectified linear unit (PReLU) activation function. In addition, both spatial context and spectral information are elegantly integrated into the C-CNN such that spatial-spectral features are learned for hyperspectral images. A companion feature-learning CNN (FL-CNN) is constructed by extracting fully connected feature layers in this C-CNN. Both supervised and unsupervised modes are designed for the proposed FL-CNN to learn sensor-specific spatial-spectral features. Extensive experimental results on four benchmark data sets from two well-known hyperspectral sensors, namely airborne visible/infrared imaging spectrometer (AVIRIS) and reflective optics system imaging spectrometer (ROSIS) sensors, demonstrate that our proposed C-CNN outperforms the state-of-the-art CNN-based classification methods, and its corresponding FL-CNN is very effective to extract sensor-specific spatial-spectral features for hyperspectral applications under both supervised and unsupervised modes.
Shaohui Mei, Jingyu Ji, Junhui Hou, Xu Li 0010, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2016 Random Forest with Suppressed Leaves for Hough Voting
Hui Liang 0003, Junhui Hou, Junsong Yuan 0001, Daniel Thalmann
ACCV (3)2
2016 Robust laplacian matrix learning for smooth graph signals
abstract
We propose a new method for robust learning Laplacian matrices from observed smooth graph signals in the presence of both Gaussian noise and random-valued impulse noise (i.e., outliers). Using the recently developed factor analysis model for representing smooth graph signals in [1], we formulate our learning process as a constrained optimization problem, and adopt the £i-norm for measuring the data fidelity in order to improve robustness. Computational results on three types of synthetic graphs demonstrate that the proposed method outperforms the state-of-the-art methods in terms of commonly used information retrieval metrics, such as F-measure, precision, recall and normalized mutual information. In particular, we observed that F-measure is improved by up to 16%.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Huanqiang Zeng
ICIP1
2016 Sparse two-dimensional singular value decomposition
abstract
In this paper, we propose a new data-driven transform, called sparse two-dimensional singular value decomposition (S2DSVD). By leveraging the advantages of discrete cosine transform and the conventional 2D SVD, we decompose a set of matrices into transform coefficient matrices with sparse and orthogonal basis functions. Such sparsity characteristic can significantly reduce their overhead, hence being beneficial to data compression. We formulate S2DSVD as a constrained optimization problem and solve it via alternative iteration. We demonstrate the efficacy of S2DSVD on image and video datasets, and observe that it can produce results with error comparable to 2D SVD whereas its space complexity is significantly smaller than 2D SVD.
Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Ying He 0001
ICME1
2016 How to fully explore the low-rank property for data recovery of hyperspectral images
abstract
The performance of hyperspectral classification is affected by within-class spectral variation since different materials may present similar spectral signatures. In this paper, we investigate how to fully use the low-rank property of hyperspectral images to alleviate spectra variation. Particulary, two effective strategies that explore the low-rank property in local spectral and spatial space are proposed. According to experimental results, we conclude that exploring the low-rank property in local spectral-spatial space can help to alleviate spectral variation and improve the performance of classification obviously for all tested data, while exploring the low-rank property in spatial space is more effective for images presenting large homogeneous areas.
Shaohui Mei, Qianqian Bi, Jingyu Ji, Junhui Hou, Qian Du 0001
IGARSS4
2016 Integrating spectral and spatial information into deep convolutional Neural Networks for hyperspectral classification
abstract
Deep convolutional neural networks (CNNs) have brought in achievements in image classification and target detection. In this paper, we propose a novel five-layer CNN for hyperspectral classification by encountering recent achievement in deep learning area, such as batch normalization, dropout, Parametric Rectified Linear Unit (PReLu) activation function. By taking advantage of the specific characteristics of hyperspectral images, spatial context and spectral information are elegantly integrated into the framework. Experimental results demonstrate that our proposed CNN out- performs the state-of-the-art methods.
Shaohui Mei, Jingyu Ji, Qianqian Bi, Junhui Hou, Qian Du 0001, Wei Li 0032
IGARSS4
2016 Low-latency compression of mocap data using learned spatial decorrelation transform
abstract
Due to the growing needs of motion capture (mocap) in movie, video games , sports, etc., it is highly desired to compress mocap data for efficient storage and transmission. Unfortunately, the existing compression methods have either high latency or poor compression performance , making them less appealing for time-critical applications and/or network with limited bandwidth . This paper presents two efficient methods to compress mocap data with low latency. The first method processes the data in a frame-by-frame manner so that it is ideal for mocap data streaming. The second one is clip-oriented and provides a flexible trade-off between latency and compression performance . It can achieve higher compression performance while keeping the latency fairly low and controllable. Observing that mocap data exhibits some unique spatial characteristics , we learn an orthogonal transform to reduce the spatial redundancy . We formulate the learning problem as the least square of reconstruction error regularized by orthogonality and sparsity , and solve it via alternating iteration. We also adopt a predictive coding and temporal DCT for temporal decorrelation in the frame- and clip-oriented methods, respectively. Experimental results show that the proposed methods can produce higher compression performance at lower computational cost and latency than the state-of-the-art methods. Moreover, our methods are general and applicable to various types of mocap data.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
Comput. Aided Geom. Des.1
2016 Spectral Variation Alleviation by Low-Rank Matrix Approximation for Hyperspectral Image Analysis
abstract
Spectral variation is profound in remotely sensed images due to variable imaging conditions. The wide presence of such spectral variation degrades the performance of hyperspectral analysis, such as classification and spectral unmixing. In this letter, 11-based low-rank matrix approximation is proposed to alleviate spectral variation for hyperspectral image analysis. Specifically, hyperspectral image data are decomposed into a low-rank matrix and a sparse matrix, and it is assumed that intrinsic spectral features are represented by the low-rank matrix and spectral variation is accommodated by the sparse matrix. As a result, the performance of image data analysis can be improved by working on the low-rank matrix. Experiments on benchmark hyperspectral data sets demonstrate the performance of classification, and spectral unmixing can be clearly improved by the proposed approach.
Shaohui Mei, Qianqian Bi, Jingyu Ji, Junhui Hou, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2016 Facial Position and Expression-Based Human-Computer Interface for Persons With Tetraplegia
abstract
A human-computer interface (namely Facial position and expression Mouse system, FM) for the persons with tetraplegia based on a monocular infrared depth camera is presented in this paper. The nose position along with the mouth status (close/open) is detected by the proposed algorithm to control and navigate the cursor as computer user input. The algorithm is based on an improved Randomized Decision Tree, which is capable of detecting the facial information efficiently and accurately. A more comfortable user experience is achieved by mapping the nose motion to the cursor motion via a nonlinear function. The infrared depth camera enables the system to be independent of illumination and color changes both from the background and on human face, which is a critical advantage over RGB camera-based options. Extensive experimental results show that the proposed system outperforms existing assistive technologies in terms of quantitative and qualitative assessments.
Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann
IEEE J. Biomed. Health Informatics2
2015 Reordering-based transform for compressing human motion capture data
abstract
This paper presents a simple yet effective algorithm for compressing human motion capture (mocap) data. With a reordering-based discrete wavelet transform and the standard discrete cosine transform, our method can effectively reduce the spatial and temporal correlation in mocap data. Our method is conceptually simple and easy to implement. Experimental results show that our method can achieve better compression performance with lower latency, compared to the state-of-the-art methods.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
ISCAS1
2015 A linear dependent rate-quantization model for scalable video enhancement layer encoding
abstract
In this paper, we propose a linear dependent rate-quantization model for video enhancement layers encoding in H.264/AVC based scalable video coding (SVC). It is noted that the proposed model is applicable for different scalable structures, such as temporal, quality, spatial and combined scalability. Leveraging the base layer information (such as bitrate and quantization parameter), proposed model can accurately predict the number of bit required for the enhancement layer encoding. Such linear model demonstrates the high accuracy for bitrate estimation at enhancement layers, with the average prediction accuracy over 94%. It has the noticeable improvement from the existing works, without requiring additional complexity increase. Meanwhile, proposed model is applied to do the rate control for enhancement layers encoding. Experimental results show that the average bitrate mismatch error can be significantly reduced compared with the existing algorithms.
Junhui Hou, Shuai Wan, Lap-Pui Chau
ISCAS1
2015 Compressing 3-D Human Motions via Keyframe-Based Geometry Videos
abstract
This paper presents keyframe-based geometry video (KGV), a novel framework for compressing 3-D human motion data by using geometry videos. Given a motion data encoded in a geometry video (GV) format, our method extracts the keyframes and produces a reconstruction matrix. Then it applies the video compression technique (e.g., H.264/Advanced Video Coding) to the reordered keyframes, which can significantly reduce the spatial and temporal redundancy in the KGV. We develop a rate distortion-based optimization algorithm to determine the parameters (i.e., the number of keyframes and quantization parameter) leading to optimal performance. Experimental results show that the proposed KGV framework significantly outperforms the existing GV techniques in terms of both the rate distortion performance and visual quality. Besides, the computational cost of the KGV is rather low at the decoder, making it highly desirable for power-constrained devices. Last but not least, our method can be easily extended to progressive compression with heterogeneous communication network.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 Fall Detection Based on Body Part Tracking Using a Depth Camera
abstract
The elderly population is increasing rapidly all over the world. One major risk for elderly people is fall accidents, especially for those living alone. In this paper, we propose a robust fall detection approach by analyzing the tracked key joints of the human body using a single depth camera. Compared to the rivals that rely on the RGB inputs, the proposed scheme is independent of illumination of the lights and can work even in a dark room. In our scheme, a pose-invariant randomized decision tree algorithm is proposed for the key joint extraction, which requires low computational cost during the training and test. Then, the support vector machine classifier is employed to determine whether a fall motion occurs, whose input is the 3-D trajectory of the head joint. The experimental results demonstrate that the proposed fall detection method is more accurate and robust compared with the state-of-the-art methods.
Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann
IEEE J. Biomed. Health Informatics2
2015 Human Motion Capture Data Tailored Transform Coding
abstract
Human motion capture (mocap) is a widely used technique for digitalizing human movements. With growing usage, compressing mocap data has received increasing attention, since compact data size enables efficient storage and transmission. Our analysis shows that mocap data have some unique characteristics that distinguish themselves from images and videos. Therefore, directly borrowing image or video compression techniques, such as discrete cosine transform, does not work well. In this paper, we propose a novel mocap-tailored transform coding algorithm that takes advantage of these features. Our algorithm segments the input mocap sequences into clips, which are represented in 2D matrices. Then it computes a set of data-dependent orthogonal bases to transform the matrices to frequency domain, in which the transform coefficients have significantly less dependency. Finally, the compression is obtained by entropy coding of the quantized coefficients and the bases. Our method has low computational cost and can be easily extended to compress mocap databases. It also requires neither training nor complicated parameter setting. Experimental results demonstrate that the proposed scheme significantly outperforms state-of-the-art algorithms in terms of compression performance and speed.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Vis. Comput. Graph.1
2015 Motion capture data recovery using skeleton constrained singular value thresholding
Cheen-Hau Tan, Junhui Hou, Lap-Pui Chau
Vis. Comput.2
2014 Low-rank based compact representation of motion capture data
abstract
In this paper, we propose a practical, elegant and effective scheme for compact mocap data representation. Guided by our analysis of the unique properties of mocap data, the input mocap sequence is optimally segmented into a set of subsequences. Then, we project the subsequences onto a pair of computational orthogonal matrices to explore strong low-rank characteristic within and among the subsequences. The experimental results show that the proposed scheme is much more effective for reducing the data size, compared with the existing techniques.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
ICIP1
2014 A fast learning algorithm for multi-layer extreme learning machine
abstract
Extreme learning machine (ELM) is an efficient training algorithm originally proposed for single-hidden layer feedforward networks (SLFNs), of which the input weights are randomly chosen and need not to be fine-tuned. In this paper, we present a new stack architecture for ELM, to further improve the learning accuracy of ELM while maintaining its advantage of training speed. By exploiting the hidden information of ELM random feature space, a recovery-based training model is developed and incorporated into the proposed ELM stack architecture. Experimental results of the MNIST handwriting dataset demonstrate that the proposed algorithm achieves better and much faster convergence than the state-of-the-art ELM and deep learning methods.
Jiexiong Tang, Chenwei Deng, Guang-Bin Huang, Junhui Hou
ICIP4
2014 Restoring corrupted motion capture data via jointly low-rank matrix completion
abstract
Motion capture (mocap) technology is widely used in various applications. The acquired mocap data usually has missing data due to occlusions or ambiguities. Therefore, restoring the missing entries of the mocap data is a fundamental issue in mocap data analysis. Based on jointly low-rank matrix completion, this paper presents a practical and highly efficient algorithm for restoring the missing mocap data. Taking advantage of the unique properties of mocap data (i.e, strong correlation among the data), we represent the corrupted data as two types of matrices, where both the local and global characteristics are taken into consideration. Then we formulate the problem as a convex optimization problem, where the missing data is recovered by solving the two matrices using the alternating direction method of multipliers algorithm. Experimental results demonstrate that the proposed scheme significantly outperforms the state-of-the-art algorithms in terms of both the quality and computational cost.
Junhui Hou, Zhen-Peng Bian, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
ICME1
2014 A novel compression framework for 3D time-varying meshes
abstract
Compression of 3D time-varying meshes (TVMs) plays a critical role in the storage and transmission of 3D contents. In this paper, we propose a novel framework for compressing 3D TVMs. In our framework, 3D TVMs are parameterized and represented by the geometry videos (GVs) through polycube parameterization. By considering the low-rank characteristic of dynamic meshes, we decompose GVs into a sequence with small frames namely EigenGV and the computed reconstruction matrix. We further apply 2D video encoder to eliminate spatial and temporal redundancy among the EigenGV. Experimental results demonstrate that the proposed method significantly outperforms the existing compression schemes in terms of both the rate distortion performance and visual quality. Besides, the proposed method naturally achieves progressive form, which is very suitable for error prone channel transmission.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
ISCAS1
2014 Human Computer Interface for Quadriplegic People Based on Face Position/gesture Detection
abstract
This paper proposes a human computer interface using a single depth camera for quadriplegic people. The nose position is employed to control the cursor along with the commands provided by mouth's status. The detection of nose position and mouth's status is based on randomized decision tree algorithm.The experimental results show that the proposed interface is comfortable, easy to use, robust, and outperforms the existing assistive technology.
Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann
ACM Multimedia2
2014 Scalable and Compact Representation for Motion Capture Data Using Tensor Decomposition
abstract
Motion capture (mocap) technology is widely used in movie and game industries. Compact representation of the mocap data is critical to efficient storage and transmission. In this letter, we propose a novel tensor decomposition based scheme for compact and progressive representation of the mocap data. Our method segments and stacks the mocap sequence locally, and generates a 3rd-order tensor, which has strong correlation within and across slices of the tensor. Then, our method iteratively applies tensor decomposition in a multi-layer structure to explore the correlation characteristic. Experimental results demonstrate that the proposed scheme significantly outperforms existing algorithms in terms of scalability and storage requirement.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Signal Process. Lett.1
2014 A Highly Efficient Compression Framework for Time-Varying 3-D Facial Expressions
abstract
The rapid recent development of 3-DTV technology has led to an increase in studies on mesh-based 3-D scene representation. Compressing 3-D time-varying meshes is critical for the storage and transmission of 3-D contents. This paper proposes a highly efficient framework for compressing time-varying 3-D facial expressions. We use the near-isometric property of human facial expressions to parameterize the 3-D dynamic faces into an expression-invariant 2-D canonical domain that will naturally generate 2-D geometry videos (GVs). Considering the intrinsic properties of GVs, we apply low-rank and sparse matrix decomposition (LRSMD) separately to three dimensions of GVs (namely, \(X, Y,\) and \(Z\) ). Based on our high precision rate and distortion models for GVs, we further compress the components from LRSMD using a video encoder in which bitrates of all components are assigned optimally according to the target bitrate. Experimental results show that the proposed scheme can significantly improve compression performance in terms of rate-distortion performance and visual quality compared with the state-of-the-art algorithms.
Junhui Hou, Lap-Pui Chau, Minqi Zhang, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Human motion capture data recovery via trajectory-based sparse representation
abstract
Motion capture is widely used in sports, entertainment and medical applications. An important issue is to recover motion capture data that has been corrupted by noise and missing data entries during acquisition. In this paper, we propose a new method to recover corrupted motion capture data through trajectory-based sparse representation. The data is firstly represented as trajectories with fixed length and high correlation. Then, based on the sparse representation theory, the original trajectories can be recovered by solving the sparse representation of the incomplete trajectories through the OMP algorithm using a dictionary learned by K-SVD. Experimental results show that the proposed algorithm achieves much better performance, especially when significant portions of data is missing, than the existing algorithms.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Jie Chen 0026, Nadia Magnenat-Thalmann
ICIP1
2013 Expression-invariant and sparse representation for mesh-based compression for 3-D face models
abstract
Compression of mesh-based 3-D models has been an important issue, which ensures efficient storage and transmission. In this paper, we present a very effective compression scheme specifically for expression variation 3-D face models. Firstly, 3-D models are mapped into 2-D parametric domain and corresponded by expression-invariant parameterizaton, leading to 2-D image format representation namely geometry images, which simplifies the 3-D model compression into 2-D image compression. Then, sparse representation with learned dictionaries via K-SVD is applied to each patch from sliced GI so that only few coefficients and their indices are needed to be encoded, leading to low datasize. Experimental results demonstrate that the proposed scheme provides significant improvement in terms of compression performance, especially at low bitrate, compared with existing algorithms.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
VCIP1
2013 Rate-Distortion Model Based Bit Allocation for 3-D Facial Compression Using Geometry Video
abstract
With the extensive applications of 3-D multimedia technology, 3-D content compression has been an important issue, which ensures its smooth transmission on the network with constrained bandwidth. In this letter, we propose a new compression framework for dynamic 3-D facial expressions. Taking advantage of the near-isometric property of human facial expressions, we parameterize the dynamic 3-D faces into an expression-invariant canonical domain, which naturally generates 2-D geometry videos and allows us to apply the well-studied video compression techniques. Due to the difference from natural videos, each dimension (i.e., X, Y and Z, respectively) of the geometry video is regarded as a video sequence and encoded separately. Meanwhile, a model-based joint bit allocation scheme is designed to allocate reasonable bitrate to each dimension by detailed analysis of rate-distortion model for geometry videos, to obtain optimal results under given target bitrate. Experimental results show that up to 25% improvement in terms of bitrate reduction can be achieved, compared to existing algorithms.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Minqi Zhang, Nadia Magnenat-Thalmann
IEEE Trans. Circuits Syst. Video Technol.1