VLDB 2026 Research / reviewers in the wild / expert
Gang He 0002
dblp:96/3183-2
· DBLP profile ↗
48ranked-venue papers
14as first author
22since 2021 · last 2026
0000-0003-0022-4357ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4Systems, architecture and hardware · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation LearningabstractHigh-Dynamic-Range Wide-Color-Gamut (HDR-WCG) technology is becoming increasingly widespread, driving a growing need for converting Standard Dynamic Range (SDR) content to HDR. Existing methods primarily rely on fixed tone mapping operators, which struggle to handle the diverse appearances and degradations commonly present in real-world SDR content. To address this limitation, we propose a generalized SDR-to-HDR framework that enhances robustness by learning attribute-disentangled representations. Central to our approach is Realistic Attribute-Disentangled Representation Learning (RealRep), which explicitly disentangles luminance and chrominance components to capture intrinsic content variations across different SDR distributions. Furthermore, we design a Luma-/Chroma-aware negative exemplar generation strategy that constructs degradation-sensitive contrastive pairs, effectively modeling tone discrepancies across SDR styles. Building on these attribute-level priors, we introduce the Degradation-Domain Aware Controlled Mapping Network (DDACMNet), a lightweight, two-stage framework that performs adaptive hierarchical mapping guided by a control-aware normalization mechanism. DDACMNet dynamically modulates the mapping process via degradation-conditioned features, enabling robust adaptation across diverse degradation domains. Extensive experiments demonstrate that RealRep consistently outperforms state-of-the-art methods in both generalization and perceptually faithful HDR color gamut reconstruction. Li Xu 0008, Kepeng Xu, Lin Zhang 0040, Gang He 0002, Yu-Wing Tai |
AAAI | 5 |
| 2026 | MV-NeRV: Compact neural representation for multi-view videos via coarse-to-fine parallax elimination
Chang Wu 0001, Gang He 0002, Xiandong Meng, Yunsong Li 0001 |
Knowl. Based Syst. | 3 |
| 2026 | Domain adapter for visual object tracking based on hyperspectral video
Langkun Chen, Gang He 0002, Weiying Xie, Yunsong Li 0001 |
Pattern Recognit. | 5 |
| 2026 | Multi-modal object tracking with detailed spatial prompt learning
Gang He 0002, Yuze Ke |
Pattern Recognit. | 1 |
| 2025 | Multi-Frame Deformable Look-Up Table for Compressed Video Quality EnhancementabstractThe rapid progress of multimedia technology has led to an increased focus on enhancing the quality of experience (QoE) for video. Specifically, the demand for low-latency and high-quality decoding has grown significantly. Compressed Video Quality Enhancement (CVQE) methods based on Deep Neural Networks (DNNs) have achieved remarkable success. However, most of the methods suffer from high computational complexity, thereby limiting their practicality in low-latency scenarios. Recently, Look-Up Table (LUT) methods have shown great efficiency, which makes them considerably promising in the field of low-latency CVQE. In this paper, we propose an efficient multi-frame deformable Look-Up Table structure for CVQE. Firstly, we design an efficient CNN to explore the inter-frame correlation and then predict the multi-scale convolution offsets. Secondly, we introduce a temporal feature extraction module and a multi-scale fusion module. We first exploit the predicted offsets to guide sampling for precise temporal alignment and extract multi-frame information. Then, higher quality frames are reconstructed from the fused multi-scale features. During the inference, we convert these two modules into LUTs to achieve a sound trade-off between model performance and computational complexity. Experiments demonstrate that our proposed method dramatically outperforms the state-of-the-art LUT-based methods, and obtains competitive performance compared to CNN-based methods with the capability to run in real-time(30fps) at 1080p resolution. Gang He 0002, Guancheng Quan, Chang Wu 0001, Dajiang Zhou, Yunsong Li 0001 |
AAAI | 1 |
| 2025 | RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancement
Gang He 0002, Guancheng Quan, Dajiang Zhou, Yunsong Li 0001 |
CVPR | 1 |
| 2025 | Beyond Feature Mapping GAP: Integrating Real HDRTV Priors for Superior SDRTV-to-HDRTV ConversionabstractThe rise of HDR-WCG display devices has highlighted the need to convert SDRTV to HDRTV, as most video sources are still in SDR. Existing methods primarily focus on designing neural networks to learn a single-style mapping from SDRTV to HDRTV. However, the limited information in SDRTV and the diversity of styles in real-world conversions render this process an ill-posed problem, thereby constraining the performance and generalization of these methods. Inspired by generative approaches, we propose a novel method for SDRTV to HDRTV conversion guided by real HDRTV priors. Despite the limited information in SDRTV, introducing real HDRTV as reference priors significantly constrains the solution space of the originally high-dimensional ill-posed problem. This shift transforms the task from solving an unreferenced prediction problem to making a referenced selection, thereby markedly enhancing the accuracy and reliability of the conversion process. Specifically, our approach comprises two stages: the first stage employs a Vector Quantized Generative Adversarial Network to capture HDRTV priors, while the second stage matches these priors to the input SDRTV content to recover realistic HDRTV outputs. We evaluate our method on public datasets, demonstrating its effectiveness with significant improvements in both objective and subjective metrics across real and synthetic datasets. Gang He 0002, Kepeng Xu, Li Xu 0008, Wenxin Yu 0001, Xianyun Wu |
IJCAI | 1 |
| 2025 | Unleashing the Potential of Transformer Flow for Photorealistic Face RestorationabstractFace restoration is a challenging task due to the need to remove artifacts and restore details. Traditional methods usually use generative model prior to achieve face restoration, but the restored results are still insufficient in terms of realism and details. In this paper, we introduce OmniFace, a novel face restoration framework that leverages Transformer-based diffusion flow. By exploiting the scaling property of Transformer, OmniFace achieves high-resolution restoration with exceptional realism and detail. The framework integrates three key components: (1) a Transformer-driven vector estimation network, (2) a representation aligned ControlNet, and (3) an adaptive training strategy for face restoration. The inherent scaling law of Transformer architectures enables the restoration of high-quality faces at high resolution. The controlnet combined with pre-trained diffusion representation can be easily trained. The adaptive training strategy provides a vector field that is more suitable for face restoration. Comprehensive experiments demonstrate that OmniFace outperforms existing techniques in terms of restoration quality across multiple benchmark datasets, especially in restoring photographic-level texture details in high-resolution scenes. Kepeng Xu, Li Xu 0008, Gang He 0002, Wei Chen 0062, Xianyun Wu, Wenxin Yu 0001 |
IJCAI | 3 |
| 2025 | RGB-D visual object tracking with transformer-based multi-modal feature fusion
Yuze Ke, Wanlin Zhao, Gang He 0002, Yunsong Li 0001 |
Knowl. Based Syst. | 6 |
| 2025 | SA-CVSR: Scale-Arbitrary Compressed Video Super-Resolution
Gang He 0002, Chang Wu 0001, Guancheng Quan, Yunsong Li 0001 |
Pattern Recognit. | 1 |
| 2025 | Hyperspectral Object Tracking With Spectral Information PromptabstractHyperspectral videos contain a larger number of spectral bands, providing extensive spectral information and material identification capabilities. This advantage confers hyperspectral trackers to achieve superior performance in challenging tracking scenarios. However, the limited availability of hyperspectral training data and the inability of existing algorithms to fully exploit hyperspectral information restrict the tracking performance. To address this issue, a novel framework, Spectral Prompt-based Hyperspectral Object Tracking (SP-HST), is proposed. SP-HST leverages a RGB tracking network as the main branch for feature extraction and tracking, which accounts for more than 98% of the total parameters and remains frozen during the training procedure. Additionally, the Spectral Prompt Learning (SPL) branch, comprising multiple lightweight prompt blocks, is introduced to generate complementary spectral representations as the prompt. The prompts contain abundant spectral information from hyperspectral data, enhancing the discriminative ability of features within the main branch. Furthermore, the Complementary Weight Learning (CWL) is employed to calculate the importance of spectral information from different prompts, enabling the features for hyperspectral object tracking to contain more spectral information that is absent in the feature of the main branch. By utilizing the spectral information as prompt, the number of trainable parameters is less than 2% of that in the tracking network, and the convergence is reached in 12 training epoch. Extensive experiments demonstrate the superiority of SP-HST, achieving a new state-of-the-art tracking performance, 71.3% of the AUC score on the HOTC dataset and 96.7% of the DP@20P score on the IMEC25 dataset. The code will be released at https://github.com/lgao001/SP-HST. Gang He 0002, Langkun Chen, Weiying Xie, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
Kepeng Xu, Li Xu 0008, Gang He 0002, Wenxin Yu 0001, Yunsong Li 0001 |
IJCAI | 3 |
| 2024 | QS-NeRV: Real-Time Quality-Scalable Decoding with Neural Representation for VideosabstractIn this paper, we propose a neural representation for videos that enables real-time quality-scalable decoding, called QS-NeRV. QS-NeRV comprises a Self-Learning Distribution Mapping Network (SDMN) and Extensible Enhancement Networks (EENs). Firstly, SDMN functions as the base layer (BL) for scalable video coding, focusing on encoding videos of lower quality. Within SDMN, we employ a methodology that minimizes the bitstream overhead to achieve efficient information exchange between the encoder and decoder instead of direct transmission. Specifically, we utilize an invertible network to map the multi-scale information obtained from the encoder to a specific distribution. Subsequently, during the decoding process, this information is recovered from a randomly sampled latent variable to assist the decoder in achieving improved reconstruction performance. Secondly, EENs serve as the enhancement layers (ELs) and are trained in an overfitting manner to obtain robust restoration capability. By integrating the fixed BL bitstream with the parameters of EEN as an extension pack, the decoder can produce higher-quality enhanced videos. Furthermore, the scalability of the method allows for adjusting the number of combined packs to accommodate diverse quality requirements. Experimental results demonstrate our proposed QS-NeRV outperforms the state-of-the-art real-time decoding INR-based methods on various datasets for video compression and interpolation tasks. Chang Wu 0001, Guancheng Quan, Gang He 0002, Yunsong Li 0001, Wenxin Yu 0001, Xianmeng Lin, Cheng Yang 0016 |
ACM Multimedia | 3 |
| 2024 | An End-to-End Real-World Camera Imaging Pipelineabstractpipeline still faces challenges including the lack of joint optimization in system components, computational redundancies, and optical distortions such as lens shading.In light of this, we propose an end-to-end camera imaging pipeline (RealCamNet) to enhance realworld camera imaging performance.Our methodology diverges from conventional, fragmented multi-stage image signal processing towards end-to-end architecture.This architecture facilitates joint optimization across the full pipeline and the restoration of coordinate-biased distortions.RealCamNet is designed for highquality conversion from RAW to RGB and compact image compression.Specifically, we deeply analyze coordinate-dependent optical distortions, e.g., vignetting and dark shading, and design a novel 2804 Kepeng Xu, Zijia Ma, Li Xu 0008, Gang He 0002, Yunsong Li 0001, Wenxin Yu 0001, Taichu Han, Cheng Yang 0016 |
ACM Multimedia | 4 |
| 2024 | PMCN: Parallax-motion collaboration network for stereo video dehazing
Chang Wu 0001, Gang He 0002, Wanlin Zhao, Yunsong Li 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Local-Global Self-Attention for Transformer-Based Object TrackingabstractTransformer-based tracking methods have been widely studied in the field of visual object tracking. The long-range information capturing ability of the transformer improves the performance of the tracking network. However, the self-attention learning procedure in the transformer module neglects the local information, the target and the background around it, which can be beneficial for trackers to handle background clutter and deformation. In this paper, the local-global self-attention (LGSA) learning is proposed for the object tracking task, which obtains the local and global information simultaneously in one attention learning block. Based on the LGSA, the encoder and the decoder are designed to fuse the features corresponding to the template and search images. Additionally, two tracking networks, LGSAT-T and LGSAT-B instantiated with the proposed encoder and decoder are introduced. Exclusive experiments on the commonly used datasets, including OTB100, GOT-10K, LaSOT, and TrackingNet, demonstrate the effectiveness of LGSA, and indicate the state-of-the-art performance of the proposed tracking network. The code will be released athttps://github.com/lgao001/LGSAT. Langkun Chen, Yunsong Li 0001, Gang He 0002, Jifeng Ning |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Global priors guided modulation network for joint super-resolution and SDRTV-to-HDRTVabstractWatching low resolution standard dynamic range (LR SDR) video on a 4K high dynamic range (HDR) TV is not the best viewing experience. Joint super-resolution (SR) and SDRTV-to-HDRTV aims to enhance the visual quality of LR SDR videos that have quality deficiencies in resolution and dynamic range. Previous methods that rely on learning local information typically cannot do well in preserving color conformity and long-range structural similarity, resulting in unnatural color transition and texture artifacts. In order to tackle these challenges, we propose a global priors guided modulation network (GPGMNet). In particular, we design a global priors extraction module (GPEM) to extract color conformity prior and structural similarity prior that are beneficial for SDRTV-to-HDRTV and SR tasks, respectively. To further exploit the global priors and preserve spatial information, we devise multiple global priors-guided spatial-wise modulation blocks (GSMBs) with a few parameters for intermediate feature modulation. In these GSMBs, the modulation parameters are generated by the shared global priors and the spatial features map from the spatial pyramid convolution block (SPCB). With these elaborate designs, the GPGMNet can achieve higher visual quality with lower computational complexity. Extensive experiments demonstrate that our proposed GPGMNet is superior to the state-of-the-art methods. Specifically, our proposed model exceeds the state-of-the-art by 0.64 dB in PSNR, with 69% fewer parameters and 3.1 × speedup. Gang He 0002, Shaoyi Long, Li Xu 0008, Chang Wu 0001, Wenxin Yu 0001, Jinjia Zhou |
Neurocomputing | 1 |
| 2023 | MPCNet: Compressed multi-view video restoration via motion-parallax complementation network
Chang Wu 0001, Gang He 0002, Yunsong Li 0001 |
Neural Networks | 2 |
| 2022 | Transcoded Video Restoration by Temporal Spatial Auxiliary NetworkabstractIn most video platforms, such as Youtube, Kwai, and TikTok, the played videos usually have undergone multiple video encodings such as hardware encoding by recording devices, software encoding by video editing apps, and single/multiple video transcoding by video application servers. Previous works in compressed video restoration typically assume the compression artifacts are caused by one-time encoding. Thus, the derived solution usually does not work very well in practice. In this paper, we propose a new method, temporal spatial auxiliary network (TSAN), for transcoded video restoration. Our method considers the unique traits between video encoding and transcoding, and we consider the initial shallow encoded videos as the intermediate labels to assist the network to conduct self-supervised attention training. In addition, we employ adjacent multi-frame information and propose the temporal deformable alignment and pyramidal spatial fusion for transcoded video restoration. The experimental results demonstrate that the performance of the proposed method is superior to that of the previous techniques. The code is available at https://github.com/icecherylXuli/TSAN. Li Xu 0008, Gang He 0002, Jinjia Zhou, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yu-Wing Tai |
AAAI | 2 |
| 2022 | SDRTV-to-HDRTV via Hierarchical Dynamic Context Feature MappingabstractIn this work, we address the task of SDR videos to HDR videos(SDRTV-to-HDRTV conversion). Previous approaches use global feature modulation for SDRTV-to-HDRTV conversion. Feature modulation scales and shifts the features in the original feature space, which has limited mapping capability. In addition, the global image mapping cannot restore detail in HDR frames due to the luminance differences in different regions of SDR frames. To resolve the appeal, we propose a two-stage solution. The first stage is a hierarchical Dynamic Context feature mapping (HDCFM) model. HDCFM learns the SDR frame to HDR frame mapping function via hierarchical feature modulation (HME and HM ) module and a dynamic context feature transformation (DYCT) module. The HME estimates the feature modulation vector, HM is capable of hierarchical feature modulation, consisting of global feature modulation in series with local feature modulation, and is capable of adaptive mapping of local image features. The DYCT module constructs a feature transformation module in conjunction with the context, which is capable of adaptively generating a feature transformation matrix for feature mapping. Compared with simple feature scaling and shifting, the DYCT module can map features into a new feature space and thus has a more excellent feature mapping capability. In the second stage, we introduce a patch discriminator-based context generation model PDCG to obtain subjective quality enhancement of over-exposed regions. The proposed method can achieve state-of-the-art objective and subjective quality results. Specifically, HDCFM achieves a PSNR gain of 0.81 dB at about 100K parameters. The number of parameters is 1/14th of the previous state-of-the-art methods. The test code will be released on https://github.com/cooperlike/HDCFM. Gang He 0002, Kepeng Xu, Li Xu 0008, Chang Wu 0001, Ming Sun 0008, Yu-Wing Tai |
ACM Multimedia | 1 |
| 2022 | Interlayer Restoration Deep Neural Network for Scalable High Efficiency Video CodingabstractThis paper applies an interlayer restoration deep neural network (IRDNN) for scalable high efficiency video coding (SHVC) to improve visual quality and coding efficiency. It is the first time to combine deep neural network (DNN) and SHVC. Considering the coding architecture of SHVC, we elaborate a multi-frame and multi-layer neural network to restore the interlayer of SHVC by utilizing both the adjacent reconstructed frames of the base layer (BL) and enhancement layer (EL). Moreover, we analyze the temporal motion relationship of frames in one layer and the compression degradation relationship of frames between different layers, and propose the synergistic mechanism of motion restoration and compression restoration in our IRDNN. The network can generate an interlayer with higher quality serving for the EL coding and thus enhance the coding efficiency. A large-scale and various-quality-degradation dataset is self-made for the task of interlayer restoration of SHVC. The experimental results show that with our implementation on SHVC, the EL Bj$\phi $ntegaard delta bit-rate (BD-BR) reduction is 9.291% and 6.007% in signal-to-noise ratio scalability and spatial scalability, respectively. The code is available athttps://github.com/icecherylXuli/IRDNN. Gang He 0002, Li Xu 0008, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yibo Fan, Jinjia Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | RR-DnCNN v2.0: Enhanced Restoration-Reconstruction Deep Neural Network for Down-Sampling-Based Video CodingabstractIntegrating deep learning techniques into the video coding framework gains significant improvement compared to the standard compression techniques, especially applying super-resolution (up-sampling) to down-sampling based video coding as post-processing. However, besides up-sampling degradation, the various artifacts brought from compression make super-resolution problem more difficult to solve. The straightforward solution is to integrate the artifact removal techniques before super-resolution. However, some helpful features may be removed together, degrading the super-resolution performance. To address this problem, we proposed an end-to-end restoration-reconstruction deep neural network (RR-DnCNN) using the degradation-aware technique, which entirely solves degradation from compression and sub-sampling. Besides, we proved that the compression degradation produced by Random Access configuration is rich enough to cover other degradation types, such as Low Delay P and All Intra, for training. Since the straightforward network RR-DnCNN with many layers as a chain has poor learning capability suffering from the gradient vanishing problem, we redesign the network architecture to let reconstruction leverages the captured features from restoration using up-sampling skip connections. Our novel architecture is called restoration-reconstruction u-shaped deep neural network (RR-DnCNN v2.0). As a result, our RR-DnCNN v2.0 outperforms the previous works and can attain 17.02% BD-rate reduction on UHD resolution for all-intra anchored by the standard H.265/HEVC. The source code is available at https://minhmanho.github.io/rrdncnn/. Man M. Ho, Jinjia Zhou, Gang He 0002 |
IEEE Trans. Image Process. | 3 |
| 2020 | Interactive Separation Network For Image InpaintingabstractImage inpainting, also known as image completion, is the process of filling in the missing region of an incomplete image to make the repaired image visually plausible. Strided convolutional layer learns high-level representations while reducing the computational complexity, but fails to preserve existing detail from the original images (eg, texture, sharp transients), therefore it degrades the generative model in image inpainting task. To reduce the erosion of high-resolution components of images meanwhile maintaining the semantic representation, this paper designs a brand-new network called Interactive Separation Network that progressively decomposites the features into two streams and fuses them. Besides, the rationality of network design and the efficiency of proposed network is demonstrated in the ablation study. To the best of our knowledge, the experimental results of proposed method are superior to state-of-the-art inpainting approaches. Siyuan Li 0004, Xin Cheng 0004, Kepeng Xu, Wenxin Yu 0001, Gang He 0002, Jinjia Zhou |
ICIP | 7 |
| 2020 | Coarse-to-Fine Attention Network via Opinion Approximate Representation for Aspect-Level Sentiment Classification
Wei Chen 0062, Wenxin Yu 0001, Gang He 0002, Ning Jiang 0002, Gang He 0001 |
ICONIP (1) | 3 |
| 2020 | Customizable GAN: Customizable Image Synthesis Based on Adversarial Learning
Wenxin Yu 0001, Jinjia Zhou, Xuewen Zhang, Jialiang Tang, Siyuan Li 0004, Ning Jiang 0002, Gang He 0001, Gang He 0002 |
ICONIP (4) | 9 |
| 2020 | No-Reference Quality Assessment Based on Spatial Statistic for Generated Images
Yunye Zhang, Xuewen Zhang, Wenxin Yu 0001, Ning Jiang 0002, Gang He 0002 |
ICONIP (4) | 6 |
| 2020 | Down-Sampling Based Video Coding with Degradation-Aware Restoration-Reconstruction Deep Neural Network
Minh-Man Ho, Gang He 0002, Jinjia Zhou |
MMM (1) | 2 |
| 2020 | Hyperspectral Image Super-Resolution via Intrafusion NetworkabstractThis article presents an intrafusion network (IFN) for hyperspectral image (HSI) super-resolution (SR). Given that the HSI is a 3-D data cube with both the spatial information and the spectral information, the key challenge to construct HSI SR is how to efficiently exploit the spectral information among consecutive low-resolution (LR) bands, besides the spatial information. The proposed IFN consists of three modules, including the spectral difference module, the parallel convolution module, and the intrafusion module, which directly utilizes both the spatial information and the spectral information for reconstructing the high-resolution HSI. Different from most of the existed methods that tackle the spatial and spectral information separately, the proposed spatial-spectral utilization is achieved in one integrated network, which opens up a new way for HSI SR. Meanwhile, applications of this three modules strategy (first spectral difference, then parallel convolution, and finally, intrafusion) on both the conventional convolutional neural network and the residual network with deeper depth have shown the generalization capacity of this proposal. Experimental results and data analysis demonstrate the effectiveness of the proposed method using three hyperspectral data sets. Jing Hu 0005, Xiuping Jia, Yunsong Li 0001, Gang He 0002, Minghua Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Semisupervised Spectral Learning With Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractLimited by the anomalous spectral vectors in unlabeled hyperspectral images (HSIs), anomaly detection methods based on background distribution estimation often suffer from the contamination of anomalies, which decreases the estimation accuracy and, thus, weakens the detection performance. To address this problem, we proposed a novel semisupervised spectral learning (SSL) for the hyperspectral anomaly detection framework based on the generative adversarial network (GAN). GAN is applied and developed to estimate the background distribution in a semisupervised manner and obtain an initial spectral feature because of its strong representational capability and adversarial training advantage. In the proposed framework, an initial spatial feature is generated via morphological attribute filtering. Finally, an exponential constrained nonlinear suppression fusion technique is adopted to suppress the background and combine the complementary information in different features to obtain a fused detection map. The performance of the proposed anomaly detection technique is evaluated on a series of HSIs. Experimental results demonstrate that our method can outperform state-of-the-art anomaly detection methods. Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Gang He 0002, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Spectral Adversarial Feature Learning for Anomaly Detection in Hyperspectral ImageryabstractTheoretically, hyperspectral images (HSIs) are capable of providing subtle spectral differences between different materials, but in fact, it is difficult to distinguish between background and anomalies because the samples of anomalous pixels in HSIs are limited and susceptible to background and noise. To explore the discriminant features, a spectral adversarial feature learning (SAFL) architecture is specially designed for hyperspectral anomaly detection in this article. In addition to reconstruction loss, SAFL also introduces spectral constraint loss and adversarial loss in the network with batch normalization to extract the intrinsic spectral features in deep latent space. To further reduce the false alarm rate, we present an iterative optimization approach by a weighted suppression function that depends on the contribution rate of each feature to the detection. In particular, the structure tensor matrix is adopted to adaptively calculate the contribution rate of each feature. Benefiting from these improvements, the proposed method is superior to the typical and state-of-the-art methods either in detection probability or false alarm rate. Weiying Xie, Baozhu Liu, Yunsong Li 0001, Jie Lei 0001, Chein-I Chang, Gang He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | SRUN: Spectral Regularized Unsupervised Networks for Hyperspectral Target DetectionabstractThe high dimensionality of a hyperspectral image (HSI) provides the possibility of deeply capturing the underlying and intrinsic characteristics in spectra, such that targets embedded in the background can be detected. However, redundant information, deteriorated bands, and other interferences from background challenge the target detection problem. In this article, an effective feature extraction method based on unsupervised networks is proposed to mine intrinsic properties underlying HSIs. Our approach, called spectral regularized unsupervised networks (SRUN), imposes spectral regularization on autoencoder (AE) and variational AE (VAE) to emphasize spectral consistency, which is more suitable for characterizing spectral information of HSIs by hidden nodes than the original AE and VAE models. Then, we conduct a simple feature selection algorithm on the hidden nodes in the deepest code to select specific nodes that contain distinguishability between target and background, which is based on the spectral angular difference between a known target spectrum and spectra of other pixels in input. The selected nodes are further weighted adaptively to obtain a discriminative map depending on the observation that each selected node provides different contribution rates to target detection. Experimental results on several data sets illustrate that the proposed SRUN-based target detection algorithm is suitable for targets at the subpixel level and those with structural information. Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001, Gang He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Text to Image Synthesis Based on Multiple Discrimination
Yunye Zhang, Wenxin Yu 0001, Jingwei Lu, Li Nie, Gang He 0001, Ning Jiang 0002, Gang He 0002, Yibo Fan |
ICANN (3) | 8 |
| 2019 | Target-Based Attention Model for Aspect-Level Sentiment Analysis
Wei Chen 0062, Wenxin Yu 0001, Yunye Zhang, Kepeng Xu, Fengwei Zhang, Yibo Fan, Gang He 0002 |
ICONIP (3) | 8 |
| 2019 | Inpainting with Sketch Reconstruction and Comprehensive Feature Selection
Siyuan Li 0004, Zhijing Li 0003, Kepeng Xu, Matthieu Claisse, Wenxin Yu 0001, Gang He 0001, Gang He 0002, Yibo Fan |
ICONIP (5) | 8 |
| 2019 | Dense Image Captioning Based on Precise Feature Extraction
Yunye Zhang, Wenxin Yu 0001, Li Nie, Gang He 0002, Yibo Fan |
ICONIP (5) | 6 |
| 2019 | Text to Image Synthesis Using Two-Stage Generation and Two-Stage Discrimination
Yunye Zhang, Wenxin Yu 0001, Gang He 0001, Ning Jiang 0002, Gang He 0002, Yibo Fan |
KSEM (2) | 6 |
| 2018 | Q Value-Based Dynamic Programming with Boltzmann Distribution by Using Neural Network
Wenxin Yu 0001, Gang He 0001, Yibo Fan, Gang He 0002, Jiu Xu |
ICONIP (7) | 5 |
| 2017 | Fast mode decision and PU size decision algorithm for intra depth coding in 3D-HEVC
Gang He 0002, Jing Hu 0005, Yunsong Li 0001, Wenxin Yu 0001, Peikun Liu, Ruixue Guo |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Fast algorithm for prediction unit and mode decisions of intra depth coding in 3D-HEVCabstractAs the state-of-the-art video coding standard for 3D video, the 3D video extension of High Efficiency Video Coding (3D-HEVC) compresses the multi-view texture videos plus depth maps. The intra depth coding consumes huge computational complexity due to the added depth modeling modes (DMMs) and its new complex processing flow. This paper proposes a fast algorithm to reduce the complexity for prediction unit (PU) and mode decisions for intra depth coding. Firstly, the early PU splitting and pruning methods are proposed to fast decide the PU size, based on the intra depth coding flow. Secondly, by analyzing the relationship between DMMs and Planar mode, a fast algorithm is used to skip the mode decision under the certain condition. Experimental results show our proposed methods together reduce 56.32% and 50.12% computational complexity for depth map and total video coding, while the performance loss is only 1.42% BD-rate increasing. Ruixue Guo, Gang He 0002, Yunsong Li 0001 |
ICIP | 2 |
| 2016 | Fast algorithm based on sole- and multi-depth measurements for HEVC intra codingabstractIn High Efficiency Video Coding (HEVC), intra coding plays an important role, but also involves huge computational complexity due to a flexible coding unit (CU) structure and a large number of prediction modes. This paper presents a fast algorithm based on the sole- and multi-depth measurements to reduce the complexity from CU and prediction mode decisions. For the CU decision, evaluation results with sole and multiple depths are utilized to judge if the CU is a heterogeneous, homogeneous, or depth prominent one, where fast CU decisions are made. For the prediction mode decision, the tendencies for different CU sizes are detected based on multiple depths. The number of searching modes is decreased adaptively for the depth with fewer tendencies. Experimental results show the proposed algorithm reduces 61.49% computational complexity, with 0.75% bit-rate increasing, which is more efficient than state-of-the-arts. Gang He 0002, Jing Hu 0005, Yunsong Li 0001, Wenxin Yu 0001 |
ICIP | 1 |
| 2016 | A fast mode selection for depth modelling modes of intra depth coding in 3D-HEVCabstractThe 3D extension of the high efficiency video coding (3D-HEVC) standard adopts new depth modelling modes (DMMs) to provide more accurate prediction for depth map intra coding, while the mode selection for DMMs causes huge computational complexity. In this paper, we develop the fast algorithm for DMMs selection to reduce the complexity. Firstly, the evaluation results of intra conventional modes are utilized to determine whether DMMs should be skipped. Secondly, golden ratio is adopted to simplify DMMs searching. Results show that golden ratio can reduce 70.44% time for DMMs searching. Experimental results show that our algorithm reduces 37.40% encoding time on average with only 0.40% increase on synthesized BD-rate. Peikun Liu, Gang He 0002, Shun Xue, Yunsong Li 0001 |
VCIP | 2 |
| 2016 | Fast algorithm based on the sole- and multi-depth texture measurements for HEVC intra coding
Jing Hu 0005, Gang He 0002, Yunsong Li 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | High-Throughput Power-Efficient VLSI Architecture of Fractional Motion Estimation for Ultra-HD HEVC Video EncodingabstractFractional motion estimation (FME) significantly enhances video compression efficiency, but its high computational complexity also limits the real-time processing capability. In this brief, we present a VLSI implementation of FME design in High Efficiency Video Coding for ultrahigh definition video applications. We first propose a bilinear quarter pixel approximation, together with a search pattern based on it to reduce the complexity of interpolation and fractional search process. Furthermore, a data reuse strategy is exploited to reduce the hardware cost of transform. In addition, using the considered pixel parallelism and dedicated access pattern for memory, we fully pipeline the computation and achieve high hardware utilization. This design has been implemented as a 65-nm CMOS chip and verified. The measured throughput reaches 995 Mpixels/s for 7680 × 4320 30 frames/s at 188 MHz, at least 4.7 times faster than prior arts. The corresponding power dissipation is 198.6 mW, with a power efficiency of 0.2 nJ/pixel. Due to the optimization, our work achieves more than 52% improvement on power efficiency, relative to previous works in H.264. Gang He 0002, Dajiang Zhou, Yunsong Li 0001, Zhixiang Chen 0002, Tianruo Zhang, Satoshi Goto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | High-Performance H.264/AVC Intra-Prediction Architecture for Ultra High Definition Video ApplicationsabstractThis paper presents an H.264/AVC intra-prediction design for ultrahigh definition (ultra-HD) video. Due to the huge throughput requirements of ultra-HD, design challenges such as complexity and data dependency, which currently exist for lower resolutions, become even more critical. To solve these problems, we first propose an interlaced block reordering scheme together with a preliminary mode decision (PMD) strategy to resolve the data dependency between intra mode decision and reconstruction. In the meantime, hardware cost is reduced by PMD. We also propose a probability-based reconstruction scheme to solve the problem of long pipeline latency. In addition, hardware reuse strategies including a shared fine decision module and processing element-reusable prediction generator, are applied to further optimize the design. As a result, the hardware complexity is reduced by 77% in terms of area and frequency, and it takes an average of 33 cycles to process a macroblock. The implementation result demonstrates that our design can support up to the specification of 7680 × 4320p 60 f/s when running at 273 MHz. The design is implemented with 451.5 k gates in 65-nm CMOS. Gang He 0002, Dajiang Zhou, Wei Fei, Zhixiang Chen 0002, Jinjia Zhou, Satoshi Goto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | A 24.5-53.6pJ/pixel 4320p 60fps H.264/AVC intra-frame video encoder chip in 65nm CMOSabstractAn H.264/AVC intra-frame video encoder is implemented in 65nm CMOS. With an efficient intra prediction design, its maximum throughput reaches 1991Mpixels/s for 7680×4320p 60fps video, 9.4x to 32x faster than previous designs. The encoder also incorporates a 1.41Gbins/s CABAC architecture that has been enhanced by 31%. Moreover, low energy consumption is achieved by the high parallelism and hardware efficiency of this design. 1080p 30fps encoding dissipates only 2mW at 0.8V and 9MHz. Dajiang Zhou, Gang He 0002, Wei Fei, Zhixiang Chen 0002, Jinjia Zhou, Satoshi Goto |
ASP-DAC | 2 |
| 2013 | Combined hole-filling with spatial and temporal predictionabstractA combined hole-filling approach with spatial and temporal prediction is presented in this paper. Depth image-based rendering (DIBR) is generally used to synthesize virtual view images in free viewpoint television (FTV) and three-dimensional (3-D) video. Limited original camera views and depth maps are used to generate the additional virtual views in the synthesizing process. One of the main problems in DIBR is that there are some regions are occluded by the foreground objects in the original views, and they will be some holes in the generated additional virtual views, especially for the view extrapolation cases. Therefore, the proposed algorithm is introduced and it can be used to fill the holes which caused by disocclusion regions and inaccurate depth values. The proposed algorithm combines the spatial and temporal prediction, and the performance is much better and more stable than the previous work. The experimental results show that the proposed method can improve the quality of the virtual views a lot compared with the previous work. The improvement is not only obvious in the objective comparison, but also in the subjective comparison. Wenxin Yu 0001, Weichen Wang 0005, Gang He 0002, Satoshi Goto |
ICIP | 3 |
| 2013 | A combined SAO and de-blocking filter architecture for HEVC video decoderabstractThe up-coming video compression standard, high efficiency video coding (HEVC), reduces 50% bit rates in encoding video sequences with same picture quality compared to H.264/AVC. In the in-loop filter (LF) part of HEVC, sample adaptive offset (SAO) is newly added and de-blocking filter (DBF) has been changed a lot. Thus how to construct a high speed and low cost VLSI architecture for HEVC SAO and de-blocking filter is a challenge. In this article, we propose a HEVC LF architecture composed of fully utilized de-blocking filter and SAO. Block based SAO and DBF are employed in this architecture to achieve seamless pipeline between them. The implementation results show that it can be synthesized to 240MHz with 65nm technology. Thus this solution can process 3.84G pixels/s and support 4320p(7680×4320)@120fps decoding. Jiayi Zhu 0001, Dajiang Zhou, Gang He 0002, Satoshi Goto |
ICIP | 3 |
| 2010 | Intra prediction architecture for H.264/AVC QFHD encoderabstractThis paper proposes a high-performance intra prediction architecture that can support H.264/AVC high profile. The proposed MB/block co-reordering can avoid data dependency and improve pipeline utilization. Therefore, the timing constraint of real-time 4k×2k encoding can be achieved with negligible quality loss. 16×16 prediction engine and 8×8 prediction engine work parallel for prediction and coefficients generating. A reordering interlaced reconstruction is also designed for fully pipelined architecture. It takes only 160 cycles to process one macroblock (MB). Hardware utilization of prediction and reconstruction modules is almost 100%. Furthermore, PE-reusable 8×8 intra predictor and hybrid SAD & SATD mode decision are proposed to save hardware cost. The design is implemented by 90nm CMOS technology with 113.2k gates and can encode 4k×2k video sequences at 60 fps with operation frequency of 310MHz. Gang He 0002, Dajiang Zhou, Jinjia Zhou, Satoshi Goto |
PCS | 1 |