EDBT 2026 Demo / reviewers in the wild / expert
Junjie Huang 0001
dblp:85/774-1 · also Jun-Jie Huang 0001
· DBLP profile ↗
30ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0003-2986-4665ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic DomainabstractCloud-based third-party multimedia services have become increasingly popular in last decade, however, they pose serious threats to users' privacy. To address this issue, in this paper, we propose a novel Adaptive Image Restoration network with Privacy protection, namely AIRPNet, which first attempts to perform image restoration in steganographic domain. Compared with existing methods, our method has significant advantages in invisibility, security and flexibility. Specifically, we first propose a wavelet lifting-based Adaptive Invertible Hiding (AIH) module to conceal the low-quality (LQ) secret image into a stego image. Then, instead of performing single type of restoration on the secret image, an adaptive secure restoration (ASR) module is developed to deal with multiple image degradations on the stego image. Finally, a high-quality (HQ) secret image can be extracted from the restored stego image. Here, since the secret image remains hidden throughout the whole image restoration process, the privacy of users can be greatly protected. The framework can be flexibly extended to multiple image restoration, which can restore multiple secret images from the same stego image. Experimental results on various datasets demonstrate that our AIRPNet outperforms existing methods in terms of restoration accuracy, invisibility and security on different image restoration tasks. Fangyuan Gao, Xin Deng 0002, Junjie Huang 0001, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Contrastive and Dual Adversarial Representation Learning for Multi-View ClusteringabstractMulti-View Clustering (MVC) has gained increasing attention due to its ability to effectively leverage the complementary information of multi-view data. Despite the success of existing MVC methods in many real-world applications, they often overlook the discrepancy of view-specific latent distribution and struggle to ensure the completeness of the multi-view data. To address these challenges and harness the powerful feature extraction capability of deep networks, we propose a novel Contrastive and Dual Adversarial Representation Learning method for Multi-view Clustering, termed as CDARL, to solve multi-view clustering problems with both complete and incomplete multi-view data. Specifically, CDARL employs alternating adversarial and contrastive learning to align the view-specific representations, driving them into the same semantic latent space to minimize the discrepancy in view-specific distributions. In addition, a consensus latent representation is learned by an adaptive fusion block that integrates information from multiple views. The consensus representation is further refined through adversarial learning modeling the transformation of the standard Gaussian distribution to the original data distribution. Moreover, the proposed method incorporates an imputation strategy designed to handle the incomplete multi-view data clustering task. This strategy utilizes both reconstructed samples and cross-view neighbors to impute missing views from the latent space and the original space, thereby preserving clustering information, which ensures the quality and feasibility of the imputed samples. Experimental results on six widely used datasets have verified the competitiveness of the proposed CDARL method against state-of-the-art methods in MVC problems with complete and incomplete multi-view data. Code is available athttps://github.com/xywy220/CDARL-MVC. Yanwanyu Xi, Chang Tang, Junjie Huang 0001, Xingchen Hu 0001, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | MM4flow: A Pre-trained Multi-modal Model for Versatile Network Traffic AnalysisabstractNetwork traffic analysis is a critical research area, playing an essential role in enhancing network security and ensuring high-quality network services. Existing methods, which primarily rely on a single modality, face two significant limitations. First, while existing approaches may achieve strong performance in specific tasks, they often lack sufficient adaptability for diverse tasks. Second, existing pre-trained models are only trained with GB-scale traffic, with which increases the risk of over-fitting and limiting the models' overall performance. To address these challenges, we propose MM4flow, a pre-trained multi-modal model designed for versatile network traffic analysis. We divide network flows into two modalities: raw byte streams and transmission patterns, which encapsulate the content and behavior information, respectively. MM4flow is composed of two key stages: uni-modal pre-training and multi-modal fine-tuning. We develop an efficient data collection scheme enabling TB-scale traffic pre-training. Leveraging a real-world traffic that exceeds 70 TB, MM4flow conducts uni-modal pre-training on each modality with a modified BERT architecture tailored for network flows. For specific downstream tasks, we introduce a modal fusion module based on cross-attention mechanisms. The fusion module facilitates effective integration of multi-modal information, enabling MM4flow to fully utilize both content and behavior cues during fine-tuning with minimal labeled dataset. We evaluate MM4flow on six public datasets covering six various tasks. Extensive experiments demonstrate that MM4flow achieves superior accuracy than baselines. Especially, compared to existing pre-trained models, MM4flow achieves an 84% improvement in accuracy for website identification under encrypted tunnels. Moreover, the pre-trained MM4flow significantly reduces the reliance on high-quality labeled training data for downstream tasks. Luming Yang, Lin Liu 0018, Junjie Huang 0001, Zhuotao Liu, Shiyu Liang, Shaojing Fu |
CCS | 3 |
| 2025 | AIM-VR: All-in-One Video Restoration via Dual-Path Mamba with Frequency Adaptive FusionabstractReal-world video-based vision systems frequently suffer concurrent degradations caused by unpredictable weather conditions such as rain, haze, and snow, severely affecting the visual quality of both human observers and the downstream computer vision tasks. In this paper, we propose All In Mamba Video Restoration (AIM-VR) model, a novel multi-degradation video restoration framework based on the Selective State Space Model. The proposed AIM-VR model effectively handles adverse weather conditions through three key innovations: a Dual-Path Mamba Modeling (DPMM) backbone with complementary Space-Time Sequence Mamba Block (SSMB) and Hilbert Sequence Mamba block (HSMB) for efficient temporal modeling, a Frequency Adaptive Fusion block (FAFB) for degradation-specific feature modulation, and a Universal Multi-degradation Contrastive Learning (UMCL) strategy for robust pattern discovery in multi-degradation scenarios. Experimental results demonstrate that the proposed AIM-VR achieves superior performance in terms of both restoration quality (0.82 dB PSNR over the state-of-the-art methods) and computational efficiency across multiple weather scenarios. Code is available at https://github.com/StephenLockhart/AIM-VR. Zhizhou Lu, Junjie Huang 0001, Xueqiong Li, Baili Xiao |
ICME | 4 |
| 2025 | VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal ModelingabstractThe exponential growth of video content necessitates efficient summarization techniques that balance local redundancy reduction and global dependency modeling. In this work, we introduce VSumMamba, an innovative video summarization approach that leverages Selective State Space Models to address the quadratic complexity limitations of Transformer based approaches meanwhile surpassing CNNs' restricted long-range modeling capabilities. The proposed framework comprises three core components: 1) a Multi-Scale Aggregator, 2) a Cascaded Temporal Modeling Module with bi-directional Mamba blocks for temporal representation enhancement, and 3) a Parallel Spatial Modeling Module employing spatial Mamba blocks, operating in concert to effectively refine spatiotemporal video representations. Through three specialized multi-scale spatial-temporal modeling schemes, VSumMamba demonstrate the ability to balance computational efficiency and summarization performance. Comprehensive evaluations on benchmarks datasets demonstrate VSumMamba's superior performance, achieving 67.5% and 56.0% F1-scores on TVSum and SumMe respectively, while maintaining lower computational cost compared to existing state-of-the-art methods. Yamiao Ding, Tianrui Liu 0001, Zhizhou Lu, Junjie Huang 0001, Xinwang Liu 0002, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2025 | SPA: A poisoning attack framework for graph neural networks through searching and pairingabstractGraph Neural Networks (GNN) have played an important role in many fields, while GNNs also suffer from adversarial attacks that aim to malfunction the GNN model by changing the adjacency matrix (i.e. generating adversarial edges) or node features (i.e. generating adversarial features) in graph data. Although the gradient-based adversarial attack methods have achieved remarkable results in DNNs, optimizing discrete adversarial edges in graph data using continuous gradients may lead to sub-optimal solutions. In order to alleviate this situation, we propose a novel Searching and Pairing Attack (SPA) method to effectively generate adversarial edges by treating each adversarial edge as a combination of a pair of adversarial nodes. The proposed SPA method generates the adversarial edges through a Node Searching step and a Node Pairing step. The proposed Node Searching Ant Colony Optimization (NS-ACO) improves the attack effect by using the ability of heuristic algorithm to quickly find the approximate optimal solution, while in the Node Pairing (NP) step we propose a generative graph convolutional network with a novel Aggregate Cooperative (AC) layer to generate a set of nodes that meet the constraints, so as to obtain the perturbation set together with the Node Searching step. The proposed SPA method outperforms the state-of-the-art adversarial attack methods and achieves a misclassification rate of 32.5% in the poisoning attack on Cora dataset with a perturbation rate of 0.5%. Xiao Liu 0031, Junjie Huang 0001 |
Mach. Learn. | 2 |
| 2025 | A Lightweight Deep Exclusion Unfolding Network for Single Image Reflection RemovalabstractSingle Image Reflection Removal (SIRR) is a canonical blind source separation problem and refers to the issue of separating a reflection-contaminated image into a transmission and a reflection image. The core challenge lies in minimizing the commonalities among different sources. Existing deep learning approaches either neglect the significance of feature interactions or rely on heuristically designed architectures. In this paper, we propose a novel Deep Exclusion unfolding Network (DExNet), a lightweight, interpretable, and effective network architecture for SIRR. DExNet is principally constructed by unfolding and parameterizing a simple iterative Sparse and Auxiliary Feature Update (i-SAFU) algorithm, which is specifically designed to solve a new model-based SIRR optimization formulation incorporating a general exclusion prior. This general exclusion prior enables the unfolded SAFU module to inherently identify and penalize commonalities between the transmission and reflection features, ensuring more accurate separation. The principled design of DExNet not only enhances its interpretability but also significantly improves its performance. Comprehensive experiments on four benchmark datasets demonstrate that DExNet achieves state-of-the-art visual and quantitative results while utilizing only approximately 8% of the parameters required by leading methods. Junjie Huang 0001, Tianrui Liu 0001, Xinwang Liu 0002, Meng Wang 0001, Pier Luigi Dragotti |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Robustness Matters: Pre-Training Can Enhance the Performance of Encrypted Traffic AnalysisabstractModels with large-scale parameters and pre-training have been leveraged for encrypted traffic analysis. However, existing researches primarily focused on accuracy, often overlooking the role of large-scale pre-trained parameters in enhancing robustness. While machine learning (ML) and deep learning (DL) models trained from scratch can achieve high accuracy, they exhibit limited robustness. When subjected to network noise in real-world, their identification results can fluctuate significantly, which is unacceptable. Unfortunately, current robustness evaluation methods neglect samples diversity and employ unreasonable noise settings. This field still lacks a reasonable quantitative description of models robustness. In this paper, we propose the PA-curve to display the distribution of sample’s correct-decision stability, which can simultaneously reflect the model’s accuracy and robustness. By calculating the area under the PA-curve, called PA-area, we enable the quantitative assessment of robustness for encrypted traffic analysis. Furthermore, we design a pre-trained model based on packet length sequence, and pre-trained it on TB-scale traffic. By fine-tuning on limited labeled training data, it can achieve downstream analysis tasks. We conduct experiments on five encrypted traffic datasets with different tasks. Besides accuracy, we analyzed the robustness of the pre-trained model and existing methods under common network disturbances, including packet loss, retransmission, and disorder. Experimental results demonstrated that, compared to ML-based and DL-based models trained from scratch, the pre-trained model can not only achieve high accuracy, but also exhibit greater resilience to network noise. The source code is available at https://anonymous.4open.science/r/BERT-ps-4630. Luming Yang, Lin Liu 0018, Junjie Huang 0001, Jiangyong Shi, Shaojing Fu, Jinshu Su |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | unFlowS: An Unsupervised Construction Scheme of Flow Spectrum for Network Traffic DetectionabstractIn recent years, the construction of behavior-based analysis models is hindered by issues such as insufficient data, difficulty in labeling, and the complexity of behavior types. In reality, specific cyber threats often require manual analysis of raw network traffic, which is a complex and inefficient process. Flow spectrum can simplify the complex analysis process of raw network flow by mapping it from a high-dimensional space to a one-dimensional spectral space. However, the existing flow spectrum cannot adapt to the open-world scenarios and behavior-based detection for unknown cyber threats. To address these challenges, we propose a new flow spectrum construction scheme, named unFlowS, to effectively represent network flows and assist analysts to understand the behaviors of network traffic. unFlowS-Net, an unsupervised flow-based detection model we designed as the core of our scheme, can transform network flows into spectral lines. It makes unFlowS possible to detect unknown cyber threats. We further build spectral vectors for spectral lines generated by network flow sets, enabling the visualization of network behaviors within a period of time and automatic behavior-based detection. Experimental results demonstrated that unFlowS-Net can achieve better performance than state-of-the-art methods on unsupervised flow-based detection. Based on spectral vectors, not only can it intuitively display the network behavior characteristic of the target host, but also automatically detect suspicious network behaviors. Luming Yang, Lin Liu 0018, Junjie Huang 0001, Jiangyong Shi, Shaojing Fu, Shize Guo |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Multi-View Clustering via Multi-Stage FusionabstractMulti-view clustering (MVC) exploits the information captured from diverse views to partition data into different groups and attracts much attention recently. Despite significant progress, most MVC methods fuse multi-view information via one-stage fusion while neglecting the merits of multi-stage fusion which causes insufficient in utilizing rich information within data and therefore degrades the clustering performance. To this end, designing a functional framework that can fully exploit multi-view information becomes a key challenge in multi-view clustering research. In this paper, we propose a novel multi-stage fusion method, which elegantly unifies the late and early fusion into one unified framework, to capture sufficient information underlying the multi-view data and to effectively reduce the effect of low-quality views. Specifically, we construct a low dimensional latent representation from multi-view data by learning proper correlation among multi-view data in the early fusion stage. The late fusion establishes a new optimal combinational data partition from base partitions constructed by spectral clustering, which suppresses the influence of low-quality basic partitions. Then we couple the low dimensional latent representation with the learned combinational data partition to share the same cluster structure by$k$-means and maximization alignment. As a result, we collaboratively learn an accurate and robust partition representation for the following clustering task. Besides, the late fusion and early fusion are jointly learned to achieve mutual collaboration for better performance. Finally, an alternating optimization algorithm is designed to solve the resultant optimization problem. Extensive experiments conducted on eight datasets show the superiority of our method in terms of effectiveness and efficiency. Yu Gan 0004, Yunning You, Junjie Huang 0001, Sen Xiang, Chang Tang, Wei Hu 0001, Shan An |
IEEE Trans. Multim. | 3 |
| 2024 | DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint PriorabstractSingle image reflection removal (SIRR) problem can be interpreted as a canonical blind source separation problem and is highly ill-posed. A parameter effective, fast learning and interpretable reflection removal algorithm is essential for many vision analysis applications. In this paper, we propose a novel model-inspired and learning-based SIRR method called Deep Unfolded Reflection Removal Network (DURRNet). It combines the merits of both model-based and learning-based paradigms, leading to a more interpretable and effective deep architecture. To achieve this, we first propose a model-based optimization approach and then obtain DURRNet by unfolding an iterative step into a Unfolded Separation Block (USB) based on proximal gradient descent. Key features of DURR-Net include the use of Invertible Neural Networks to impose the transform-based exclusion prior on the basis of natural image prior, as well as a coarse-to-fine architecture to fine-grain the reflection removal process. Extensive experiments on public datasets demonstrate that DURRNet achieves state-of-the-art results not only visually, quantitatively, but also effectively. Junjie Huang 0001, Tianrui Liu 0001, Jingyuan Xia, Meng Wang 0001, Pier Luigi Dragotti |
ICASSP | 1 |
| 2024 | Automatic and Aligned Anchor Learning Strategy for Multi-View ClusteringabstractMulti-View Clustering (MVC) commonly utilizes the anchor technique to mitigate the computational complexity. Existing methods generally assume a pre-selection of anchors to facilitate subsequent clustering tasks. However, the determination of the optimal number of anchors is often non-trivial and necessitates their treatment as a tunable parameter, incurring additional computational overhead. Moreover, it is not reasonable to assume an identical number of anchors across all views, as this assumption restricts the representational capacity of anchors in individual views. To address the above issues, we propose a view adaptive anchor multi-view clustering called Multi-view Clustering with Automatic and Aligned Anchor (3AMVC). We introduce a Hierarchical Bipartite Neighbor Clustering (HBNC) strategy to adaptively select a suitable number of representative anchors in each view. Specifically, when the representative difference of anchors lies in a acceptable and satisfactory range, the HBNC process is halted and picks out the final anchors. Moreover, we propose an innovative anchor alignment strategy in response to the varying quantities of anchors across different views. This approach initially evaluates the quality of anchors on each view based on the intra-cluster distance criterion and then proceeds to align based on the view with the highest-quality anchors. The carefully organized experiments well validate the effectiveness and strengthens of 3AMVC. Siwei Wang 0001, Shengju Yu, Suyuan Liu, Junjie Huang 0001, Huijun Wu 0001, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 5 |
| 2024 | DeMPAA: Deployable Multi-Mini-Patch Adversarial Attack for Remote Sensing Image ClassificationabstractDeep Neural Networks (DNNs) have demonstrated excellent performance in image classification, yet remain vulnerable to adversarial attacks. Generating deployable adversarial patches represents a promising approach to safeguard critical facilities against DNN-based classifiers used for Remote Sensing Images (RSI). While existing adversarial patch attack methods are designed for natural images, they typically generate a single and large patch which is impractically oversize for RSI applications. In this paper, we propose a Deployable Multi-Mini-Patch Adversarial Attack (DeMPAA) method for RSI classification task, which deploys multiple small adversarial patches on key locations considering both the feasibility and the effectiveness. The proposed DeMPAA method formulates the problem as a constrained optimization problem that jointly optimizes patch locations and adversarial patches. The proposed DeMPAA method takes a searching and optimization strategy to tackle it. The DeMPAA framework consists of a Feasible and Effective Map Generation (FEMG) module and a Patch Generation (PG) module. The FEMG module generates a location map to guide the adversarial patch location sampling by excluding the infeasible locations and considering the location effectiveness. In the PG module, a Probability guided Random Sampling based patch location selection (PRSamp) method is used to search better locations, then we optimize the adversarial patches using gradient descent with respect to an adversarial classification loss and an imperceptibility loss. Extensive experimental results conducted on Aerial Image Dataset show that the proposed DeMPAA method achieves 94.80% attacking success rate against ResNet50 using 16 small patches, which significantly outperforms other adversarial patch methods. Junjie Huang 0001, Tianrui Liu 0001, Wenhan Luo, Meng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Super-Resolution for Macro X-Ray Fluorescence Data Collected from Old Master PaintingsabstractMacro X-ray fluorescence (MA-XRF) scanning is commonly used to non-invasively analyse Old Master paintings by mapping the distribution of the chemical elements present in the artworks. The visual quality of the element distribution maps is very important for characterising the materials and understanding the execution and condition of the painting. However, this quality is limited by the acquisition time for the XRF datacube, resulting in a trade-off between signal-to-noise ratio (SNR) and spatial resolution. To solve this problem we propose to enhance the spatial resolution of the XRF datacube of a painting leveraging a corresponding high-resolution (HR) RGB image. We achieve that by introducing a method based on coupled dictionary learning along with a similarity constraint based on mutual information. In particular, we divide the RGB image and the XRF datacube into a common part and a unique part based on whether the information is shared or not, and then transfer the HR information between the two common parts, resulting in high-quality reconstructions. Numerical results show that our XRF super-resolution method outperforms the other state-of-the-art approaches. Su Yan 0003, Herman Verinaz-Jadan, Junjie Huang 0001, Nathan Daly, Catherine Higgitt, Pier Luigi Dragotti |
ICASSP | 3 |
| 2023 | Multi-view subspace clustering via adaptive graph learning and late fusion alignment
Chuan Tang, Kun Sun 0002, Chang Tang, Xinwang Liu 0002, Junjie Huang 0001, Wei Zhang 0049 |
Neural Networks | 6 |
| 2023 | Metalearning-Based Alternating Minimization Algorithm for Nonconvex OptimizationabstractIn this article, we propose a novel solution for nonconvex problems of multiple variables, especially for those typically solved by an alternating minimization (AM) strategy that splits the original optimization problem into a set of subproblems corresponding to each variable and then iteratively optimizes each subproblem using a fixed updating rule. However, due to the intrinsic nonconvexity of the original optimization problem, the optimization can be trapped into a spurious local minimum even when each subproblem can be optimally solved at each iteration. Meanwhile, learning-based approaches, such as deep unfolding algorithms, have gained popularity for nonconvex optimization; however, they are highly limited by the availability of labeled data and insufficient explainability. To tackle these issues, we propose a meta-learning based alternating minimization (MLAM) method that aims to minimize a part of the global losses over iterations instead of carrying minimization on each subproblem, and it tends to learn an adaptive strategy to replace the handcrafted counterpart resulting in advance on superior performance. The proposed MLAM maintains the original algorithmic principle, providing certain interpretability. We evaluate the proposed method on two representative problems, namely, bilinear inverse problem: matrix completion and nonlinear problem: Gaussian mixture models. The experimental results validate the proposed approach outperforms AM-based methods. Jingyuan Xia, Shengxi Li, Junjie Huang 0001, Zhixiong Yang 0001, Imad Jaimoukha, Deniz Gündüz |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Stability and Generalization of Kernel Clustering: from Single Kernel to Multiple KernelabstractMultiple kernel clustering (MKC) is an important research topic that has been widely studied for decades. However, current methods still face two problems: inefficient when handling out-of-sample data points and lack of theoretical study of the stability and generalization of clustering. In this paper, we propose a novel method that can efficiently compute the embedding of out-of-sample data with a solid generalization guarantee. Specifically, we approximate the eigen functions of the integral operator associated with the linear combination of base kernel functions to construct low-dimensional embeddings of out-of-sample points for efficient multiple kernel clustering. In addition, we, for the first time, theoretically study the stability of clustering algorithms and prove that the single-view version of the proposed method has uniform stability as $\mathcal{O}\left(Kn^{-3/2}\right)$ and establish an upper bound of excess risk as $\widetilde{\mathcal{O}}\left(Kn^{-3/2}+n^{-1/2}\right)$, where $K$ is the cluster number and $n$ is the number of samples. We then extend the theoretical results to multiple kernel scenarios and find that the stability of MKC depends on kernel weights. As an example, we apply our method to a novel MKC algorithm termed SimpleMKKM and derive the upper bound of its excess clustering risk, which is tighter than the current results. Extensive experimental results validate the effectiveness and efficiency of the proposed method. Weixuan Liang, Xinwang Liu 0002, Yong Liu 0018, Sihang Zhou 0001, Junjie Huang 0001, Siwei Wang 0001, Jiyuan Liu 0003, Yi Zhang 0104, En Zhu |
NeurIPS | 5 |
| 2022 | WINNet: Wavelet-Inspired Invertible Network for Image DenoisingabstractImage denoising aims to restore a clean image from an observed noisy one. Model-based image denoising approaches can achieve good generalization ability over different noise levels and are with high interpretability. Learning-based approaches are able to achieve better results, but usually with weaker generalization ability and interpretability. In this paper, we propose a wavelet-inspired invertible network (WINNet) to combine the merits of the wavelet-based approaches and learning-based approaches. The proposed WINNet consists of K -scale of lifting inspired invertible neural networks (LINNs) and sparsity-driven denoising networks together with a noise estimation network. The network architecture of LINNs is inspired by the lifting scheme in wavelets. LINNs are used to learn a non-linear redundant transform with perfect reconstruction property to facilitate noise removal. The denoising network implements a sparse coding process for denoising. The noise estimation network estimates the noise level from the input image which will be used to adaptively adjust the soft-thresholds in LINNs. The forward transform of LINNs produces a redundant multi-scale representation for denoising. The denoised image is reconstructed using the inverse transform of LINNs with the denoised detail channels and the original coarse channel. The simulation results show that the proposed WINNet method is highly interpretable and has strong generalization ability to unseen noise levels. It also achieves competitive results in the non-blind/blind image denoising and in image deblurring. Junjie Huang 0001, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 1 |
| 2022 | Video Summarization Through Reinforcement Learning With a 3D Spatio-Temporal U-NetabstractIntelligent video summarization algorithms allow to quickly convey the most relevant information in videos through the identification of the most essential and explanatory content while removing redundant video frames. In this paper, we introduce the 3DST-UNet-RL framework for video summarization. A 3D spatio-temporal U-Net is used to efficiently encode spatio-temporal information of the input videos for downstream reinforcement learning (RL). An RL agent learns from spatio-temporal latent scores and predicts actions for keeping or rejecting a video frame in a video summary. We investigate if real/inflated 3D spatio-temporal CNN features are better suited to learn representations from videos than commonly used 2D image features. Our framework can operate in both, a fully unsupervised mode and a supervised training mode. We analyse the impact of prescribed summary lengths and show experimental evidence for the effectiveness of 3DST-UNet-RL on two commonly used general video summarization benchmarks. We also applied our method on a medical video summarization task. The proposed video summarization method has the potential to save storage costs of ultrasound screening videos as well as to increase efficiency when browsing patient video data during retrospective analysis or audit without loosing essential information. Tianrui Liu 0001, Qingjie Meng, Junjie Huang 0001, Athanasios Vlontzos, Daniel Rueckert, Bernhard Kainz |
IEEE Trans. Image Process. | 3 |
| 2022 | Mixed X-Ray Image Separation for Artworks With Concealed DesignsabstractIn this paper, we focus on X-ray images (X-radiographs) of paintings with concealed sub-surface designs (e.g., deriving from reuse of the painting support or revision of a composition by the artist), which therefore include contributions from both the surface painting and the concealed features. In particular, we propose a self-supervised deep learning-based image separation approach that can be applied to the X-ray images from such paintings to separate them into two hypothetical X-ray images. One of these reconstructed images is related to the X-ray image of the concealed painting, while the second one contains only information related to the X-ray image of the visible painting. The proposed separation network consists of two components: the analysis and the synthesis sub-networks. The analysis sub-network is based on learned coupled iterative shrinkage thresholding algorithms (LCISTA) designed using algorithm unrolling techniques, and the synthesis sub-network consists of several linear mappings. The learning algorithm operates in a totally self-supervised fashion without requiring a sample set that contains both the mixed X-ray images and the separated ones. The proposed method is demonstrated on a real painting with concealed content, Do na Isabel de Porcel by Francisco de Goya, to show its effectiveness. Junjie Huang 0001, Barak Sober, Nathan Daly, Catherine Higgitt, Ingrid Daubechies, Pier Luigi Dragotti, Miguel R. D. Rodrigues |
IEEE Trans. Image Process. | 2 |
| 2021 | Deep phase retrieval: Analyzing over-parameterization in phase retrieval
Junjie Huang 0001, Jubo Zhu, Wei Dai 0001, Pier Luigi Dragotti |
Signal Process. | 2 |
| 2021 | Coupled Network for Robust Pedestrian Detection With Gated Multi-Layer Feature Extraction and Deformable Occlusion HandlingabstractPedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, detecting ismall-scaled pedestrians and occluded pedestrians remains a challenging problem. In this paper, we propose a pedestrian detection method with a couple-network to simultaneously address these two issues. One of the sub-networks, the gated multi-layer feature extraction sub-network, aims to adaptively generate discriminative features for pedestrian candidates in order to robustly detect pedestrians with large variations on scale. The second sub-network targets on handling the occlusion problem of pedestrian detection by using deformable regional region of interest (RoI)-pooling. We investigate two different gate units for the gated sub-network, namely, the channel-wise gate unit and the spatio-wise gate unit, which can enhance the representation ability of the regional convolutional features among the channel dimensions or across the spatial domain, repetitively. Ablation studies have validated the effectiveness of both the proposed gated multi-layer feature extraction sub-network and the deformable occlusion handling sub-network. With the coupled framework, our proposed pedestrian detector achieves promising results on both two pedestrian datasets, especially on detecting small or occluded pedestrians. On the CityPersons dataset, the proposed detector achieves the lowest missing rates (i.e. 40.78% and 34.60%) on detecting small and occluded pedestrians, surpassing the second best comparison method by 6.0% and 5.87%, respectively. Tianrui Liu 0001, Wenhan Luo, Lin Ma 0002, Junjie Huang 0001, Tania Stathaki, Tianhong Dai |
IEEE Trans. Image Process. | 4 |
| 2020 | Reconstruction of Fri Signals Using Deep Neural Network ApproachesabstractFinite Rate of Innovation (FRI) theory considers sampling and reconstruction of classes of non-bandlimited continuous signals that have a small number of free parameters, such as a stream of Diracs. The task of reconstructing FRI signals from discrete samples is often transformed into a spectral estimation problem and solved using Prony's method and matrix pencil method which involve estimating signal subspaces. They achieve an optimal performance given by the Cramer-Rao bound yet break down at a certain peak signal-to-́ noise ratio (PSNR). This is probably due to the so-called subspace swap event. In this paper, we aim to alleviate the subspace swap problem and investigate alternative approaches including directly estimating FRI parameters using deep neural networks and utilising deep neural networks as denoisers to reduce the noise in the samples. Simulations show significant improvements on the breakdown PSNR over existing FRI methods, which still outperform learning-based approaches in medium to high PSNR regimes. Vincent C. H. Leung, Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 2 |
| 2020 | Gated Multi-Layer Convolutional Feature Extraction Network for Robust Pedestrian DetectionabstractPedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, it remains a challenging problem how to robustly detect pedestrians of varied sizes and with occlusions. In this paper, we propose a gated multi-layer convolutional feature extraction method which can adaptively generate discriminative features for candidate pedestrian regions. The proposed gated feature extraction framework consists of squeeze units, gate units and concatenation layers which perform feature dimension squeezing, feature manipulation and features combination from multiple CNN layers, respectively. We proposed two different gate models that can manipulate the regional feature maps in a channel-wise selection manner and a spatial-wise selection manner, respectively. Experiments on the challenging CityPersons dataset demonstrate the effectiveness of the proposed method, especially on detecting small-size and occluded pedestrians. Tianrui Liu 0001, Junjie Huang 0001, Tianhong Dai, Guangyu Ren, Tania Stathaki |
ICASSP | 2 |
| 2020 | Revealing Hidden Drawings in Leonardo's 'the Virgin of the Rocks' from Macro X-Ray Fluorescence Scanning Data through Element Line LocalisationabstractMacro X-Ray Fluorescence (XRF) scanning is an increasingly widely used imaging technique for the non-invasive detection and mapping of chemical elements in Old Master paintings. Existing approaches for XRF signal analysis require varying degrees of expert user input. They are mainly based on peak fitting at fixed energies associated with each element and require the target elements to be selected manually. In this paper, we propose a new method that can process macro XRF scanning data from paintings fully automatically. The method consists of two parts: 1) detecting pulses in an XRF spectrum using Finite Rate of Innovation (FRI) theory; 2) producing the distribution maps for each element automatically identified in the painting. The results presented show the ability of our method to detect weak or partially overlapping signals and more excitingly to have visualisation of underdrawing in a masterpiece by Leonardo da Vinci. Su Yan 0003, Junjie Huang 0001, Nathan Daly, Catherine Higgitt, Pier Luigi Dragotti |
ICASSP | 2 |
| 2019 | A Deep Dictionary Model to Preserve and Disentangle Key Features in a SignalabstractWe propose a deep dictionary model for single image super-resolution (SISR) made of multiple layers of analysis dictionaries interlaced with corresponding soft-thresholding operations and a single synthesis dictionary. In this paper, we introduce a novel method for learning analysis dictionary and thresholding pairs as building block for the deep dictionary model. Each analysis dictionary contains two sub-dictionaries: an information preserving analysis dictionary (IPAD) and a clustering analysis dictionary (CAD). The IPAD and thresholding pair passes the key information from the previous layer, while the CAD and thresholding pair gives a sparse representation of its input data that facilitates discrimination of key features. Simulation results show that the proposed deep dictionary model achieves comparable performance with a deep neural network which has the same structure and is optimized using backpropagation. Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 1 |
| 2018 | U-Fresh: An Fri-Based Single Image Super Resolution Algorithm and An Application in Image CompressionabstractLearning based single image super resolution (SISR) methods have achieved notable results, however, they require large datasets for training, and may struggle when there is a mismatch between the testing and training data. To overcome these drawbacks, we propose an approach, named U - FRESH, which only requires a small dataset but can achieve state-of-the-art performance also in the presence of training and testing mismatches. We accomplish this by leveraging a method called FRESH, which enhances the image resolution using FRI theory. We start upscaling from the FRESH generated low resolution image. To minimize the reconstruction error, we propose a new regression selection technique to make the mapping more reliable and robust, and a wavelet based back projection technique to improve the quality of the reconstructed image. Based on U - FRESH, we also propose a new framework based on JPEG 2000 for image compression. Numerical results show that our U-FRESH method achieves state-of-the-art performance in SISR and provides better compression results than JPEG 2000. Xin Deng 0002, Junjie Huang 0001, Mengying Liu, Pier Luigi Dragotti |
ICASSP | 2 |
| 2018 | A Deep Dictionary Model for Image Super-ResolutionabstractInspired by the recent success of deep neural network architectures and the recent effort to develop multi-layer sparse models, we propose a novel deep dictionary learning architecture which is optimized to address a specific regression task known as single image super-resolution. Contrary to other multi-layer dictionaries, our architecture contains L-1 analysis dictionaries to extract high-level features and one synthesis dictionary which is designed to optimize the regression task. We propose a variation of an existing method to learn the analysis dictionaries and we update them without the need to use a back-propagation approach. Results on image super-resolution are satisfactory. Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 1 |
| 2018 | Photo Realistic Image Completion via Dense CorrespondenceabstractIn this paper, we propose an image completion algorithm based on dense correspondence between the input image and an exemplar image retrieved from the Internet. Contrary to traditional methods which register two images according to sparse correspondence, in this paper, we propose a hierarchical PatchMatch method that progressively estimates a dense correspondence, which is able to capture small deformations between images. The estimated dense correspondence has usually large occlusion areas that correspond to the regions to be completed. A nearest neighbor field (NNF) interpolation algorithm interpolates a smooth and accurate NNF over the occluded region. Given the calculated NNF, the correct image content from the exemplar image is transferred to the input image. Finally, as there could be a color difference between the completed content and the input image, a color correction algorithm is applied to remove the visual artifacts. Numerical results show that our proposed image completion method can achieve photo realistic image completion results. Junjie Huang 0001, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 1 |
| 2017 | ProSparse extension: Prony's based sparse pattern recovery with extended dictionariesabstractProSparse is a Prony's based method that solves the sparse representation problem of signals in the union of Fourier and canonical bases. By exploiting the structure of the dictionary, ProSparse is able to reconstruct sparse signals beyond the recovery bound of Basis Pursuit. We generalize this framework for a broader class of dictionaries which are still formed from the union of two bases. The proposed algorithm achieves perfect reconstruction over a lower sparsity level than Basis Pursuit in noiseless cases. In the presence of noise, we extend the ProSparse Denoise algorithm to the generalized dictionaries by considering their intrinsic structure. The original ProSparse can be viewed as a special case of our proposed algorithm. From simulation results, our approach outperforms state-of-the-art algorithms. Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 1 |