EDBT 2026 Demo / reviewers in the wild / expert
Jun Xiao 0010
dblp:71/2308-10
· DBLP profile ↗
23ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-4935-7866ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SfM-free 3D Gaussian Splatting from extremely sparse view
Zongqi He, Hanmin Li, Kin-Chung Chan, Yushen Zuo, Zhe Xiao 0001, Jun Xiao 0010, Xiaoyang Bai, Kin-Man Lam 0001 |
Comput. Graph. | 7 |
| 2025 | See In Detail: Enhancing Sparse-view 3D Gaussian Splatting with Local Depth and Semantic Regularizationabstract3D Gaussian Splatting (3DGS) has shown remarkable performance in novel view synthesis. However, its rendering quality deteriorates with sparse inphut views, leading to distorted content and reduced details. This limitation hinders its practical application. To address this issue, we propose a sparse-view 3DGS method. Given the inherently ill-posed nature of sparse-view rendering, incorporating prior information is crucial. We propose a semantic regularization technique, using features extracted from the pretrained DINO-ViT model, to ensure multi-view semantic consistency. Additionally, we propose local depth regularization, which constrains depth values to improve generalization on unseen views. Our method outperforms state-of-the-art novel view synthesis approaches, achieving up to 0.4dB improvement in terms of PSNR on the LLFF dataset, with reduced distortion and enhanced visual quality. Zongqi He, Zhe Xiao 0001, Kin-Chung Chan, Yushen Zuo, Jun Xiao 0010, Kin-Man Lam 0001 |
ICASSP | 5 |
| 2025 | Geometric Distortion Guided Transformer for Omnidirectional Image Super-ResolutionabstractAs virtual and augmented reality applications gain popularity, omnidirectional image (ODI) super-resolution has become increasingly important. Unlike 2D plain images that are formed on a plane, ODIs are projected onto spherical surfaces. Applying established image super-resolution methods to ODIs, therefore, requires performing equirectangular projection (ERP) to map the ODIs onto a plane. ODI super-resolution needs to take into account geometric distortion resulting from ERP. However, without considering such geometric distortion of ERP images, previous methods only utilize a limited range of pixels and may easily miss self-similar textures for reconstruction. In this paper, we introduce a novel Geometric Distortion Guided Transformer for Omnidirectional image Super-Resolution (GDGT-OSR). Specifically, a distortion modulated rectangle-window selfattention mechanism, integrated with deformable self-attention, is proposed to better perceive the distortion and thus involve more self-similar textures. Distortion modulation is achieved through a newly devised distortion guidance generator that produces guidance for the rectangular windows by exploiting the variability of distortion across latitudes. Furthermore, we propose a dynamic feature aggregation scheme to adaptively fuse the features from different self-attention modules. We present extensive experimental results on public datasets and show that the new GDGT-OSR outperforms methods in existing literature. Cuixin Yang, Rongkang Dong, Jun Xiao 0010, Kin-Man Lam 0001, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Towards Progressive Multi-Frequency Representation for Image WarpingabstractImage warping, a classic task in computer vision, aims to use geometric transformations to change the appearance of images. Recent methods learn the resampling kernels for warping through neural networks to estimate missing values in irregular grids, which, however, fail to capture local variations in deformed content and produce images with distortion and less high-frequency details. To address this issue, this paper proposes an effective method, namely MFR, to learn Multi-Frequency Representations from in-put images for image warping. Specifically, we propose a progressive filtering network to learn image representations from different frequency subbands and generate deformable images in a coarse-to-fine manner. Furthermore, we employ learnable Gabor wavelet filters to improve the model's capability to learn local spatial-frequency representations. Comprehensive experiments, including homography trans-formation, equirectangular to perspective projection, and asymmetric image super-resolution, demonstrate that the proposed MFR significantly outperforms state-of-the-art image warping methods. Our method also showcases superior generalization to out-of-distribution domains, where the generated images are equipped with rich details and less distortion, thereby high visual quality. The source code is available at https://github.com/junxiao01/MFR. Jun Xiao 0010, Zihang Lyu, Yakun Ju, Changjian Shui, Kin-Man Lam 0001 |
CVPR | 1 |
| 2024 | Learning Equilibrium Transformation for Gamut Expansion and Color Restoration
Jun Xiao 0010, Changjian Shui, Kin-Man Lam 0001 |
ECCV (71) | 1 |
| 2024 | Hierarchical Vertex-Wise Intensification Graph Convolution for Skeleton-Based Activity RecognitionabstractGraph convolutional networks (GCNs), which can effectively captures the spatial and temporal relationships between skeleton joints through graph topology, have shown promising performances in skeleton-based activity recognition in recent years. These methods typically learn the semantic features of the vertices of a skeleton and the associated adjacency matrix. However, how to efficiently establish relationships between vertices still remains a substantial problem. To solve this problem, we propose a novel Hierarchical Vertex-wise Intensification Graph Convolution Network (HVI-GCN) for skeleton-based action recognition. The proposed module dilates input features into higher dimensions to broaden the temporal horizon, and builds a vertex-wise topology based on self-adaptively learned attention. With the adjacency matrix, features from other positions can be collected to aid the prediction of the current position. The proposed module provides a better receptive field and semantic understanding of both the spatial and temporal domains than related methods. Experiments were mainly conducted on the at NTU-RGB-D, NTU-GRB-D 120, and NW-UCLA datasets with joint and bone integrated with motion sequences. Experimental results show that HVI-GCN can improve accuracy by up to 1.1% on the RGB-D 120 dataset. Meanwhile, the accuracy on RGB-D 60 dataset and NW-UCLA dataset can be boosted by 1.4% and 1.2%, respectively. Jun Xiao 0010, Tianshan Liu, Kin-Man Lam 0001 |
ICIP | 3 |
| 2024 | AI-Generated Image Detection With Wasserstein Distance Compression and Dynamic AggregationabstractWith the rapid advancement of generative models, image detectors for AI-generated content have become an increasingly necessary technology in computer vision, attracting significant attention from researchers. This technology aims to detect whether an image is naturally generated by imaging systems (e.g., digital cameras) or generated by advanced AI techniques. Despite the promising performance achieved by recent fake detection methods, they are typically trained on millions of redundant images with similar characteristics, leading to inefficient training. Furthermore, the performances of existing detectors often deteriorate when the training datasets are imbalanced. To address these challenges, we propose a novel AI-generated image detector based on dynamic aggregation and information compression with the Wasserstein distance. Experimental results show that our proposed method significantly outperforms state-of-the-art models that generalize across different generative models, with an increase of $\mathbf{+ 1. 8 6 \%}$ average accuracy and $\mathbf{+ 0. 1 4 \%}$ average precision, while substantially reducing the training time. On imbalanced datasets, our proposed method leads to a $\mathbf{+ 1 4. 4 6 \%}$ accuracy improvement, clearly demonstrating its robustness on imbalanced datasets. Zihang Lyu, Jun Xiao 0010, Kin-Man Lam 0001 |
ICIP | 2 |
| 2024 | Point Cloud Densification for 3D Gaussian Splatting from Sparse Input ViewsabstractThe technique of 3D Gaussian splatting (3DGS) has demonstrated its effectiveness and efficiency in rendering photo-realistic images for novel view synthesis. However, 3DGS requires a high density of camera coverage, and its performance inevitably degrades with sparse training views, which significantly limits its applicability in real-world scenarios. In recent years, many researchers have explored the use of depth information to alleviate this problem, but the performance of their methods is sensitive to the accuracy of depth estimation. To this end, we propose an efficient method to enhance the performance of 3DGS with sparse training views. Specifically, instead of applying depth maps for regularization, we propose a densification method that generates high-quality point clouds, providing a superior initialization for 3D Gaussians. Furthermore, we propose Systematically Angle of View Sampling (SAOVS), which employs Spherical Linear Interpolation (SLERP) and linear interpolation for side view sampling, to determine unseen views outside the training data for semantic pseudo-label regularization. Experiments show that our proposed method significantly outperforms other leading 3D rendering models on the ScanNet dataset and the LLFF dataset. In particular, compared with the conventional 3DGS method, our proposed method achieves performance gains of up to 1.71dB in PSNR and 0.07 in SSIM. In addition, the novel view synthesis produced by our method demonstrates the highest visual quality with minimal distortions. Kin-Chung Chan, Jun Xiao 0010, Hana Lebeta Goshu, Kin-Man Lam 0001 |
ACM Multimedia | 2 |
| 2024 | Deep progressive feature aggregation network for multi-frame high dynamic range imaging
Jun Xiao 0010, Tianshan Liu, Kin-Man Lam 0001 |
Neurocomputing | 1 |
| 2024 | Integrally Mixing Pyramid Representations for Anchor-Free Object Detection in Aerial ImageryabstractAnchor-free object detectors have recently received increasing research attention in the field of aerial scene object detection, due to their high flexibility and practicality. Anchor-free detectors typically depend on the feature pyramid network (FPN) to alleviate the challenge of significant variations in object scales in aerial contexts. Despite establishing a multi-scale feature pyramid, existing FPN-based methods treat each aerial object as an indivisible entity solely managed by a single-scale representation. However, they fail to take into account the distinct characteristics of various components within an instance. To this end, this letter proposes a novel anchor-free detector, namely IMPR-Det, which can integrally mix multi-scale pyramid representations for different components of an instance, thus boosting the fine-grained object representation capability. Specifically, IMPR-Det fundamentally introduces a more advanced detection head with an adaptive routing mechanism for pixel-level multi-scale feature assignment, instead of previous instance-level assignment. Experimental results demonstrate the superiority of the proposed method over its counterparts, in terms of both accuracy and efficiency, for object detection in aerial images. Jun Xiao 0010, Cuixin Yang, Jingchun Zhou, Kin-Man Lam 0001, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Deep multi-scale feature mixture model for image super-resolution with multiple-focal-length degradation
Jun Xiao 0010, Rui Zhao 0012, Kin-Man Lam 0001, Kao Wan |
Signal Process. Image Commun. | 1 |
| 2023 | Efficient Feature Fusion for Learning-Based Photometric StereoabstractHow to handle an arbitrary number for input images is a fundamental problem of learning-based photometric stereo methods. Existing approaches adopt max-pooling or observation map to fuse an arbitrary number of extracted features. However, these methods discard a large amount of the features from the input images, impacting the utilization and accuracy, or ignore the constraints from the intra-image spatial domain. In this paper, we explore how to efficiently fuse features from a variable number of input images. First, we propose a bilateral extraction module, which categorizes features into positive and negative, to maximally keep the useful feature in the fusion stage. Second, we adopt a top-k pooling to both the bilateral information, which selects the k maximum response value from all features. These two modules proposed are "plug-and-play" and can be used in different fusion tasks. We further propose a hierarchical photometric stereo network, namely HPS-Net, to handle bilateral extraction and top-k pooling for multiscale features. Experiments in the widely used benchmark illustrate the improvement of our proposed framework in the conventional max-pooling method and the proposed HPS-Net outperforms existing learning-based photometric stereo methods. Yakun Ju, Kin-Man Lam 0001, Jun Xiao 0010, Cuixin Yang, Junyu Dong |
ICASSP | 3 |
| 2023 | Improving Robustness of Single Image Super-Resolution Models with Monte Carlo MethodabstractDeep learning-based methods have achieved promising results in single image super-resolution (SISR). However, the performance of existing deep SISR methods is very sensitive to image degradation. In addition, these methods are deterministic and do not introduce any uncertainty to the generated images, so we have no way of knowing the reliability of these generated images. To address these two challenging issues, we propose a model-agnostic approach for existing deep SISR networks to improve their robustness under various degradations. Our proposed method follows a probabilistic framework and applies Monte Carlo dropout to existing deep SISR methods. Instead of performing point estimation, the proposed method predicts the posterior distribution of super-resolved images. Based on this, we can determine the uncertainty of the generated images. Experiment results show that the proposed method can effectively improve the robustness of existing deep SISR methods, leading to state-of-the-art performance when applied to images having different degradations. The code is available at https://github.com/YangTracy/MCD-SR. Cuixin Yang, Jun Xiao 0010, Yakun Ju, Guoping Qiu, Kin-Man Lam 0001 |
ICIP | 2 |
| 2023 | Online Video Super-Resolution With Convolutional Kernel Bypass GraftsabstractDeep learning-based models have achieved remarkable performance in video super-resolution (VSR) in recent years, but most of these models are less applicable to online video applications. These methods solely consider the distortion quality and ignore crucial requirements for online applications, e.g., low latency and low model complexity. In this paper, we focus on online video transmission in which VSR algorithms are required to generate high-resolution video sequences frame by frame in real time. To address such challenges, we propose an extremely low-latency VSR algorithm based on a novel kernel knowledge transfer method, named the convolutional kernel bypass graft (CKBG). First, we design a lightweight network structure that does not require future frames as inputs and saves extra time for caching these frames. Then, our proposed CKBG method enhances this lightweight base model by bypassing the original network with “kernel grafts”, which are extra convolutional kernels containing the prior knowledge of the external pretrained image SR models. During the testing phase, we further accelerate the grafted multibranch network by converting it into a simple single-path structure. The experimental results show that our proposed method can process online video sequences up to 110 FPS with very low model complexity and competitive SR performance. Jun Xiao 0010, Xinyang Jiang, Ningxin Zheng, Huan Yang 0005, Yifan Yang 0004, Yuqing Yang 0001, Dongsheng Li 0002, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Feature Redundancy Mining: Deep Light-Weight Image Super-Resolution ModelabstractDespite the great success achieved by deep convolutional neural network (CNN)-based models in the single image super-resolution (SISR) problem, the requirement of high computational complexity, accompanied with the deep CNN models, makes it less applicable in embedded devices, e.g., mobile phones. Recently, deep light-weight models for the SISR problem have been in demand for industrial applications, and have caught the attention of many researchers. The strategies of cascading several small networks and multi-path feature extraction have shown their effectiveness in most of the existing methods. In this paper, by considering the correlation and redundancy of feature maps, we propose a feature information mining network to efficiently investigate the features, for the SISR problem. Experiment results show that our proposed model achieves the best balance between the performance and the model size, compared with other competitive deep SR models. Jun Xiao 0010, Wenqi Jia 0001, Kin-Man Lam 0001 |
ICASSP | 1 |
| 2021 | Self-feature Learning: An Efficient Deep Lightweight Network for Image Super-resolutionabstractDeep learning-based models have achieved unprecedented performance in single image super-resolution (SISR). However, existing deep learning-based models usually require high computational complexity to generate high-quality images, which limits their applications in edge devices, e.g., mobile phones. To address this issue, we propose a dynamic, channel-agnostic filtering method in this paper. The proposed method not only adaptively generates convolutional kernels based on the local information of each position, but also can significantly reduce the cost of computing the inter-channel redundancy. Based on this, we further propose a simple, yet effective, deep lightweight model for SISR. Experiment results show that our proposed model outperforms other state-of-the-art deep lightweight SISR models, leading to the best trade-off between the performance and the number of model parameters. Jun Xiao 0010, Rui Zhao 0012, Kin-Man Lam 0001, Kao Wan |
ACM Multimedia | 1 |
| 2021 | Progressive and Selective Fusion Network for High Dynamic Range ImagingabstractThis paper considers the problem of generating an HDR image of a scene from its LDR images. Recent studies employ deep learning and solve the problem in an end-to-end fashion, leading to significant performance improvements. However, it is still hard to generate a good quality image from LDR images of a dynamic scene captured by a hand-held camera, e.g., occlusion due to the large motion of foreground objects, causing ghosting artifacts. The key to success relies on how well we can fuse the input images in their feature space, where we wish to remove the factors leading to low-quality image generation while performing the fundamental computations for HDR image generation, e.g., selecting the best-exposed image/region. We propose a novel method that can better fuse the features based on two ideas. One is multi-step feature fusion; our network gradually fuses the features in a stack of blocks having the same structure. The other is the design of the component block that effectively performs two operations essential to the problem, i.e., comparing and selecting appropriate images/regions. Experimental results show that the proposed method outperforms the previous state-of-the-art methods on the standard benchmark tests. Jun Xiao 0010, Kin-Man Lam 0001, Takayuki Okatani |
ACM Multimedia | 2 |
| 2021 | Balanced distortion and perception in single-image super-resolution based on optimal transport in wavelet domain
Jun Xiao 0010, Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001 |
Neurocomputing | 1 |
| 2021 | Bayesian sparse hierarchical model for image denoising
Jun Xiao 0010, Rui Zhao 0012, Kin-Man Lam 0001 |
Signal Process. Image Commun. | 1 |
| 2021 | Invertible Image DecolorizationabstractInvertible image decolorization is a useful color compression technique to reduce the cost in multimedia systems. Invertible decolorization aims to synthesize faithful grayscales from color images, which can be fully restored to the original color version. In this paper, we propose a novel color compression method to produce invertible grayscale images using invertible neural networks (INNs). Our key idea is to separate the color information from color images, and encode the color information into a set of Gaussian distributed latent variables via INNs. By this means, we force the color information lost in grayscale generation to be independent of the input color image. Therefore, the original color version can be efficiently recovered by randomly re-sampling a new set of Gaussian distributed variables, together with the synthetic grayscale, through the reverse mapping of INNs. To effectively learn the invertible grayscale, we introduce the wavelet transformation into a UNet-like INN architecture, and further present a quantization embedding to prevent the information omission in format conversion, which improves the generalizability of the framework in real-world scenarios. Extensive experiments on three widely used benchmarks demonstrate that the proposed method achieves a state-of-the-art performance in terms of both qualitative and quantitative results, which shows its superiority in multimedia communication and storage systems. Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Deep Multi-task Learning for Facial Expression Recognition and Synthesis Based on Selective Feature SharingabstractMulti-task learning is an effective learning strategy for deep-learning-based facial expression recognition tasks. However, most existing methods take into limited consideration the feature selection, when transferring information between different tasks, which may lead to task interference when training the multi-task networks. To address this problem, we propose a novel selective feature-sharing method, and establish a multi-task network for facial expression recognition and facial expression synthesis. The proposed method can effectively transfer beneficial features between different tasks, while filtering out useless and harmful information. Moreover, we employ the facial expression synthesis task to enlarge and balance the training dataset to further enhance the generalization ability of the proposed method. Experimental results show that the proposed method achieves state-of-the-art performance on those commonly used facial expression recognition benchmarks, which makes it a potential solution to real-world facial expression recognition problems. Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
ICPR | 3 |
| 2020 | Progressive Motion Representation Distillation With Two-Branch Networks for Egocentric Activity RecognitionabstractVideo-based egocentric activity recognition involves fine-grained spatio-temporal human-object interactions. State-of-the-art methods, based on the two-branch-based architecture, rely on pre-calculated optical flows to provide motion information. However, this two-stage strategy is computationally intensive, storage demanding, and not task-oriented, which hampers it from being deployed in real-world applications. Albeit there have been numerous attempts to explore other motion representations to replace optical flows, most of the methods were designed for third-person activities, without capturing fine-grained cues. To tackle these issues, in this letter, we propose a progressive motion representation distillation (PMRD) method, based on two-branch networks, for egocentric activity recognition. We exploit a generalized knowledge distillation framework to train a hallucination network, which receives RGB frames as input and produces motion cues guided by the optical-flow network. Specifically, we propose a progressive metric loss, which aims to distill local fine-grained motion patterns in terms of each temporal progress level. To further enforce the proposed distillation framework to concentrate on those informative frames, we integrate a temporal attention mechanism into the metric loss. Moreover, a multi-stage training procedure is employed for the efficient learning of the hallucination network. Experimental results on three egocentric activity benchmarks demonstrate the state-of-the-art performance of the proposed method. Tianshan Liu, Rui Zhao 0012, Jun Xiao 0010, Kin-Man Lam 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Deep Progressive Convolutional Neural Network for Blind Super-Resolution With Multiple DegradationsabstractBlind super-resolution (SR) of blurry and noisy low-resolution (LR) images is still a challenging problem in single image super-resolution (SISR). The performance of most existing convolutional neural network (CNN)-based models is inevitably degraded when LR images are corrupted by both blur and noise. For those blind SR methods based on kernel estimation, accurate estimation is barely attained under complex degradations and this gives rise to poor-quality results. To address these problems, we propose a deep progressive network under a probabilistic framework and a novel up-sampling method for blind super-resolution with multiple degradations, which effectively utilizes image priors across scales. Experimental results show that the proposed method achieves promising performance on images with multiple degradations. Jun Xiao 0010, Rui Zhao 0012, Shun-Cheung Lai, Wenqi Jia 0001, Kin-Man Lam 0001 |
ICIP | 1 |