EDBT 2026 Demo / reviewers in the wild / expert
Xiwu Shang
dblp:142/0070
· DBLP profile ↗
18ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0001-9266-715XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast multi-type tree partitioning via lightweight multilayer perceptron for video-based point cloud compression
Peizhi Cheng, Xiwu Shang, Chenjie Hu, Xiaoli Zhao 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | TCGFNet: Multi-scale Transformer-Convolution with Geometry-Guided Feedback for Robust Point Cloud Denoising
Kekun Jin, Xiwu Shang |
ICIG (1) | 2 |
| 2023 | Super-resolution network with dynamic cleanup and temporal-spatial attention for compressed videos
Zhenting Zhou, Xiwu Shang |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Low complexity inter coding scheme for Versatile Video Coding (VVC)
Xiwu Shang, Xiaoli Zhao 0003, Yifan Zuo 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Fast CU size decision algorithm for VVC intra coding
Xiwu Shang, Xiaoli Zhao 0003, Hua Han 0002, Yifan Zuo 0001 |
Multim. Tools Appl. | 1 |
| 2022 | Color-Sensitivity-Based Rate-Distortion Optimization for H.265/HEVCabstractRate-Distortion Optimization (RDO) is an important step in video coding to achieve the best quality under a certain compression ratio constraint. The traditional RDO assigns equal importance to different color components. However, Human Visual System (HVS) has different sensitivities to different components. In this paper, the color-sensitivity-based combined PSNR (CSPSNR) is utilized as the distortion measurement in the process of RDO, where the characteristics of the color sensitivities of HVS are taken into account. Firstly, the distortion weights of luma and chroma components are derived from the criterion of maximizing CSPSNR. Then Lagrange multiplier and quantization parameter (QP) are adjusted according to the variation of distortion weights among different components. Finally, the CSPSNR-based RDO (CSRDO) adaptively calculates the RD costs of luma and chroma components under different sampling rates to improve the coding efficiency of the whole sequence. Experimental results in H.265/HEVC demonstrate that the proposed method can achieve 3.11% and 3.58% BD-RATE gain for AI and RA configurations in terms of CSPSNR on average. Xiwu Shang, Jie Liang 0001, Xiaoli Zhao 0003, Hua Han 0002, Yifan Zuo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | MIG-Net: Multi-Scale Network Alternatively Guided by Intensity and Gradient Features for Depth Map Super-ResolutionabstractThe studies of previous decades have shown that the quality of depth maps can be significantly lifted by introducing the guidance from intensity images describing the same scenes. With the rising of deep convolutional neural network, the performance of guided depth map super-resolution is further improved. The variants always consider deep structure, optimized gradient flow and feature reusing. Nevertheless, it is difficult to obtain sufficient and appropriate guidance from intensity features without any prior. In fact, features in the gradient domain, e.g., edges, present strong correlations between the intensity image and the corresponding depth map. Therefore, the guidance in the gradient domain can be more efficiently explored. In this paper, the depth features are iteratively upsampled by 2×. In each upsampling stage, the low-quality depth features and the corresponding gradient features are iteratively refined by the guidance from the intensity features via two parallel streams. Then, to make full use of depth features in the image and gradient domains, the depth features and gradient features are alternatively complemented with each other. Compared with state-of-the-art counterparts, the sufficient experimental results show improvements according to the objective and subjective assessments. The code is available athttps://github.com/Yifan-Zuo/MIG-net-gradient_guided_depth_enhancement. Yifan Zuo 0001, Yuming Fang 0001, Xiaoshui Huang, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | AFLNet: Adversarial focal loss network for RGB-D salient object detection
Xiaoli Zhao 0003, Jenq-Neng Hwang, Xiwu Shang |
Signal Process. Image Commun. | 4 |
| 2021 | KISS+ for Rapid and Accurate Pedestrian Re-IdentificationabstractPedestrian re-identification (Re-ID) is a very challenging and unavoidable problem in the field of multi-camera surveillance in smart transportation. Among many ways to solve this problem, keep it simple and straightforward (KISS) metric learning (KISSME) stands out since it has unbeatable advantages in running time while maintaining highly acceptable matching rate. It can be used to realize effective pedestrian Re-ID in an open world. Although it has achieved highly acceptable performance in some applications, it encounters a small sample size (S3) problem that causes too small eigenvalues of its covariance matrix, thus resulting in an instability issue. Its large eigenvalues are overestimated; while its small ones are underestimated. In order to solve this problem, we use an orthogonal basis vector to generate virtual samples to overcome the S3problem. The resulting algorithm named KISS+ is experimentally shown to have the eigenvalues of its covariance matrix significantly larger than those of the original KISSME. In order to show its advantage in pedestrian Re-ID, this work uses multi-feature fusion to extract more discriminant features, and obtain a low-dimensional expression of features through dimension reduction. Experiments based on several well-known databases show that our method can improve the matching rate, while maintaining the advantage of fast computation. Compared with deep learning algorithms, our algorithm does not achieve their matching rate, but it is highly suitable for real-time pedestrian Re-ID of an open world due to its simplicity, easy operation and fast execution. Hua Han 0002, MengChu Zhou, Xiwu Shang, Abdullah Abusorrah |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Frequency-Dependent Depth Map Enhancement via Iterative Depth-Guided Affine Transformation and Intensity-Guided RefinementabstractRecently, deep convolutional neural network sho-ws significant improvement for intensity-guided depth map enhancement. The most networks focus on either increasing depth or easing features propagation via residual learning and dense connection. However, it has not been explicitly considered yet to mitigate the artifacts caused by the differences of the distributions between the depth map and the corresponding color image, e.g., edge misalignment. In this paper, a novel depth-guided affine transformation is used to filter out the unrelated intensity features, which is further used to refine the depth features. Since the quality of initial depth features is low, the depth-guided intensity features filtering and the intensity-guided depth features refinement are iteratively performed, which progressively promotes effects of such tasks. To make full use of the iterations, all the refined depth features are dense connected followed by a 1 × 1 convolution layer. In addition, to improve the performance in the case of large upsampling factors (e.g., 16×), the depth features are enhanced from coarse to fine. In each frequency-dependent refinement of the depth features, the above iterative subnetwork as well as the residual learning are introduced. The proposed method is tested for the noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Ping An 0001, Xiwu Shang, Junnan Yang |
IEEE Trans. Multim. | 4 |
| 2020 | A New Deep Learning Method Based on Unsupervised Domain Adaptation and Re-ranking in Person Re-identificationabstractPerson re-identification (Re-ID) is a research hot spot in the field of intelligent video analysis, and it is also a challenging task. As the number of samples grows larger, traditional metric and feature learning methods fall into bottleneck, while it just meets the needs of deep learning algorithm, which perform very well in person re-identification. Although they have achieved good results in the field of supervised learning, their application in real-world scenarios is not very satisfactory. This is mainly because in the real world, a huge number of labeled images are hard to obtain, and even if they are obtained, the cost is expensive. Meanwhile, the performance of deep learning in unsupervised metrics is not ideal. For solving the problem, we propose a new method based on unsupervised domain adaptation (UDA) and re-ranking, and name it UDA[Formula: see text]. As for this method, we first train a camera-aware style transfer model to gain camstyle images. Then we further reduce the difference between the domain of the target and source by using invariant feature, and further improve their commonality. In addition, re-ranking is also introduced to optimize the matching results. This method can not only reduce the cost of obtaining labeled data, but also improve the accuracy. Experimental results show that our method can outperform the most advanced method by 4% on Rank-1 and 14% on mAP. The results also better confirm the effectiveness of Re-ranking module and provide a new idea for domain adaptation by unsupervised methods in the future. Hua Han 0002, Xiwu Shang, Xiaoli Zhao 0003 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2020 | Depth Map Enhancement by Revisiting Multi-Scale Intensity Guidance Within Coarse-to-Fine StagesabstractBeing different from the most methods of guided depth map enhancement based on deep convolutional neural network which focus on increasing the depth of networks, this paper is to improve the effectiveness of intensity guidance when the network goes deep. Overall, the proposed network upsamples the low-resolution depth maps from coarse to fine. Within each refinement stage of certain-scale depth features, the current-scale and all coarse-scales of the guidance features are revisited by dense connection. Therefore, the multi-scale guidance is efficiently maintained as the propagation of features. Furthermore, the proposed network maintains the intensity features in the high-resolution domain from which the multi-scale guidance is directly extracted. This design further improves the quality of intensity guidance. In addition, the shallow depth features upsampled via transposed convolution layer are directly transferred to the final depth features for reconstruction, which is called global residual learning in feature domain. Similarly, the global residual learning in pixel domain learns the difference between the depth ground truth and the coarsely upsampled depth map. Also, the local residual learning is to maintain the low frequency within each refinement stage and progressively recover the high frequency. The proposed method is tested for noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Residual dense network for intensity-guided depth map enhancement
Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang |
Inf. Sci. | 4 |
| 2019 | Color-Sensitivity-Based Combined PSNR for Objective Video Quality AssessmentabstractThe peak signal-to-noise ratio (PSNR) has been widely employed as an objective video quality assessment (VQA) metric. Usually, videos are represented in the YCbCr color space, which results in three PSNR values for each video frame. Several VQA metrics have been proposed to measure the video quality with a single combined PSNR. However, these metrics are derived heuristically without theoretical justification. In this paper, based on our extensive subjective tests on the sensitivity of the human visual system to different color components, we derive the optimal weighting coefficients of a color-sensitivity-based combined PSNR (CSPSNR). Moreover, to verify the performance of the combined PSNR, test sequences with different levels of combined PSNRs are used to evaluate the quality of the videos. However, no such database is currently available for measuring the effectiveness of different methods regarding combined PSNRs. In this paper, we design a novel coding scheme to produce sequences whose PSNRs are the combinations of different levels of PSNRs of YCbCr, with which the correlation between the subjective score and the combined PSNR is analyzed. Experiment results and statistical analysis demonstrate that the proposed CSPSNR correlates better with the mean opinion score than the existing methods. Xiwu Shang, Jie Liang 0001, Haiwu Zhao, Chengjia Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Octagonal Mapping Scheme for Panoramic Video EncodingabstractAs modern video coding standards are not designed to code panoramic videos, the pixels on the sphere need to be sampled onto a rectangle, and this process is called mapping. The mapping schemes generally include two steps: sampling points on sphere and arranging points into one compression-friendly rectangular frame. Traditional mapping schemes including equirectangular and cubic mapping have high sampling density on some sampling areas, which result in wasted pixels. In this letter, we propose a novel octagonal mapping scheme, which can decrease the oversampling areas and arrange points into an octagon. The octagon can be reshaped and rearranged into a rectangle before encoding. Experimental results demonstrate that the proposed mapping scheme saves more bitrates, compared to the existing mapping schemes. Chengjia Wu, Haiwu Zhao, Xiwu Shang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | A new combined PSNR for objective video quality assessmentabstractIn video coding, quality evaluation is important for improving the coding efficiency. Usually Peak Signal-to-Noise Ratio (PSNR) is utilized to measure the performance of different coding techniques. During the video coding process in YCbCr color space, there are three PSNRs, one for each color component. Sometimes they may contradict to each other, which poses a problem for evaluating the coding performance. Several video quality assessment (VQA) metrics have been proposed to measure the video quality with a combined PSNR. However, these combined PSNRs are obtained heuristically without theoretical justification. In this paper, we propose a color-sensitivity-based combined PSNR (CSP-SNR) based on extensive subjective tests on the sensitivity of human visual system (HVS) to different color components. Subjective experiment results demonstrate that the proposed combined PSNR correlates well with the mean opinion score (MOS) than existing methods. Xiwu Shang, Haiwu Zhao, Jie Liang 0001, Chengjia Wu |
ICME | 1 |
| 2015 | Fast CU size decision and PU mode decision algorithm in HEVC intra codingabstractHigh Efficiency Video Coding (HEVC) is a new video compression standard for high resolution video content, which only needs 50% bit rate of the H.264/AVC at the same perceptual quality. However, high computational complexity increases dramatically for adopting quad-tree structured Coding Unit (CU). In this paper, a fast CU size decision and Prediction Unit (PU) mode decision algorithm is presented for HEVC intra coding. It exploits the depth information of neighboring CUs to make an early CU split decision or CU pruning decision. Moreover, there are correlations between the higher layer prediction mode and current layer mode in PUs. By using these correlations, we can terminate some prediction modes which are rarely selected as the optimal mode. Experimental results illustrate that the proposed method can save 37.91% computational complexity on average as compared with the current HM with only 0.66% increase in BDBR and 0.03 dB loss in BDPSNR. Xiwu Shang |
ICIP | 1 |
| 2013 | Perceptual multiview video coding based on foveated just noticeable distortion profile in DCT domainabstractRecently just noticeable distortion (JND) has been highly successful in improving the video coding efficiency. Foveated JND (FJND) is an extension of the JND by further exploiting the human vision characteristic. However, there is a challenge to quickly estimate the foveation point and accurately combine the foveation factor with the spatio-temporal JND model of DCT domain for FJND. In this paper, a new FJND model in DCT domain is proposed, which adaptively searches the foveation point by exploiting the property of image signature and builds a foveation model based on the contrast threshold. Experimental results demonstrate that the proposal model remarkably reduces the complexity of FJND in searching foveation. In a number of video coding experiments, we find that, in terms of coding efficiency, the proposed perceptual coding based on FJND for multiview video coding (MVC) significantly outperforms the existing algorithm. Xiwu Shang, Lidong Luo, Zhaoyang Zhang 0002 |
ICIP | 1 |