Jun Sun 0005

dblp:s/JunSun5 · DBLP profile ↗
← Back
74ranked-venue papers
1as first author
19since 2021 · last 2026
0009-0001-1956-694XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Computer networks · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Text-Driven Relation Manipulation of Diffusion Imagery
abstract
Text-guided image manipulation has recently attracted significant attention. Prevailing algorithms predominantly focus on modifying the appearances of existing instances, such as texture and attribute editing, while they often fail to address the interactions between different instances or achieve fundamental structural changes, such as multi-object editing. This paper introduces a novel text-guided manipulation task named "relation manipulation", aimed at fundamentally altering the structure of images. This task is capable of modifying the quantity of instances and, more importantly, enhancing the understanding and editing of interactions among diverse instances. Our approach comprises two main components: relation customization and multi-region guided diffusion. Relation customization fine-tunes specific relationships using a compact dataset of exemplary relations, facilitating nuanced understanding and implementation of instance interactions. Multi-region guided diffusion employs gradient optimization to update the generation process across multiple regions, integrating a fine-grained attention control strategy to minimize regional interference and conflict. Additionally, the demonstrated applications of our method in multi-region inversion underline its potential in practical scenarios, such as relation manipulation of real images and consecutive image manipulation. Compatible with different variants of Stable Diffusion models, our approach seamlessly integrates into the Stable Diffusion WebUI, enabling high-quality image generation and exceptional control over extensive manipulation. This makes it a robust tool for both academic research and creative industries. Code is available at https://github.com/liyiming09/RMD.
Peng Zhou 0010, Hongwei Hu, Xiaokang Qin, Jun Sun 0005, Yi Xu 0001
IEEE Trans. Image Process.5
2026 Signed Relation Graph Based Dynamical Interacting System Modeling for Multi-Agent Trajectory Prediction
abstract
Many complex systems prevalent in nature and society, from particle physics systems to social networks and team sports, can be viewed as dynamical interacting systems. Understanding the underlying interactions of agents in the system is the key task for predicting future behaviors of agents, which can be applied in various applications, e.g., autonomous vehicles and smart video surveillance. Since the interaction patterns between agents in the system can be dynamic and heterogeneous rather than fixed and homogeneous, it is very challenging to model interacting systems. In this paper, we design a novel graph structure called Signed Relation Graph (SRG) to model dynamical interacting systems. Since collective behaviors are very common in real-world scenes, our method is a group based model that takes heterogeneous relationships between agents into consideration, and achieves jointly modeling inter-group interactions and intra-group interactions. To assign signs on SRG, an unsupervised method called Relationship Reasoning Network is proposed. The relationship categories are reasoned explicitly, which makes handling multi-agent systems with multiple and dynamic interactions available. Further, Group Interaction Attention Graph Neural Network is proposed to aggregate information on SRG, which achieves not only reasoning the intensity of different interaction patterns but also modeling the trade-off between inter-group interactions and intra-group interactions. Our interacting systems modeling method can be used to predict multi-agent future trajectories in a variety of scenes with hard scenarios, including dense and drastic scenarios. Experimental results on three widely used human trajectory prediction datasets, including ETH and UCY in traffic scenes and NBA SportVU in sports scenes, demonstrate the effectiveness of our proposed model.
Cunyan Li, Hua Yang 0001, Jun Sun 0005
IEEE Trans. Multim.3
2025 Neural Adaptive Contextual Video Streaming
abstract
Video streaming services typically employ traditional codecs, such as H.264, to encode videos into multiple bitrate representations. These codecs are tightly limited by discrete quantization parameters (QPs), resulting in encoded rates that do not align with the target bitrate. Additionally, the subpar video quality produced by conventional codecs does not meet the demands of high-resolution communication. Considering the limitations of traditional codecs, we take a fresh new approach to video streaming by leveraging advanced deep learning-based video codecs. Specifically, we develop a neural adaptive contextual video streaming framework that incorporates: 1) an ensemble deep reinforcement learning based adaptive bitrate algorithm named TSAC that enables continuous bitrate adjustment to varying network conditions 2) a two-stage proportional-integral-derivative-based rate control module that dynamically fine-tunes QPs to ensure the encoded bitrate aligning with the target bitrate. Furthermore, we implement intra-GoP and inter-GoP techniques to accelerate the inference process of the contextual video codec for real-time processing needs. Our experiments demonstrate that the average relative error in bitrate remains below 2%, the quality of experience provided by our TSAC agents surpasses that of existing discrete algorithms by 13%-20%. Our optimization techniques enable real-time decoding at approximately 24 frames per second for quad high definition videos.
Jianchao Yang, Mufan Liu, Puyue Hou, Yiling Xu, Jun Sun 0005
ICASSP5
2025 Position-LoRA: Enhanced Relation Customization through Structural Prior in Initial Latent Noise
abstract
Recent advancements in concept customization via diffusion models have significantly enhanced controllability and quality. However, precise relation customization, which controls the position of interactions among multiple instances, remains challenging due to unpredictable initial latent noise. Existing methods primarily rely on conditional prompts and attention control, overlooking the structured potential of initial noise. This paper introduces Position-LoRA, a novel framework leveraging structural prior in initial noise to improve relation customization and layout control. Position-LoRA employs a differential fine-tuning scheme and a latent noise encoder. The guided fine-tuning enhances generation tendencies from structured initial noise, embedding explicit relationship-specific spatial information. The latent noise encoder dynamically manipulates latent noises, enabling precise spatial control and flexibility in relational image generation. Furthermore, a fine-grained guidance and control strategy is employed during generation to enhance the image-text alignment and layout alignment. Experiments demonstrate that Position-LoRA improves stability, controllability, and fidelity in relational image generation with layout control, surpassing existing concept customization and layout-to-image methods in qualitative and quantitative evaluations. Code is available at https://github.com/liyiming09/Position-LoRA.
Peng Zhou 0010, Xiaokang Qin, Hongwei Hu, Jun Sun 0005, Yi Xu 0001
ACM Multimedia5
2024 Consistent GT-Proposal Assignment for Challenging Pedestrian Detection
abstract
Accurate pedestrian classification and localization has garnered significant attention due to their extensive applications in various multimedia applications such as security monitoring, autonomous driving, and more. We have observed that the commonly employed Intersection over Union (IoU) metric in many pedestrian detectors is susceptible to an inconsistent GT-Proposal assignment issue. This issue arises when spatially adjacent proposals, which have highly similar features, are assigned to distinct ground-truth boxes, leading to confusion during the training process and an increased number of false positives during inference. To address this challenge, our work presents a novel algorithm namedDirectionalAssignmentStrategy (DAS). Firstly, in conjunction with depth distribution, our approach transforms the assignment metric from a two-dimensional (2D) view into a three-dimensional (3D) space, enabling the optimization of the regression head under the constraint of depth direction. Secondly, in contrast to the conventional IoU-basedone-to-oneassignment of one proposal to one ground-truth box, our method aims to establish a more reasoned matching between sets of proposals and ground-truth boxes. By doing so, the detector is less reliant on the setting of a specific threshold. Leveraging this strategy as a plug-in module within state-of-the-art pedestrian detectors, we demonstrate a notable improvement in performance.
Yan Luo 0003, Muming Zhao, Jun Sun 0005, Guangtao Zhai
IEEE Trans. Multim.3
2024 MLE-Based Device Activity Detection Under Rician Fading for Massive Grant-Free Access With Perfect and Imperfect Synchronization
abstract
Most existing studies on massive grant-free access, proposed to support massive machine-type communications (mMTC) for the Internet of things (IoT), assume Rayleigh fading and perfect synchronization for simplicity. However, in practice, line-of-sight (LoS) components generally exist, and time and frequency synchronization are usually imperfect. This paper systematically investigates maximum likelihood estimation (MLE)-based device activity detection under Rician fading for massive grant-free access with perfect and imperfect synchronization. We assume that the large-scale fading powers, Rician factors, and normalized LoS components can be estimated offline. We formulate device activity detection in the synchronous case and joint device activity and offset detection in three asynchronous cases (i.e., time, frequency, and time and frequency asynchronous cases) as MLE problems. In the synchronous case, we propose an iterative algorithm to obtain a stationary point of the MLE problem. In each asynchronous case, we propose two iterative algorithms with identical detection performance but different computational complexities. In particular, one is computationally efficient for small ranges of offsets, whereas the other one, relying on fast Fourier transform (FFT) and inverse FFT, is computationally efficient for large ranges of offsets. The proposed algorithms generalize the existing MLE-based methods for Rayleigh fading and perfect synchronization. Numerical results show that the proposed algorithm for the synchronous case can reduce the detection error probability by up to 50.4% at a 78.6% computation time increase, compared to the MLE-based state-of-the-art, and the proposed algorithms for the three asynchronous cases can reduce the detection error probabilities and computation times by up to 65.8% and 92.0%, respectively, compared to the MLE-based state-of-the-arts.
Ying Cui 0001, Feng Yang 0006, Lianghui Ding, Jun Sun 0005
IEEE Trans. Wirel. Commun.5
2023 One-for-All: Proposal Masked Cross-Class Anomaly Detection
abstract
One of the most challenges for anomaly detection (AD) is how to learn one unified and generalizable model to adapt to multi-class especially cross-class settings: the model is trained with normal samples from seen classes with the objective to detect anomalies from both seen and unseen classes. In this work, we propose a novel Proposal Masked Anomaly Detection (PMAD) approach for such challenging multi- and cross-class anomaly detection. The proposed PMAD can be adapted to seen and unseen classes by two key designs: MAE-based patch-level reconstruction and prototype-guided proposal masking. First, motivated by MAE (Masked AutoEncoder), we develop a patch-level reconstruction model rather than the image-level reconstruction adopted in most AD methods for this reason: the masked patches in unseen classes can be reconstructed well by using the visible patches and the adaptive reconstruction capability of MAE. Moreover, we improve MAE by ViT encoder-decoder architecture, combinational masking, and visual tokens as reconstruction objectives to make it more suitable for anomaly detection. Second, we develop a two-stage anomaly detection manner during inference. In the proposal masking stage, the prototype-guided proposal masking module is utilized to generate proposals for suspicious anomalies as much as possible, then masked patches can be generated from the proposal regions. By masking most likely anomalous patches, the “shortcut reconstruction” issue (i.e., anomalous regions can be well reconstructed) can be mostly avoided. In the reconstruction stage, these masked patches are then reconstructed by the trained patch-level reconstruction model to determine if they are anomalies. Extensive experiments show that the proposed PMAD can outperform current state-of-the-art models significantly under the multi- and especially cross-class settings. Code will be publicly available at https://github.com/xcyao00/PMAD.
Xincheng Yao, Ruoqi Li, Jun Sun 0005
AAAI4
2023 Explicit Boundary Guided Semi-Push-Pull Contrastive Learning for Supervised Anomaly Detection
abstract
Most anomaly detection (AD) models are learned using only normal samples in an unsupervised way, which may result in ambiguous decision boundary and insufficient discriminability. In fact, a few anomaly samples are often available in real-world applications, the valuable knowledge of known anomalies should also be effectively exploited. However, utilizing a few known anomalies during training may cause another issue that the model may be biased by those known anomalies and fail to generalize to unseen anomalies. In this paper, we tackle supervised anomaly detection, i.e., we learn AD models using a few available anomalies with the objective to detect both the seen and unseen anomalies. We propose a novel explicit boundary guided semi-push-pull contrastive learning mechanism, which can enhance model's discriminability while mitigating the bias issue. Our approach is based on two core designs: First, we find an explicit and compact separating boundary as the guidance for further feature learning. As the boundary only relies on the normal feature distribution, the bias problem caused by a few known anomalies can be alleviated. Second, a boundary guided semi-push-pull loss is developed to only pull the normal features together while pushing the abnormal features apart from the separating boundary beyond a certain margin region. In this way, our model can form a more explicit and discriminative decision boundary to distinguish known and also unseen anomalies from normal samples more effectively. Code will be available at https://github.com/xcyao00/BGAD.
Xincheng Yao, Ruoqi Li, Jun Sun 0005
CVPR4
2023 Improving Point Cloud Quality Metrics with Noticeable Possibility Maps
abstract
Point cloud quality assessment (PCQA) plays a vital role in the quality of experience (QoE) oriented data processing. To reflect the visual degradation introduced by various distortions, many PCQA metrics have been proposed in recent years. However, these metrics often take all distortions indiscriminately into account, ignoring the fact that some distortions are below the noticeable threshold and thus do not affect subjective perception. To solve this problem, we involve the characteristic of just noticeable difference (JND) into PCQA. Specifically, we first rotate the reference and distorted point cloud repeatedly to obtain multiple perspectives, then utilize the 2D JND models to derive the 3D noticeable possibility maps (NPM) to infer the possibility that the distortion is perceivable at each point. Utilizing the generated NPM, we modify current point-wise and structure-wise quality metrics to help them correlate better with subjective perception. Extensive experiments show the universal effectiveness of the proposed NPM in improving PCQA metrics. Code will be available at https://github.com/NekoooooOoi/NPM.
Qi Yang 0003, Yiling Xu, Jun Sun 0005, Shan Liu 0001
ICME6
2023 Point Cloud Quality Assessment using 3D Saliency Maps
abstract
Point cloud quality assessment (PCQA) has become an appealing research field in recent days. Considering the importance of saliency detection in quality assessment, we propose an effective full-reference PCQA metric which makes an attempt to utilize the saliency information to facilitate quality prediction, called point cloud quality assessment using 3D saliency maps (PQSM). Specifically, we first propose a projectionbased point cloud saliency map generation method, in which depth information is introduced to better reflect the geometric characteristics of point clouds. Then, we construct point cloud local neighborhoods to derive three structural descriptors to indicate the geometry, color and saliency discrepancies. Finally, a saliency-based pooling strategy is proposed to generate the final quality score. Extensive experiments are performed on four independent PCQA databases. The results demonstrate that the proposed PQSM shows competitive performances compared to multiple state-of-the-art PCQA metrics.
Qi Yang 0003, Yiling Xu, Jun Sun 0005, Shan Liu 0001
VCIP5
2023 MPED: Quantifying Point Cloud Distortion Based on Multiscale Potential Energy Discrepancy
abstract
In this article, we propose a new distortion quantification method for point clouds, the multiscale potential energy discrepancy (MPED). Currently, there is a lack of effective distortion quantification for a variety of point cloud perception tasks. Specifically, in human vision tasks, a distortion quantification method is used to predict human subjective scores and optimize the selection of human perception task parameters, such as dense point cloud compression and enhancement. In machine vision tasks, a distortion quantification method usually serves as loss function to guide the training of deep neural networks for unsupervised learning tasks (e.g., sparse point cloud reconstruction, completion, and upsampling). Therefore, an effective distortion quantification should be differentiable, distortion discriminable, and have low computational complexity. However, current distortion quantification cannot satisfy all three conditions. To fill this gap, we propose a new point cloud feature description method, the point potential energy (PPE), inspired by classical physics. We regard the point clouds are systems that have potential energy and the distortion can change the total potential energy. By evaluating various neighborhood sizes, the proposed MPED achieves global-local tradeoffs, capturing distortion in a multiscale fashion. We further theoretically show that classical Chamfer distance is a special case of our MPED. Extensive experiments show that the proposed MPED is superior to current methods on both human and machine perception tasks. Our code is available at https://github.com/Qi-Yangsjtu/MPED.
Qi Yang 0003, Siheng Chen, Yiling Xu, Jun Sun 0005, Zhan Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 No-Reference Point Cloud Quality Assessment via Domain Adaptation
abstract
We present a novel no-reference quality assessment metric, the image transferred point cloud quality assessment (IT-PCQA), for 3D point clouds. For quality assessment, deep neural network (DNN) has shown compelling performance on no-reference metric design. However, the most challenging issue for no-reference PCQA is that we lack large-scale subjective databases to drive robust networks. Our motivation is that the human visual system (HVS) is the decision-maker regardless of the type of media for quality assessment. Leveraging the rich subjective scores of the natural images, we can quest the evaluation criteria of human perception via DNN and transfer the capability of prediction to 3D point clouds. In particular, we treat natural images as the source domain and point clouds as the target domain, and infer point cloud quality via unsupervised adversarial domain adaptation. To extract effective latent features and minimize the domain discrepancy, we propose a hierarchical feature encoder and a conditional-discriminative network. Considering that the ultimate pur-pose is regressing objective score, we introduce a novel con-ditional cross entropy loss in the conditional-discriminative network to penalize the negative samples which hinder the convergence of the quality regression network. Experi-mental results show that the proposed method can achieve higher performance than traditional no-reference metrics, even comparable results with full-reference metrics. The proposed method also suggests the feasibility of assessing the quality of specific media content without the expensive and cumbersome subjective evaluations. Code is available at https://github.com/Qi-Yangsjtu/IT-PCQA.
Qi Yang 0003, Yipeng Liu 0003, Siheng Chen, Yiling Xu, Jun Sun 0005
CVPR5
2022 Mask-Guided Transformer for Human-Object Interaction Detection
abstract
Human-object interaction (HOI) detection is a meaningful research topic on human activity understanding. Recent works have made significant progress by focusing on efficient triplet matching and leveraging image-wide features based on encoder-decoder architecture. However, the ability to gather relevant contextual information about human is limited and different sub-tasks in HOI detection are not differentiated by specific decoupling in previous methods. To this end, we propose a new transformer-based method for HOI detection, namely, Mask-Guided Transformer (MGT). Our model, which is composed of five parallel decoders with a shared encoder, not only emphasizes interactive regions by applying body features, but also disentangles the prediction of instance and interaction. We achieve a favorable result at 63.3 mAP on the well-known HOI detection dataset V-COCO.
Daocheng Ying, Hua Yang 0001, Jun Sun 0005
VCIP3
2022 Inferring Point Cloud Quality via Graph Similarity
abstract
Objective quality estimation of media content plays a vital role in a wide range of applications. Though numerous metrics exist for 2D images and videos, similar metrics are missing for 3D point clouds with unstructured and non-uniformly distributed points. In this paper, we propose [Formula: see text]-a metric to accurately and quantitatively predict the human perception of point cloud with superimposed geometry and color impairments. Human vision system is more sensitive to the high spatial-frequency components (e.g., contours and edges), and weighs local structural variations more than individual point intensities. Motivated by this fact, we use graph signal gradient as a quality index to evaluate point cloud distortions. Specifically, we first extract geometric keypoints by resampling the reference point cloud geometry information to form an object skeleton. Then, we construct local graphs centered at these keypoints for both reference and distorted point clouds. Next, we compute three moments of color gradients between centered keypoint and all other points in the same local graph for local significance similarity feature. Finally, we obtain similarity index by pooling the local graph significance across all color channels and averaging across all graphs. We evaluate [Formula: see text] on two large and independent point cloud assessment datasets that involve a wide range of impairments (e.g., re-sampling, compression, and additive noise). [Formula: see text] provides state-of-the-art performance for all distortions with noticeable gains in predicting the subjective mean opinion score (MOS) in comparison with point-wise distance-based metrics adopted in standardized reference software. Ablation studies further show that [Formula: see text] can be generalized to various scenarios with consistent performance by adjusting its key modules and parameters. Models and associated materials will be made available at https://njuvision.github.io/GraphSIM or http://smt.sjtu.edu.cn/papers/GraphSIM.
Qi Yang 0003, Zhan Ma 0001, Yiling Xu, Zhu Li 0001, Jun Sun 0005
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Joint Optimization of Preamble Selection and Access Barring for Random Access in MTC With General Device Activities
abstract
Most existing random access schemes for machine-type communications (MTC) simply adopt a uniform preamble selection distribution, irrespective of the underlying device activity distributions. Hence, they may yield unsatisfactory access efficiency. In this paper, we model device activities for MTC as multiple Bernoulli random variables following an arbitrary multivariate Bernoulli distribution which can reflect both dependent and independent device activities. Then, we optimize preamble selection and access barring for random access in MTC according to the underlying joint device activity distribution. Specifically, we investigate three cases of the joint device activity distribution, i.e., the cases of perfect, imperfect, and unknown joint device activity distributions, and formulate the average, worst-case average, and sample average throughput maximization problems, respectively. The problems in the three cases are challenging nonconvex problems. In the case of perfect joint device activity distribution, we develop an iterative algorithm and a low-complexity iterative algorithm to obtain stationary points of the original problem and an approximate problem, respectively. In the case of imperfect joint device activity distribution, we develop an iterative algorithm and a low-complexity iterative algorithm to obtain a Karush-Kuhn-Tucker (KKT) point of an equivalent problem and a stationary point of an approximate problem, respectively. Finally, in the case of unknown joint device activity distribution, we develop an iterative algorithm to obtain a stationary point. The proposed solutions are widely applicable and outperform existing solutions for dependent and independent device activities.
Ying Cui 0001, Feng Yang 0006, Lianghui Ding, Jun Sun 0005
IEEE Trans. Commun.5
2022 Sequential Attention-Based Distinct Part Modeling for Balanced Pedestrian Detection
abstract
Despite pedestrian detectors having made significant progress by introducing convolutional neural networks, their performance still suffers degradation, especially in occlusion scenes with more false positives (FPs) and false negatives (FNs). To alleviate the problem, we propose a novel Sequential Attention-based Distinct Part Modeling (SA-DPM) for balanced pedestrian detection. It takes one step further in constructing more robust representation that supports detection with fewer FNs and FPs. Specifically, the Sequential Attention serves as one internal perception process that captures several distinct part areas step by step from each pedestrian proposal (full-body). Different from the previous either-or feature selection, the following Joint Learning attempts to seek a reasonable trade-off between part and full-body features, and combines both features for more accurate classification and regression. Evaluation on the widely used pedestrian datasets including Caltech and Citypersons shows that the proposed SA-DPM achieves promising performance for both non-occluded and occluded pedestrian detection tasks, especially on Caltech Heavy Occlusion set, which yields a new state-of-the-art miss rate by 30.18% and outperforms the second best detector by 6.32%.
Yan Luo 0003, Weiyao Lin, Xiaokang Yang 0001, Jun Sun 0005
IEEE Trans. Intell. Transp. Syst.5
2021 Point Cloud Geometry Compression Via Neural Graph Sampling
abstract
Compressing point cloud geometry (PCG) efficiently is of great interests for enabling abundant networked applications, because PCG is a promising representation to precisely illustrate arbitrary-shaped 3D objects and relevant physical scenes. To well exploit the unconstrained geometric correlation of input PCG, a three-step neural graph sampling (NGS) is developed. First, we construct the local graph of each point using its K nearest neighbors according to the Euclidean distance metric; Second, for each local graph, its graph center point expands associated feature attribute by aggregating neighbor weights via point-wise dynamic filter; We then perform attention-based sampling to select a subset of points to well represent input points. The proposed NGS is embedded into an end-to-end analysis/synthesis-based variational autoencoder (VAE), with which the encoder applies multiscale NGS to extract latent keypoints that are augmented with neighbor structures and compressed at bottleneck leveraging the hyperpriors for accurate entropy modeling, and the decoder directly uses layered convolutions to refine progressively for the reconstruction of final point cloud. Note that all computations are fulfilled using point-wise convolution, making our solution an attractive approach in practice. Experimental results demonstrate that the proposed method using NGS mechanism outperforms the state-of-the-art point-based PCG compression methods by more than $2\times \mathrm{B}\mathrm{D}$-Rate (Bjûntegaard Delta Rate) gains, and several orders of magnitude gains over the MPEG G-PCC across all testing categories on ShapeNetCorev2 dataset.
Linyao Gao, Tingyu Fan, Jianqiang Wan, Yiling Xu, Jun Sun 0005, Zhan Ma 0001
ICIP5
2021 A Dual Camera System for High Spatiotemporal Resolution Video Acquisition
abstract
This paper presents a dual camera system for high spatiotemporal resolution (HSTR) video acquisition, where one camera shoots a video with high spatial resolution and low frame rate (HSR-LFR) and another one captures a low spatial resolution and high frame rate (LSR-HFR) video. Our main goal is to combine videos from LSR-HFR and HSR-LFR cameras to create an HSTR video. We propose an end-to-end learning framework, AWnet, mainly consisting of a FlowNet and a FusionNet that learn an adaptive weighting function in pixel domain to combine inputs in a frame recurrent fashion. To improve the reconstruction quality for cameras used in reality, we also introduce noise regularization under the same framework. Our method has demonstrated noticeable performance gains in terms of both objective PSNR measurement in simulation with different publicly available video and light-field datasets and subjective evaluation with real data captured by dual iPhone 7 and Grasshopper3 cameras. Ablation studies are further conducted to investigate and explore various aspects, such as reference structure, camera parallax, exposure time, etc) of our system to fully understand its capability for potential applications.
Zhan Ma 0001, Muhammad Salman Asif, Yiling Xu, Wenbo Bao, Jun Sun 0005
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 Predicting the Perceptual Quality of Point Cloud: A 3D-to-2D Projection-Based Exploration
abstract
Point cloud is emerged as a promising media format to represent realistic 3D objects or scenes in applications, such as virtual reality, teleportation, etc. How to accurately quantify the subjective point cloud quality for application-driven optimization, however, is still a challenging and open problem. In this paper, we attempt to tackle this problem in a systematic means. First, we produce a fairly large point cloud dataset where ten popular point clouds are augmented with seven types of impairments (e.g., compression, photometry/color noise, geometry noise, scaling) at six different distortion levels, and organize a formal subjective assessment with tens of subjects to collect mean opinion scores (MOS) for all 420 processed point cloud samples (PPCS). We then try to develop an objective metric that can accurately estimate the subjective quality. Towards this goal, we choose to project the 3D point cloud onto six perpendicular image planes of a cube for the color texture image and corresponding depth image, and aggregate image-based global (e.g., Jensen-Shannon (JS) divergence) and local features (e.g., edge, depth, pixel-wise similarity, complexity) among all projected planes for a final objective index. Model parameters are fixed constants after performing the regression using a small and independent dataset previously published. The proposed metric has demonstrated the state-of-the-art performance for predicting the subjective point cloud quality compared with multiple full-reference and no-reference models, e.g., the weighted peak signal-to-noise ratio (PSNR), structural similarity (SSIM), feature similarity (FSIM) and natural image quality evaluator (NIQE). The dataset is made publicly accessible athttp://smt.sjtu.edu.cnorhttp://vision.nju.edu.cnfor all interested audiences.
Qi Yang 0003, Hao Chen 0036, Zhan Ma 0001, Yiling Xu, Rongjun Tang, Jun Sun 0005
IEEE Trans. Multim.6
2020 Multilevel Interaction Reasoning For Complex Event Recognition
abstract
Event as a complex process contains many factors. Objects, environments, and their interactions vary with time. Recognizing event remains a challenging task in computer vision. In this paper, a multilevel interaction reasoning framework is proposed for complex event recognition. Firstly, we construct a 3D ConvNet to extract the spatial-temporal feature to present global scene. Then a graph ConvNet is built to reasoning about the multilevel interaction: object-object and object-environment interaction, via graphs that contain object feature in video, and projection of global scene feature in 3D ConvNet. The proposed method effectively explore the nature of events occurring and developing, and experimental results on challenging UCF-Crime dataset achieve state-of-the-art with a 3.5% gain over other models.
Hua Yang 0001, Jun Sun 0005
ICIP3
2020 CSpA-DN: Channel and Spatial Attention Dense Network for Fusing PET and MRI Images
abstract
In this paper, we propose a novel fusion framework based on a dense network with channel and spatial attention (CSpA-DN) for PET and MR images. In our approach, an encoder composed of the densely connected neural network is constructed to extract features from source images, and a decoder network is leveraged to yield the fused image from these features. Simultaneously, a self-attention mechanism is introduced in the encoder and decoder to further integrate local features along with their global dependencies adaptively. The extracted feature of each spatial position is synthesized by a weighted summation of those features at the same row and column with this position via a spatial attention module. Meanwhile, the interdependent relationship of all feature maps is integrated by a channel attention module. The summation of the outputs of these two attention modules is fed into the decoder and the fused image is generated. Experimental results illustrate the superiorities of our proposed CSpA-DN model compared with state-of-the-art methods in PET and MR images fusion according to both visual perception and objective assessment.
Bicao Li, Zhoufeng Liu, Jenq-Neng Hwang, Jun Sun 0005, Zongmin Wang
ICPR5
2020 Fast Video Saliency Detection based on Feature Competition
abstract
In this paper, we propose a light video saliency prediction model, named SalFCM, which achieves fixation detection rate of 110fps. It is known that the human attention is captured by objects that have always been present or newly appeared. To model this dynamic change, we propose an Inter-frame Feature Competition Module (IFCM) to make an adaptive choice between correlated and differential features of consecutive frames. Besides, it is noted that saliency is better explained by low-level rather than high-level features in some visual scenes. Hence, we design a Hierarchical Feature Competition Module (HFCM) to balance the influence of low-level and high-level features. Our model achieves a good trade-off between precision and processing speed. The developed SalFCM is evaluated on three video saliency datasets: DHF1K, Hollywood-2 and UCF-sports. We conduct ablation studies to verify the effectiveness of the proposed model.
Hang Yan 0006, Yiling Xu, Jun Sun 0005, Le Yang 0001, Wei Huang 0012
VCIP3
2019 Safeguarded Dynamic Label Regression for Noisy Supervision
abstract
Learning with noisy labels is imperative in the Big Data era since it reduces expensive labor on accurate annotations. Previous method, learning with noise transition, has enjoyed theoretical guarantees when it is applied to the scenario with the class-conditional noise. However, this approach critically depends on an accurate pre-estimated noise transition, which is usually impractical. Subsequent improvement adapts the preestimation in the form of a Softmax layer along with the training progress. However, the parameters in the Softmax layer are highly tweaked for the fragile performance and easily get stuck into undesired local minimums. To overcome this issue, we propose a Latent Class-Conditional Noise model (LCCN) that models the noise transition in a Bayesian form. By projecting the noise transition into a Dirichlet-distributed space, the learning is constrained on a simplex instead of some adhoc parametric space. Furthermore, we specially deduce a dynamic label regression method for LCCN to iteratively infer the latent true labels and jointly train the classifier and model the noise. Our approach theoretically safeguards the bounded update of the noise transition, which avoids arbitrarily tuning via a batch of samples. Extensive experiments have been conducted on controllable noise data with CIFAR10 and CIFAR-100 datasets, and the agnostic noise data with Clothing1M and WebVision17 datasets. Experimental results have demonstrated that the proposed model outperforms several state-of-the-art methods.
Jiangchao Yao, Hao Wu 0075, Ya Zhang 0002, Ivor W. Tsang, Jun Sun 0005
AAAI5
2019 Deep Learning From Noisy Image Labels With Quality Embedding
abstract
There is an emerging trend to leverage noisy image datasets in many visual recognition tasks. However, the label noise among datasets severely degenerates the performance of deep learning approaches. Recently, one mainstream is to introduce the latent label to handle label noise, which has shown promising improvement in the network designs. Nevertheless, the mismatch between latent labels and noisy labels still affects the predictions in such methods. To address this issue, we propose a probabilistic model, which explicitly introduces an extra variable to represent the trustworthiness of noisy labels, termed as the quality variable. Our key idea is to identify the mismatch between the latent and noisy labels by embedding the quality variables into different subspaces, which effectively minimizes the influence of label noise. At the same time, reliable labels are still able to be applied for training. To instantiate the model, we further propose a Contrastive-Additive Noise network (CAN), which consists of two important layers: (1) the contrastive layer that estimates the quality variable in the embedding space to reduce the influence of noisy labels; and (2) the additive layer that aggregates the prior prediction and noisy labels as the posterior to train the classifier. Moreover, to tackle the challenges in optimization, we deduce an SGD algorithm with the reparameterization tricks, which makes our method scalable to big data.We validate the proposed method on a range of noisy image datasets. Comprehensive results have demonstrated that CAN outperforms the state-of-the-art deep learning approaches.
Jiangchao Yao, Ivor W. Tsang, Ya Zhang 0002, Jun Sun 0005, Chengqi Zhang
IEEE Trans. Image Process.5
2018 Selective Convolutional Features based Generalized-mean Pooling for Fine-grained Image Retrieval
abstract
Image retrieval with convolutional neural network (CNN) has obtained a lot of attention. In this paper, we focus on a more challenging task: fine-grained image retrieval. We propose a simple and effective feature aggregation method using generalized-mean pooling (GeM pooling), which can make better use of information from the output tensor of the convolutional layer. In addition, we propose a simple feature selection scheme to remove noise and background. Experimental results demonstrate that our aggregation method not only outperformed state-of-the-art aggregation methods for general image retrieval, but also reach up to the same level of existing aggregation method for fine-grained image retrieval, with more compact representation and less memory cost.
Zhuoqun Wang, Zhu Li 0001, Jun Sun 0005, Yiling Xu
VCIP3
2018 Joint Latent Dirichlet Allocation for Social Tags
abstract
Social tags, serving as a textual source of simple but useful semantic metadata to reflect the user preference or describe the web objects, has been widely used in many applications. However, social tags have several unique characteristics, i.e., sparseness and data coupling (i.e., non-IIDness), which makes existing text analysis methods such as LDA not directly applicable. In this paper, we propose a new generative algorithm for social tag analysis named joint latent Dirichlet allocation, which models the generation of tags based on both the users and the objects, and thus accounts for the coupling relationships among social tags. The model introduces two latent factors that jointly influence tag generation: the user's latent interest factor and the object's latent topic factor, formulated as user-topic distribution matrix and object-topic distribution matrix, respectively. A Gibbs sampling approach is adopted to simultaneously infer the above two matrices as well as a topic-word distribution matrix. Experimental results on four social tagging datasets have shown that our model is able to capture more reasonable topics and achieves better performance than five state-of-the-art topic models in terms of the widely used point-wise mutual information metric. In addition, we analyze the learnt topics showing that our model recovers more themes from social tags while LDA may lead the topic vanishing problems, and demonstrate its advantages in the social recommendation by evaluating the retrieval results with mean reciprocal rank metric. Finally, we explore the joint procedure of our model in depth to show the non-IID characteristic of social tagging process.
Jiangchao Yao, Yanfeng Wang 0001, Ya Zhang 0002, Jun Sun 0005, Jun Zhou 0007
IEEE Trans. Multim.4
2017 Discovering User Interests from Social Images
Jiangchao Yao, Ya Zhang 0002, Ivor W. Tsang, Jun Sun 0005
MMM (2)4
2017 Compact scalable hash from deep learning features aggregation for content de-duplication
abstract
Unprecedented growth in media content generation, communication and consumption has taken over the vast majority of storage spaces in devices, network caches, and clouds. How to identify duplications from network caches is an important issue for fast and efficient content delivery network (CDN) communication and storage. In this work, we developed a novel hash scheme which is scalable and robust to typical CDN induced transcoding and manipulations. Scalable hash design is constructed in essentially two stages: images are first represented as 512 channels of thumbnail images from the deep learning VGG-16 networks, and then a Fisher Vector aggregation is performed on the features which offer scalability in both underlying Gaussian Mixture Model (GMM) PCA embedding and component posterior likelihood. Hash is generated by direct binarizing the Fisher Vector with component/dimensionality priority optimization. Simulation results have demonstrated that this is a very compact and accurate scheme for CDN content de-duplication.
Shan Feng, Zhu Li 0001, Yiling Xu, Jun Sun 0005
MMSP4
2017 NDMP - An emerging MPEG standard for network distributed media processing
abstract
In this paper, we introduce a novel video codec and distribution standard developed by the Moving Picture Experts Group (MPEG), called network-distributed media processing (NDMP). Compared with existing video codec and distribution standards, which assume one encoder and one decoder, NDMP inserts a media processing unit, based on network entities, which implements media processing in a distributed manner, between the encoder and the decoder. NDMP has advantages in the following aspects:(1) it reduces the need of storage resources in the network. (2) it reduces the occupation of the bandwidth in backhaul network and (3) it reduces Round-Trip Time of remote media processing.
Yiling Xu, Jun Sun 0005, Jaeyeon Song, Kyungmo Park
VCIP3
2017 Visual discomfort prediction on stereoscopic 3D images without explicit disparities
Jun Zhou 0007, Jun Sun 0005, Alan C. Bovik
Signal Process. Image Commun.3
2017 A Convolutional Neural Network-Based Chinese Text Detection Algorithm via Text Structure Modeling
abstract
Text detection in a natural environment plays an important role in many computer vision applications. While existing text detection methods are focused on English characters, there are strong application demands on text detection in other languages, such as Chinese. In this paper, we present a novel text detection algorithm for Chinese characters based on a specific designed convolutional neural network (CNN). The CNN contains a text structure component detector layer, a spatial pyramid layer, and a multi-input-layer deep belief network (DBN). The CNN is pre-trained via a convolutional sparse auto-encoder, specifically designed for extracting complex features from Chinese characters. In particular, the text structure component detectors enhance the accuracy and uniqueness of feature descriptors by extracting multiple text structure components in various ways. The spatial pyramid layer enhances the scale invariability of the CNN for detecting texts in multiple scales. Finally, the multi-input-layer DBN replaces the fully connected layers in the CNN to ensure features from multiple scales are comparable. A multilingual text detection dataset, in which texts in Chinese, English, and digits are labeled separately, is set up to evaluate the proposed text detection algorithm. The proposed algorithm shows a significant performance improvement over the baseline CNN algorithms. In addition the proposed algorithm is evaluated over a public multilingual benchmark and achieves state-of-the-art result under multiple languages. Furthermore, a simplified version of the proposed algorithm with only general components is evaluated on the ICDAR 2011 and 2013 datasets, showing comparable detection performance to the existing general text detection algorithms.
Xiaohang Ren, Yi Zhou 0003, Jianhua He 0001, Kai Chen 0006, Xiaokang Yang 0001, Jun Sun 0005
IEEE Trans. Multim.6
2016 A novel text structure feature extractor for Chinese scene text detection and recognition
abstract
Scene text information extraction plays an important role in many computer vision applications. Unlike most existing text extraction algorithms for English texts, in this paper, we focus on Chinese texts, which are more complex in stroke and structure. To tackle this challenging problem, we propose a novel convolutional neural network (CNN) based text structure feature extractor for Chinese texts. Each Chinese character contains its specific types and combination of text structure components, which is rarely seen in backgrounds. Thus, different from the features only applicable to one text extraction stage (text detection or text recognition), the text structure component feature is suitable for both Chinese text detection and recognition. A text structure component detector (TSCD) layer is designed to detect the large amount of component types, which is the most challenging part of extracting text structure component features. Through statistical classification various types of text structure component are detected by their specially designed convolutional units in the TSCD layer. With the TSCD layer, the CNN has improvements in the accuracy and uniqueness of text feature description. In the evaluation, both text detection and recognition algorithms based on the proposed text structure feature extractor achieve state-of-the-art results in two datasets.
Xiaohang Ren, Kai Chen 0006, Xiaokang Yang 0001, Yi Zhou 0003, Jianhua He 0001, Jun Sun 0005
ICPR6
2016 A novel scene text detection algorithm based on convolutional neural network
abstract
Candidate text region extraction plays a critical role in convolutional neural network (CNN) based text detection from natural images. In this paper, we propose a CNN based scene text detection algorithm with a new text region extractor. The so called candidate text region extractor I-MSER is based on Maximally Stable Extremal Region (MSER), which can improve the independency and completeness of the extracted candidate text regions. Design of I-MSER is motivated by the observation that text MSERs have high similarity and are close to each other. The independency of candidate text regions obtained by I-MSER is guaranteed by selecting the most representative regions from a MSER tree which is generated according to the spatial overlapping relationship among the MSERs. A multi-layer CNN model is trained to score the confidence value of the extracted regions extracted by the I-MSER for text detection. The new text detection algorithm based on I-MSER is evaluated with wide-used ICDAR 2011 and 2013 datasets and shows improved detection performance compared to the existing algorithms.
Xiaohang Ren, Kai Chen 0006, Xiaokang Yang 0001, Yi Zhou 0003, Jianhua He 0001, Jun Sun 0005
VCIP6
2015 Joint Latent Dirichlet Allocation for non-iid social tags
abstract
Topic models have been widely used for analyzing text corpora and achieved great success in applications including content organization and information retrieval. However, different from traditional text data, social tags in the web containers are usually of small amounts, unordered, and non-iid, i.e., it is highly dependent on contextual information such as users and objects. Considering the specific characteristics of social tags, we here introduce a new model named Joint Latent Dirichlet Allocation (JLDA) to capture the relationships among users, objects, and tags. The model assumes that the latent topics of users and those of objects jointly influence the generation of tags. The latent distributions is then inferred with Gibbs sampling. Experiments on two social tag data sets have demonstrated that the model achieves a lower predictive error and generates more reasonable topics. We also present an interesting application of this model to object recommendation.
Jiangchao Yao, Ya Zhang 0002, Zhe Xu 0003, Jun Sun 0005, Jun Zhou 0007, Xiao Gu 0001
ICME4
2014 Binocular mismatch induced by luminance discrepancies on stereoscopic images
abstract
Luminance discrepancies between image pairs occur owing to inconsistent parameters between stereoscopic camera devices and from imperfect capture conditions. Such discrepancies induce binocular mismatches and affect the visual comfort that is felt by viewers, as well as their ability to fuse stereoscopic. To better understand and observe this effect, we built a stereoscopic images database of 240 luminance discrepancy images and 30 natural images with subjective scores of visual discomfort and fusion difficulty. Two features, binocular contrast and luminance similarity were extracted to analyze the relationship between the subjective scores and the luminance discrepancies. Structural dissimilarity and average luminance are used to predict the effects of binocular mismatches. The experimental results show that the combination of binocular contrast, structural dissimilarity and average luminance exhibits high consistency with subjective scores of visual discomfort, fusion difficulty and overall binocular mismatches in terms of Spearman's Rank Ordered Correlation Coefficient.
Jun Zhou 0007, Jun Sun 0005, Alan C. Bovik
ICME3
2014 Analysis and optimization of x265 encoder
abstract
x265 is an open-source encoder project which aims to deliver the world's fastest and most computationally efficient HEVC encoder. Although x265 has been developed efficiently with many optimization techniques, it is still not able to encode HD videos in real time even at its faster setting. In this paper, we deeply investigate the encoding framework and computational complexity of x265, and find that RDO process is the most time consuming part. Then, an efficient prediction scheme is proposed which includes decreasing the number of RDO times, early skip detection and fast intra mode decision. Experimental results show that the proposed method improves the speed of x265 from 19.86fps to 37.76fps for HD test sequences, i.e., 47.44% complexity reduction, with only 1.37% BDBR coding performance loss.
Qiang Hu 0003, Xiaoyun Zhang 0001, Jun Sun 0005
VCIP4
2014 Query-expanded collaborative representation based classification with class-specific prototypes for object recognition
Jun Zhou 0007, Jun Sun 0005
Pattern Recognit.3
2014 On Two-Stage Sequential Coding of Correlated Sources
abstract
We study the problem of two-stage sequential coding (TSSC), which is an extension of sequential coding of correlated sources. Let X and Y be dependent random variables. The network contains two encoders and two decoders: 1) a Y encoder with input Y; 2) an X encoder with inputs X and Y; 3) a Y decoder that reconstructs Y; and 4) an X decoder that reconstructs X. The first stage is traditional sequential coding, where the Y encoder describes Y to both the X decoder and Y decoder, and the X encoder describes X and Y to the X decoder. At the second stage, the Y encoder refines the description of Y, and the X encoder refines the description of X. The TSSC model is a theoretical abstraction of scalable video coding; here, Y and X represent successive frames of a video sequence, and the two stages together give an embedded description that allows the video to be decoded at two distinct rates. We give an inner bound on the rate distortion region for this TSSC model. The tight bound on the rate distortion region is derived when Y must be reconstructed losslessly (in the usual Shannon sense) in the second stage. We also study the minimum total rate of the TSSC model and show that the minimum total rate of one-stage sequential coding cannot be achieved at both stages for jointly Gaussian sources. This theoretical result can shed light on the rate-distortion performance behavior of scalable video coding widely noted by practitioners.
Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu
IEEE Trans. Inf. Theory3
2013 Foreground detection: Combining background subspace learning with object smoothing model
abstract
Foreground detection is a challenging problem in complex scenes. In this paper, a novel foreground detection method is proposed which combines background subspace learning with object smoothing model. Considering background scenes in consecutive frames are almost the same, they are approximated using an efficient subspace learning technique which is based on 2D images. Due to the pixels of objects are usually clustered, an object smoothing model is adopted where a spatial smoothing constraint is imposed on its values during the estimation, and then it can be solved as a regularized matrix restoration problem with a spatial smoothing constraint. As a result, isolated noises can be suppressed while clustered foreground pixels can be preserved. We test our method on some challenging sequences and compare it with some other techniques. Experimental results show its effectiveness and robustness.
Gengjian Xue, Li Song 0001, Jun Sun 0005, Jun Zhou 0007
ICME3
2013 Face Recognition Using Multi-scale ICA Texture Pattern and Farthest Prototype Representation Classification
Jun Zhou 0007, Jun Sun 0005
MMM (2)3
2013 Face recogntion in open world environment
abstract
Face recognition in open world environment is a very challenging task due to variant appearances of the target persons and a large scale of unregistered probe faces. In this paper we combine two parallel classifiers, one based on the Local Binary Pattern (LBP) feature and the other based on the Gabor features, to build a specific face recognizer for each target person. Faces used for training are borderline patterns obtained through a morphing procedure combing target faces and random non-target ones. Grid-search is applied to find an optimal morphing-degree-pair. By using an AND operator to integrate the prediction of the two complementary parallel classifiers, many false positives are eliminated in the final results. The proposed algorithm is compared with the Robust Sparse Coding method, using selected celebrities as the target persons and the images from FERET as the non-target faces. Experimental results suggest that the proposed approach is better at tolerating the distortion of the target person's appearance and has a lower false alarm rate.
Jieqiong Qiu, Ya Zhang 0002, Jun Sun 0005
VCIP3
2013 On the Generalization of Natural Type Selection to Multiple Description Coding
abstract
Natural type selection was originally proposed by Zamir and Rose for universal single description coding. In this paper, we generalize this principle to universal multiple description coding (MDC). Two schemes based on random codebooks are proposed: one is of fixed distortion and the other is of fixed weight. Their operational sum-rate-distortion functions are derived, which coincide with the EGC (El Gamal-Cover) sum-rate bound if the parameters of the schemes are optimized. It is also shown that in both schemes the joint type of reconstruction codewords can be used to improve the rate-distortion (R-D) performance. Based on our theoretical results, a practical universal scheme is proposed by leveraging the MDC methods based on low-density generator matrix (LDGM) codes. The performance of this scheme is compared experimentally with the EGC bound, which shows its effectiveness.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Jun Chen 0005
IEEE Trans. Commun.3
2013 Foreground Estimation Based on Linear Regression Model With Fused Sparsity on Outliers
abstract
Foreground detection is an important task in computer vision applications. In this paper, we present an efficient foreground detection method based on a robust linear regression model. First, a novel framework is proposed where foreground detection has been cast as an outlier signal estimation problem in a linear regression model. We regularize this problem by imposing a so-called fused sparsity constraint, which encourages both sparsity and smoothness of vector coefficients, on the outlier signal. Second, we convert this outlier signal estimation problem into an equivalent Fused Lasso problem, and then use existing solutions to obtain an optimized solution. Third, a new foreground detection method is presented to apply this new model to the 2-D image domain by merging the results from different vectorizations. Experiments on a set of challenging sequences show that the proposed method is not only superior to many state-of-the-art techniques, but also robust to noise.
Gengjian Xue, Li Song 0001, Jun Sun 0005
IEEE Trans. Circuits Syst. Video Technol.3
2012 Learning a Mahalanobis distance metric via regularized LDA for scene recognition
abstract
Constructing a suitable distance metric for scene recognition is a very challenging task due to the huge intra-class variations. In this paper, we propose a novel framework for learning a full parameter matrix in Mahalanobis metric, where the learning process is formulated as a non-negatively constrained minimization problem in a projected space. To fully capture the structure of scenes, we first apply multiple regularized linear discriminant analysis (LDA) to form a candidate projection pool. Second, we adopt the pairwise squared differences of the projected samples as the learning instances. Finally, the diagonal selection matrix is learned through least squares with non-negative L2-norm regularization. Experiments on two datasets in scene recognition show the effectiveness and efficiency of our approach.
Jun Zhou 0007, Jun Sun 0005
ICIP3
2012 Background subtraction based on phase feature and distance transform
Gengjian Xue, Jun Sun 0005, Li Song 0001
Pattern Recognit. Lett.2
2012 Multiple Description Image Coding Based on Delta-Sigma Quantization With Rate-Distortion Optimization
abstract
Recently, Østergaard and Zamir revealed the connection between multiple description coding and delta-sigma quantization (DSQ). The principle has been applied to image coding, with main focus on the framework where each block is processed separately. In this brief, we propose a two-channel multiple description image coding scheme that performs inter-block processing. The source image is first rearranged into a block sequence. Then, vector DSQ is performed with a bank of noise-shaping filters. Their coefficients as well as the quantization steps are chosen by a rate-distortion optimization algorithm. A post-processing algorithm is proposed for side decoding. Experiment results show the improvement achieved by the proposed scheme in terms of both peak signal-to-noise ratio values and subjective quality.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005
IEEE Trans. Image Process.3
2011 On Natural Type Selection in Universal Multiple Description Coding
abstract
In this paper, we generalize the concept of natural type selection, initially proposed by R.Zamir et.al., to multiple description coding. We prove results that are parallel to those of adaptive single description coding. A universal multiple description coding scheme is then proposed. The scheme codes a source vector at a time, and updates its codebooks based on the joint type of the codewords reconstructed previously. Based on our theoretical results, it can be shown that the rate-distortion performance of the proposed scheme gradually improves as coding proceeds.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005
DCC3
2011 Foreground estimation based on robust linear regression model
abstract
Background subtraction is a basic task for many computer vision applications, yet in dynamic scenes it is still a challenging problem. In this paper, we propose a new method to deal with this difficulty. Our approach is based on robust linear regression model and casts background subtraction as a outlier signal estimation problem. In our linear regression model, we explicitly model the error term as a combination of two components: foreground outlier and background noise. The foreground outlier is sparse and can be arbitrarily large in most cases, while the background noise is relatively small and dispersed. In order to reliably estimate the coefficients under the constraint of sparse foreground outlier, we propose a new objective function. Then we transform the function to fit our problem by only estimating the foreground outlier and give the solution method. Experimental results demonstrate the effectiveness of our method.
Gengjian Xue, Li Song 0001, Jun Sun 0005
ICIP3
2011 Adaptive fast DIRECT mode decision algorithm using mode and Lagrangian cost prediction for B frame in H.264/AVC
abstract
In this paper, a fast spatial DIRECT mode decision method for B frame in H.264/AVC is proposed. It is based on a statistical analysis on multiple video sequences, and the strong relationship of mode selection and rate-distortion (RD) cost between the current DIRECT macroblock (MB) and the co-located MBs is observed. With the check of mode condition and adaptive threshold of RD cost, the complex mode decision process can be released at an early stage even for small QP cases. Simulation results demonstrate the proposed method can achieve much better performance than the original exhaustive rate-distortion optimization (RDO) based mode decision algorithm by reducing up to 57.1 % of motion estimation (ME) time for IBPBP picture group with only negligible bit increment and quality degradation.
Xiaocong Jin, Jun Sun 0005, Jun Zhou 0007, Yiqing Huang 0002, Takeshi Ikenaga
ICME2
2011 Hybrid center-symmetric local pattern for dynamic background subtraction
abstract
Effective foreground detection in dynamic scenes is a challenging task in computer vision applications. In this paper, we propose a novel background modeling method to tackle this problem. First, we propose a second-order center-symmetric local derivative pattern (CS-LDP) which extracts more detail information compared with the first-order center-symmetric local binary pattern (CS-LBP). Then by concatenating the CS-LBP and CS-LDP histograms, a new hybrid histogram feature is presented. The length of this histogram is much shorter than the local binary pattern (LBP) histogram. Based on this hybrid feature, a novel background modeling method is proposed where the pixel process is modeled with a group of adaptive hybrid histograms. The major advantage of our method is its low complexity. Experiments on three challenging sequences demonstrate that the proposed method is effective and fast, producing comparable results to state-of-art algorithm while reducing the computation time greatly.
Gengjian Xue, Li Song 0001, Jun Sun 0005
ICME3
2011 Distributed Multiple Description Video Coding on Packet Loss Channels
abstract
In this paper, we are to solve the drift problem of multiple description video coding on packet loss channels by using state-of-the-art distributed techniques. We first present an asymptotically optimal code design of multiple descriptions in the Wyner-Ziv (MDWZ) setting. Then we propose a distributed multiple description video coding (DMDVC) scheme, which performs MDWZ coding on each nonintra coded frame. Instead of the prediction loops used in traditional multiple description video coding, Slepian-Wolf based coding is used to exploit interframe correlations. A bitplane extraction scheme is proposed to improve the balance between two descriptions, so that side informations can be interchanged between the side decoders of DMDVC with negligible quality degradation, which is crucial to robust transmission over packet loss channels. Experiment results demonstrate the robustness of our scheme, especially at high packet loss rates.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005
IEEE Trans. Image Process.3
2010 Backward Adaptive Pre/Post-Filtered DPCM with Near-Optimal Rate-Distortion Performance
abstract
In this paper, we propose two backward adaptive coding algorithms based on a recently invented pre/post-filtered DPCM (Differential Pulse-Coded Modulation) codec. The pre/post-filters and the predictor are adapted jointly. Source statistics are assumed unknown a priori. One of the algorithms is based on power spectrum estimation, and the other gradient descent. Some properties of the algorithms are analyzed. Experiment results show that both algorithms achieve near-optimal rate-distortion performance, significantly outperforming the adaptive DPCM without adaptive pre/post-filtering.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Rong Xie 0004
ICC3
2010 A multiple description codec based on combinatorial optimization and its application to image coding
abstract
We propose a general N-channel MDC (Multiple Description Coding) framework which can integrate the advantages of various low-dimensional MDC schemes. Given the operational rate-distortion functions of low-dimensional codecs, we show how to optimize the proposed MDC framework. We prove that the optimization problem can be reduced to a combinatorial problem which in certain cases admits solutions. We then apply the proposed optimization algorithm to multiple description image coding. Experiment results show the effectiveness of our approach.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Cheng Zhi
ICIP3
2010 Texture-based color constancy using local regression
abstract
Color constancy endows the machines with the ability of identifying the color regardless of the illuminant. Considering none of single algorithms available is universal, this paper presents a novel combination approach based on local texture features and local regression. To better represent images, local texture features based on integrated Weibull distribution are firstly extracted on the overlapping patches of the images. Then we define a new image distance metric to search for K most similar images of the test image. Incorporating a priori knowledge into the data-driven strategy, we finally combine individual algorithms using a local penalized regression according to the frequency ratio of best single algorithms. Experiment on a widely used dataset shows that the proposed approach outperforms the state-of-the-art single algorithms as well as popular combination approaches.
Jun Zhou 0007, Jun Sun 0005, Gengjian Xue
ICIP3
2010 Background subtraction based on phase and distance transform under sudden illumination change
abstract
Effective foreground detection under sudden illumination change is an active research topic. However, most existing background subtraction approaches, which are intensity based, fail to handle this situation. In this paper, we propose a novel background modeling method that overcomes this limitation by relying on statistical models which use pixel phase instead of intensities. We first extract the phase feature of the pixel using Gabor filters. Then, a phase based background subtraction approach is proposed. In this approach, each phase feature is modeled independently by a mixture of Gaussian models and updated with a novel scheme. Since foreground pixels are scattered in the preliminary detection result, distance transform is implemented on the binary image which transforms the image into a distance map. We segment the distance image with a threshold and get the final result. Experiments on two challenging sequences demonstrate the effectiveness and robustness of our method.
Gengjian Xue, Jun Sun 0005, Li Song 0001
ICIP2
2010 Dynamic background subtraction based on spatial extended center-symmetric local binary pattern
abstract
Moving objects detection in dynamic scenes is a challenging task in many computer vision applications. Traditional background modeling methods do not work well in these situations since they assume a nearly static background. In this paper, a novel operator named spatial extended center-symmetric local binary pattern (SCS-LBP) for background modeling is proposed. It extracts spatial and temporal information simultaneously while has low complexity compared to the local binary pattern (LBP) operator. Then combining this operator with an improved temporal distribution estimation scheme, we propose a new background subtraction method. In our method, each pixel is modeled by a group of adaptive SCS-LBP histograms, which provides us with many advantages compared to traditional ones. Experimental results demonstrate the effectiveness and robustness of our method.
Gengjian Xue, Jun Sun 0005, Li Song 0001
ICME2
2009 An improved block size selection method based on macroblock movement characteristic
Jun Sun 0005, Rong Xie 0004, Songyu Yu, Wenjun Zhang 0001
Multim. Tools Appl.2
2008 A Novel Multiple Description Video Codec Based on Slepian-Wolf Coding
abstract
One major task in multiple description video coding is to prevent drift on packet loss channels, where transmission errors occur in each description. We propose a distributed multiple description video coding (DMDVC) scheme excluding any prediction loops. The new codec suffers from no drift problem. In the two-channel mode of symmetry side informations (SI), one side decoder can use the SI of the other without any decoding quality degradation. Thereby, DMDVC achieves high robustness on packet loss channels.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Peng Wang 0026, Songyu Yu
DCC3
2008 Variable block size selection for a transcoder based on MB movement information
abstract
To improve coding performance, the new video compression standard H.264 employs seven variable block sizes for one Macro Block (MB) to conduct motion estimation and compensation. MPEG-2 only has one size 16×16. This paper presents a novel fast variable block size selection method for an inter-MB in a video transcoder from MPEG-2 to H.264 with downscaling by a factor two in each dimension based on the MB motion information. Without conducting motion re-estimation, an optimal block size is decided. Experiment results show that this method saves the transcoder complexity dramatically with little compression performance degradation.
Jun Sun 0005, Rong Xie 0004, Shibao Zheng, Songyu Yu
ICME2
2008 A projection method for derivation of non-Shannon-type information inequalities
abstract
In 1998, Zhang and Yeung found the first unconditional non-Shannon-type information inequality. Recently, Dougherty, Freiling and Zeger gave six new unconditional non-Shannon-type information inequalities. This work generalizes their work and provides a method to systematically derive non-Shannon-type information inequalities. An application of this method reveals new 4-variable non-Shannon-type information inequalities.
Weidong Xu, Jia Wang 0004, Jun Sun 0005
ISIT3
2007 On Multi-Stage Sequential Coding of Correlated Sources
abstract
We study the problem of multi-stage sequential coding (MSSC), which is an extension of sequential coding of correlated sources. Consider two correlated random variables X and Y to be coded in two stages. The first stage is sequential coding as referred to in the existing literature. At the second stage, the Y encoder refines the information of Y without any knowledge of X, and X encoder refines the information of X with the knowledge of Y, while all previous outputs are known at the decoder. As the sequential coding problem provides a theoretical abstraction of video coding, the MSSC model is a theoretical abstraction of scalable video coding, which is an important application of network communications. We give an achievable region for the MSSC system. The given achievable region is tight when Y is required to be reconstructed perfectly in the usual Shannon sense at the second stage. We also study the minimum total rate MSSC problem, and derive the minimum total rate for Gaussian sources. This result disproves the possibility that the minimum total rate of one stage sequential coding can be achieved at both stages even for correlated Gaussian sources. Thus we offer a theoretical explanation for the performance loss of scalable video coding widely noted by practitioners
Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu
DCC3
2007 Ringing Artifact Reduction for JPEG2000 Images
Jinyong Fang, Jun Sun 0005
ICIC (3)2
2007 Rate-distortion Optimized Trellis-Coded Quantization
abstract
In this paper, a rate-distortion optimized trellis-coded quantization (RDOTCQ) algorithm is presented. Based on rate-distortion criterion, the proposed algorithm improves coding efficiency by adaptively optimizing quantization levels of small signals. In addition, it has no overhead and is fully compatible with the standard decoder. The proposed algorithm can be applied to any TCQ-based systems. In particular, the algorithm has been verified on the platform of the quadtree classified and trellis coded quantized (QTCQ) wavelet image compression system. Experimental results show that the proposed algorithm has a better rate-distortion performance than the QTCQ and some other compression methods.
Jun Sun 0005, Jia Wang 0004
ICME2
2007 A Rate-distortion Based Quantization Level Adjustment Algorithm in Block-based Video Compression
abstract
In this paper, a rate-distortion based quantization level adjustment (RDQLA) algorithm is presented. Based on rate-distortion criterion, the quantization level adjustment algorithm effectively improves coding efficiency by adaptively optimizing quantization levels of the signals near the boundaries of quantization cells and adjusting quantization levels per block. The proposed algorithm can be applied in any block based image and video coding method. In addition, it has no overhead and is fully compatible with the existing compression standards. The algorithm has been verified on the platform of H.263 and H.264. Experimental results show that the proposed algorithm improves the performance substantially. It is shown that the proposed algorithm has a gain of 1 dB comparing with the newest H.264 reference code and more than 1 dB comparing with H.263 standard for high bit rates.
Jun Sun 0005, Jia Wang 0004, Xiaokang Yang 0001
ICME2
2007 Fast Mode Decision for H.264 Video Encoder Based on MB Motion Characteristic
abstract
Seven variable block sizes are adopted for inter-frame MB (macroblock) coding in H.264. This new feature achieves significant coding gain. However, the computation complexity of the mode decision is extremely high when RD (rate distortion) algorithm is used. In this paper, we propose a fast mode decision algorithm with fast coding block size selection based on MB motion characteristic for inter frames. In this algorithm, the residual_MB is first gotten through motion estimation for the MB. Then, the motion characteristic of the MB is extracted through careful analysis of the texture in the residual_MB. At last, the optimization mode for the MB is decided according to the motion characteristic. Experimental results show that the mode decision can save significantly computational complexity at cost of a little degradation of RD performance.
Jun Sun 0005, Yunqiang Liu
ICME2
2007 An OWE-based Algorithm for Line Scratches Restoration in Old Movies
abstract
Line scratch is the primary artifact in old films. In this paper a new algorithm for line scratch detection and removal is proposed. First, we establish an effective model to represent line scratches in the domain of over-complete wavelet expansion (OWE), which offers more precise position description for scratches than traditional downsampled wavelet transform. We then use it to detect and locate line scratches accurately. After that, an adaptive restoration method is adopted to remove line scratches by replacing the wavelet coefficients in line scratches of each scale with new interpolated wavelet coefficients of corresponding width computed from the line scratch model. Experiments show that the proposed method can detect more line artifacts with less false detection and remove the line scratches effectively as compared with the classic algorithm.
Jinghuo Guan, Jun Sun 0005, Guangtao Zhai, Zhengguo Li
ISCAS4
2007 Multiple Descriptions with Side Informations Also Known At the Encoder
abstract
We propose a new scheme of multiple descriptions with side information (SI). The two side decoders of the system use two different SI streams. Both SI streams are available to the central decoder and to the encoder. We give an inner bound for this system for general source and SI. The tight bound is obtained for the quadratic Gaussian case. This result is compared with our previous result of the MDWZ (multiple descriptions in the Wyner-Ziv setting) problem in which none of the side information is known at the encoder. It is shown that when side information is absent at the encoder, there is a performance loss. The proposed scheme and its achievable region have practical significance. It offers theoretical insight into the multiple description video coding (MDVC) and suggests an optimal coding strategy.
Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005
ISIT4
2006 Image Resolution Scaling with Arbitrary Unequal Ratios in each Direction in the DCT Domain
abstract
To meet with the different client end devices and network bandwidth transmission requirements, it is necessary to convert the high definition resolution images and videos to standard definition resolution format. A novel approach to convert image resolution with arbitrary ratios in the DCT (discrete cosine transform) domain is proposed, which exploits the relationship of a block and its subblocks with differing size. It can realize arbitrary unequal ratios in the horizontal and vertical direction, respectively. Unlike the present methods, it does not need upsampling-downsampling process. It can perform downsizing directly to the original data and confirms to the standard decoder. The proposed approach is computationally fast and memory efficient and produces visually better images with higher PSNR compared to the spatial methods
Kebin An, Jun Sun 0005
ICASSP (3)2
2006 Multiple Descriptions in the Wyner-Ziv Setting
abstract
We propose a new scheme of multiple descriptions in the Wyner-Ziv setting (MD-WZ). The two side decoders of MD-WZ use two different side information (SI) streams. Both SI streams are available to the central decoder, but none to the encoder. We derive an achievable region (inner bound) for this MD-WZ system for general source and SI. If the source and SI are correlated Gaussian and for quadratic distortion metric, the tight bound is obtained. Our result is an extension of Ozarow's result on multiple descriptions of Gaussian source without SI. The MD-WZ coding scheme is shown to have a property of practical significance. For symmetric case where the joint distributions of the source and the two SI are the same and the two channels are balanced, interchanging the two channels causes no performance loss for Gaussian source. Considering that the existing multi-description video coding methods suffer from the notorious drifting problem induced by channels interchange, this work lends a theoretical support to distributed multi-description video coding in the Wyner-Ziv setting
Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005
ISIT4
2006 Transform domain transcoding from MPEG-2 to H.264 with interpolation drift-error compensation
abstract
With the increasingly extensive applications of the new emerging video coding standard H.264, it inspires an urgent need to transcode the widely available MPEG-2 compressed video to H.264 format. In this paper, we investigate the issues on transcoding MPEG-2 into H.264 in transform domain with consideration of drift error due to the mismatch of motion compensation, and propose a transform domain solution to transcode MPEG-2 into H.264. We first analyze two major kinds of drifting error resulting from the mismatch of motion compensation: interpolation error and quantization error. The former is caused by the difference between the interpolation filters adopted in these two standards, thus very unique to the transform domain transcoding from MPEG-2 to H.264. Furthermore, it is identified as the dominant factor of the video quality degradation, especially in the case of small quantization parameter at high bit-rate, by extensive experimental results. As a major contribution of this paper, the close form of interpolation error is derived from transform domain. We then proposed the transcoding scheme based on quantization error drifting compensation and Interpolation error drifting compensation. Experimental results show that the proposed transform domain transcoding scheme achieves very promising performance in terms of low computational complexity and high transcoded video quality. Most importantly, its peak signal-to-noise ratio is very close to the cascaded transcoding architecture with time-consuming decoding and recoding process in pixel domain.
Tuanjie Qian, Jun Sun 0005, Xiaokang Yang 0001, Jia Wang 0004
IEEE Trans. Circuits Syst. Video Technol.2
2003 1-D and 2-D transforms from integers to integers
abstract
Substituting a real valued linear transform with an integer-to-integer mapping has become very important in lots of applications. This paper introduces a new kind of matrix decomposition method called lifting-like factorization, which leads to a theorem: every 2/sup n/-order real matrix with determinant norm 1 can be expressed as the product of one permutation matrix and at most three unit triangular matrices. Rounding error of this method is analyzed. Realization of 2D integer transform is also studied and it is shown that a 2D integer-to-integer transform cannot be realized by performing two 1D integer transforms separately. Left and right permutation matrices are introduced to reduce rounding error and an application of this method to intDCT is discussed.
Jia Wang 0004, Jun Sun 0005, Songyu Yu
ICASSP (2)2
2002 Modified wavelet coding of arbitrarily shaped objects based on extrapolation and reflection (EAR)
Jia Wang 0004, Jun Sun 0005, Songyu Yu
VCIP2
2002 Wavelet image coding based on directional dilation
Jia Wang 0004, Songyu Yu, Jun Sun 0005
VCIP3
2000 Design and implementation of the second-generation HDTV prototype video encoder of China
Jun Sun 0005, Zhenghua Yu, Songyu Yu
VCIP1