VLDB 2026 Research / reviewers in the wild / expert
Deming Zhai
dblp:69/8937
· DBLP profile ↗
58ranked-venue papers
11as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 12 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSystems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Variation-Bounded Loss for Noise-Tolerant LearningabstractMitigating the negative impact of noisy labels has been a perennial issue in supervised learning. Robust loss functions have emerged as a prevalent solution to this problem. In this work, we introduce the Variation Ratio as a novel property related to the robustness of loss functions, and propose a new family of robust loss functions, termed Variation-Bounded Loss (VBL), which is characterized by a bounded variation ratio. We provide theoretical analyses of the variation radio, proving that a smaller variation ratio would lead to better robustness. Furthermore, we reveal that the variation ratio provides a feasible method to relax the symmetric condition and offers a more concise path to achieve the asymmetric condition. Based on the variation ratio, we reformulate several commonly used loss functions into a variation-bounded form for pract ical applications. Positive experiments on various datasets exhibit the effectiveness and flexibility of our approach. Jialiang Wang 0003, Xianming Liu 0005, Gangfeng Hu, Deming Zhai, Junjun Jiang, Haoliang Li |
AAAI | 5 |
| 2026 | SGCNeRF: Few-Shot Neural Rendering via Sparse Geometric Consistency GuidanceabstractNeural Radiance Field (NeRF) technology has made significant strides in creating novel viewpoints. However, its effectiveness is hampered when working with sparsely available views, often leading to performance dips due to overfitting. FreeNeRF attempts to overcome this limitation by integrating implicit geometry regularization, which incrementally improves both geometry and textures. Nonetheless, an initial low positional encoding bandwidth results in the exclusion of high-frequency elements. The quest for a holistic approach that simultaneously addresses overfitting and the preservation of high-frequency details remains ongoing. This study presents a novel feature-matching-based sparse geometry regularization module, enhanced by a spatially consistent geometry filtering mechanism and a frequency-guided geometric regularization strategy. This module excels at accurately identifying high-frequency keypoints, effectively preserving fine structural details. Through progressive refinement of geometry and textures across NeRF iterations, we unveil an effective few-shot neural rendering architecture, designated as SGCNeRF, for enhanced novel view synthesis. Our experiments demonstrate that SGCNeRF not only achieves superior geometry-consistent outcomes but also surpasses FreeNeRF, with improvements of 0.7 dB in PSNR on LLFF and DTU. Yuru Xiao, Xianming Liu 0005, Deming Zhai, Kui Jiang, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | 3D-SLARM: Practical Lossless Volumetric Image Compression via a 3D-Scanning Lightweight Autoregressive ModelabstractVolumetric images often encapsulate critical information, making it essential to employ lossless compression to preserve data integrity. Although various learned methods have demonstrated effective lossless compression for volumetric images, balancing high compression ratios with rapid coding speeds and lightweight architectures remains challenging. In this paper, we propose a 3D-scanning lightweight autoregressive model (3D-SLARM) for practical lossless volumetric image compression. 3D-SLARM integrates a novel 3D plane scanning module, a lightweight feature extraction (FE) module, and a lightweight distribution parameter and adaptive range predictor (DPARP) module. Initially, 3D-SLARM leverages a 3D plane scanning module to determine the scanning order of each voxel, allowing parallel coding of voxels within the same plane. Next, the lightweight FE module captures both intra-slice and inter-slice dependencies in the receptive field defined by the 3D plane scanning module. By incorporating our proposed serial re-parameterization (SerRep) technology alongside non-centric masked convolution (NCMC), the FE module attains a lightweight design while effectively capturing complex dependencies. Finally, 3D-SLARM employs a lightweight DPARP module to compute distribution parameters for both 8-bit and high bit-depth volumetric images. For high bit-depth images, the module further generates an adaptive probability range for each voxel, resulting in compact, voxel-specific PMF tables that facilitate efficient compression. Extensive experiments demonstrate that our 3D-SLARM achieves state-of-the-art lossless compression performance on majority volumetric image datasets and maintains fast coding speed with a lightweight design, underscoring its practical applicability. Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Image Process. | 4 |
| 2025 | Spatial Annealing for Efficient Few-shot Neural RenderingabstractNeural Radiance Fields (NeRF) with hybrid representations have shown impressive capabilities for novel view synthesis, delivering high efficiency. Nonetheless, their performance significantly drops with sparse input views. Various regularization strategies have been devised to address these challenges. However, these strategies either require additional rendering costs or involve complex pipeline designs, leading to a loss of training efficiency. Although FreeNeRF has introduced an efficient frequency annealing strategy, its operation on frequency positional encoding is incompatible with the efficient hybrid representations. In this paper, we introduce an accurate and efficient few-shot neural rendering method named Spatial Annealing regularized NeRF (SANeRF), which adopts the pre-filtering design of a hybrid representation. We initially establish the analytical formulation of the frequency band limit for a hybrid architecture by deducing its filtering process. Based on this analysis, we propose a universal form of frequency annealing in the spatial domain, which can be implemented by modulating the sampling kernel to exponentially shrink from an initial one with a narrow grid tangent kernel spectrum. This methodology is crucial for stabilizing the early stages of the training phase and significantly contributes to enhancing the subsequent process of detail refinement. Our extensive experiments reveal that, by adding merely one line of code, SANeRF delivers superior rendering quality and much faster reconstruction speed compared to current few-shot neural rendering methods. Notably, SANeRF outperforms FreeNeRF on the Blender dataset, achieving 700X faster reconstruction speed. Yuru Xiao, Deming Zhai, Wenbo Zhao 0004, Kui Jiang, Junjun Jiang, Xianming Liu 0005 |
AAAI | 2 |
| 2025 | Joint Asymmetric Loss for Learning with Noisy LabelsabstractLearning with noisy labels is a crucial task for training accurate deep neural networks. To mitigate label noise, prior studies have proposed various robust loss functions, particularly symmetric losses. Nevertheless, symmetric losses usually suffer from the underfitting issue due to the overly strict constraint. To address this problem, the Active Passive Loss (APL) jointly optimizes an active and a passive loss to mutually enhance the overall fitting ability. Within APL, symmetric losses have been successfully extended, yielding advanced robust loss functions. Despite these advancements, emerging theoretical analyses indicate that asymmetric losses, a new class of robust loss functions, possess superior properties compared to symmetric losses. However, existing asymmetric losses are not compatible with advanced optimization frameworks such as APL, limiting their potential and applicability. Motivated by this theoretical gap and the prospect of asymmetric losses, we extend the asymmetric loss to the more complex passive loss scenario and propose the Asymetric Mean Square Error (AMSE), a novel asymmetric loss. We rigorously establish the necessary and sufficient condition under which AMSE satisfies the asymmetric condition. By substituting the traditional symmetric passive loss in APL with our proposed AMSE, we introduce a novel robust loss framework termed Joint Asymmetric Loss (JAL). Extensive experiments demonstrate the effectiveness of our method in mitigating label noise. Code available at: https://github.com/cswjl/joint-asymmetric-loss Jialiang Wang 0003, Xianming Liu 0005, Gangfeng Hu, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICCV | 5 |
| 2025 | Zero6DOT: Zero-Shot 6D Object Pose Tracking With Monocular RGB Videoabstract6D object tracking plays an important role in various applications, including robotic manipulation and virtual reality. While current methodologies have achieved significant advancements through the use of CAD models, multi-modal sensor data, and category-level assumptions, such resources are often inaccessible in open-world scenarios. Consequently, tracking 6D object poses using only RGB data in such scenarios remains a challenging task. In this paper, we introduce Zero6DOT, an innovative and efficient method for real-time tracking of unknown 6D object poses in monocular RGB video sequences at 8Hz. Our approach requires only the mask of the initial frame, eliminating the need for additional data. The core of Zero6DOT lies in its ability to establish high-quality correspondences across images, from which accurate poses are derived. To achieve this, we employ a transformer-based neural network to predict initial long-term correspondences across frames and integrate a robust Dynamic Units System to refine these predictions. This combination facilitates precise pose tracking while maintaining both efficiency and robustness, even under challenging conditions such as object disappearance, reappearance, and handheld motion. The effectiveness of our approach has been rigorously evaluated through both qualitative and quantitative analyses on the OnePose, YCB-V, and RBOT datasets. The results demonstrate the potential of our proposed Zero6DOT to redefine 6D object pose tracking for real-world scenarios. Deming Zhai, Jianan Zhen, Guofeng Zhang 0001, Xianming Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Learning Lossless Compression for High Bit-Depth Volumetric Medical ImageabstractRecent advances in learning-based methods have markedly enhanced the capabilities of image compression. However, these methods struggle with high bit-depth volumetric medical images, facing issues such as degraded performance, increased memory demand, and reduced processing speed. To address these challenges, this paper presents the Bit-Division based Lossless Volumetric Image Compression (BD-LVIC) framework, which is tailored for high bit-depth medical volume compression. The BD-LVIC framework skillfully divides the high bit-depth volume into two lower bit-depth segments: the Most Significant Bit-Volume (MSBV) and the Least Significant Bit-Volume (LSBV). The MSBV concentrates on the most significant bits of the volumetric medical image, capturing vital structural details in a compact manner. This reduction in complexity greatly improves compression efficiency using traditional codecs. Conversely, the LSBV deals with the least significant bits, which encapsulate intricate texture details. To compress this detailed information effectively, we introduce an effective learning-based compression model equipped with a Transformer-Based Feature Alignment Module, which exploits both intra-slice and inter-slice redundancies to accurately align features. Subsequently, a Parallel Autoregressive Coding Module merges these features to precisely estimate the probability distribution of the least significant bit-planes. Our extensive testing demonstrates that the BD-LVIC framework not only sets new performance benchmarks across various datasets but also maintains a competitive coding speed, highlighting its significant potential and practical utility in the realm of volumetric medical image compression. Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
IEEE Trans. Image Process. | 4 |
| 2025 | FAST: Flexibly Controllable Arbitrary Style Transfer via Latent Diffusion ModelsabstractThe goal of Arbitrary Style Transfer (AST) is injecting the artistic features of a style reference into a given image/video. Existing methods usually pursue the balance between style and content by adjusting general coarse-level stylized strength, thereby leading to unsatisfactory results and hindering their practical application. To address this critical issue, a novel AST approach namely Flexibly Controllable Arbitrary Style Transfer (FAST) is proposed, which is capable of explicitly customizing the stylization results according to various sources of semantic clues. In the specific, our model is constructed based on Latent Diffusion Model (LDM) and elaborately designed to absorb content and style instances as conditions of LDM. It is characterized by introducing Style-Adapter , which allows users to flexibly manipulate the stylization results via aligning multi-level style control information and intrinsic knowledge in LDM, meanwhile enhancing the model with improved capacity to harmonize content detail retention and stylization strength. Lastly, our model is extended to handle video AST task. A novel learning objective is leveraged for video diffusion model training, which considerably improves cross-frame temporal consistency on the premise of maintaining stylization strength. Qualitative and quantitative comparisons as well as user studies demonstrate our presented approach outperforms the existing SoTA methods in generating visually plausible stylization results. The project homepage for the article is available at: https://fast-ldm.github.io/ . Haoran Wang 0004, Zhongrui Yu, Mingming Sun 0001, Junjun Jiang, Xianming Liu 0004, Deming Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Fast and Accurate 6-D Object Pose Refinement via Implicit Surface OptimizationabstractAligning a point cloud to a fixed 3D model is a crucial task in many applications, such as 6D pose estimation for robotic grasping. Typically, an initial pose is estimated by analyzing both the point cloud and the 3D model, after which the Iterative Closest Point (ICP) algorithm is used to refine the pose, reducing large errors and improving accuracy. In this paper, we propose an accurate and efficient alternative to ICP. Our method encodes the fixed 3D model into an implicit neural network, which is trained offline as a one-time process in just a few minutes, requiring only the CAD model of the object. The network takes the point cloud and pose as inputs and outputs the signed distance field (SDF) value. By minimizing the absolute SDF value with the fixed point cloud and network weights, while optimizing the pose, we obtain the final, precise alignment. The key advantage of our method is that it eliminates the need to explicitly establish one-to-one correspondences between the point cloud and the 3D model, a necessary step in ICP and its variants. This enables our framework to avoid local optima and makes it more robust to challenging conditions such as large initial pose gaps, noisy data, variations in scale, occlusions, and reflections. Furthermore, the end-to-end network of our framework offers significant runtime efficiency. We validate the superior performance of our approach through extensive comparisons with various ICP variants on both synthetic and real-world datasets.The source code of the proposed method is available athttps://github.com/pangbo1997/SDFR. Deming Zhai, Jianan Zhen, Xianming Liu 0005 |
IEEE Trans. Robotics | 2 |
| 2024 | Zero-Mean Regularized Spectral Contrastive Learning: Implicitly Mitigating Wrong Connections in Positive-Pair GraphsabstractContrastive learning has emerged as a popular paradigm of self-supervised learning that learns representations by encouraging representations of positive pairs to be similar while representations of negative pairs to be far apart. The spectral contrastive loss, in synergy with the notion of positive-pair graphs, offers valuable theoretical insights into the empirical successes of contrastive learning. In this paper, we propose incorporating an additive factor into the term of spectral contrastive loss involving negative pairs. This simple modification can be equivalently viewed as introducing a regularization term that enforces the mean of representations to be zero, which thus is referred to as *zero-mean regularization*. It intuitively relaxes the orthogonality of representations between negative pairs and implicitly alleviates the adverse effect of wrong connections in the positive-pair graph, leading to better performance and robustness. To clarify this, we thoroughly investigate the role of zero-mean regularized spectral contrastive loss in both unsupervised and supervised scenarios with respect to theoretical analysis and quantitative evaluation. These results highlight the potential of zero-mean regularized spectral contrastive learning to be a promising approach in various tasks. Xianming Liu 0005, Feilong Zhang 0002, Gang Wu 0010, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICLR | 5 |
| 2024 | $\epsilon$-Softmax: Approximating One-Hot Vectors for Mitigating Label NoiseabstractNoisy labels pose a common challenge for training accurate deep neural networks. To mitigate label noise, prior studies have proposed various robust loss functions to achieve noise tolerance in the presence of label noise, particularly symmetric losses. However, they usually suffer from the underfitting issue due to the overly strict symmetric condition. In this work, we propose a simple yet effective approach for relaxing the symmetric condition, namely **$\epsilon$-softmax**, which simply modifies the outputs of the softmax layer to approximate one-hot vectors with a controllable error $\epsilon$. Essentially, ***$\epsilon$-softmax** not only acts as an alternative for the softmax layer, but also implicitly plays the crucial role in modifying the loss function.* We prove theoretically that **$\epsilon$-softmax** can achieve noise-tolerant learning with controllable excess risk bound for almost any loss function. Recognizing that **$\epsilon$-softmax**-enhanced losses may slightly reduce fitting ability on clean datasets, we further incorporate them with one symmetric loss, thereby achieving a better trade-off between robustness and effective learning. Extensive experiments demonstrate the superiority of our method in mitigating synthetic and real-world label noise. Jialiang Wang 0003, Deming Zhai, Junjun Jiang, Xiangyang Ji, Xianming Liu 0005 |
NeurIPS | 3 |
| 2024 | Enhancing Privacy-Utility Tradeoff with Few-Round Strategy in Heterogeneous Federated LearningabstractFederated learning inherently provides a certain level of privacy protection, which however is often inadequate in many real-world scenarios. Existing privacy-preserving methods frequently incur unbearable time overheads or result in non-negligible deterioration to model performance, thus suffering from the tradeoff between performance and privacy. In this work, we propose a novel Federated Privacy-Preserving Knowledge Transfer framework, namely FedPPKT, which employs data-free knowledge distillation in a meta-learning manner to rapidly generates pseudo data and performs privacy-preserving knowledge transfer. FedPPKT establishes a protective barrier between the original private data and the federated model, thereby ensuring user privacy. Furthermore, leveraging the few-round strategy of FedPPKT, it has the capability to reduce the number of communication rounds, further mitigating the risk of privacy exposure for user data. With the help of the meta generator, the problem of uneven local label distribution on clients is alleviated, mitigating data heterogeneity and improving model performance. Experiments show that FedPPKT outperforms the state-of-the-art privacy-preserving federated learning methods. Our code is publicly available at https://github.com/HIT-weiqb/FedPPKT. Qingbin Wei, Feilong Zhang 0002, Yuanchao Bai, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
VCIP | 4 |
| 2024 | Illumination-Aware Low-Light Image Enhancement with Transformer and Auto-Knee CurveabstractImages captured under low-light conditions suffer from several combined degradation factors, including low brightness, low contrast, noise, and color bias. Many learning-based techniques attempt to learn the low-to-clear mapping between low-light and normal-light images. However, they often fall short when applied to low-light images taken in wide-contrast scenes because uneven illumination brings illumination-varying noise and the enhanced images are easily over-saturated in highlight areas. In this article, we present a novel two-stage method to tackle the problem of uneven illumination distribution in low-light images. Under the assumption that noise varies with illumination, we design an illumination-aware transformer network for the first stage of image restoration. In this stage, we introduce the Illumination-aware Attention Block featured with Illumination-aware Multi-head Self-attention, which incorporates different scales of illumination features to guide the attention module, thereby enhancing the denoising and reconstruction capabilities of the restoration network. In the second stage, we innovatively introduce a cubic auto-knee curve transfer with a global parameter predictor to alleviate the over-exposure caused by uneven illumination. We also adopt a white balance correction module to address color bias issues at this stage. Extensive experiments on various benchmarks demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively. Jinwang Pan, Xianming Liu 0005, Yuanchao Bai, Deming Zhai, Junjun Jiang, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Learning Lossless Compression for High Bit-Depth Medical ImagingabstractWe propose a learned lossless image compression method for high bit-depth medical imaging (up to 16 bit-depths). Instead of compressing a high bit-depth medical image as a whole, we split it into two low bit-depth subimages, i.e., the most significant bytes (MSB) subimage and the least significant bytes (LSB) subimage, respectively. The MSB subimage depicts piece-wise smooth structure information that is relatively easy to compress. We thus use traditional lossless codecs for low complexity. The LSB subimage depicts the complementary texture information that is more challenging to compress. We design an autoregressive entropy model conditioned on the MSB subimage that models the probability distribution of the LSB subimage and effectively reduces the redundancy between the MSB and LSB subimages. We then encode the LSB subimage to bitstreams based on the learned entropy model. The compressed high bit-depth medical image is finally stored including the bitstreams of the MSB and LSB subimages. Experimental results demonstrate the state-of-the-art compression performance of the proposed method on high bit-depth medical images, compared with both existing traditional and learned lossless image codecs. Kai Wang 0070, Yuanchao Bai, Deming Zhai, Daxin Li, Junjun Jiang, Xianming Liu 0005 |
ICME | 3 |
| 2023 | On the Dynamics Under the Unhinged Loss and BeyondabstractRecent works have studied implicit biases in deep learning, especially the behavior of last-layer features and classifier weights. However, they usually need to simplify the intermediate dynamics under gradient flow or gradient descent due to the intractability of loss functions and model architectures. In this paper, we introduce the unhinged loss, a concise loss function, that offers more mathematical opportunities to analyze the closed-form dynamics while requiring as few simplifications or assumptions as possible. The unhinged loss allows for considering more practical techniques, such as time-vary learning rates and feature normalization. Based on the layer-peeled model that views last-layer features as free optimization variables, we conduct a thorough analysis in the unconstrained, regularized, and spherical constrained cases, as well as the case where the neural tangent kernel remains invariant. To bridge the performance of the unhinged loss to that of Cross-Entropy (CE), we investigate the scenario of fixing classifier weights with a specific structure, (e.g., a simplex equiangular tight frame). Our analysis shows that these dynamics converge exponentially fast to a solution depending on the initialization of features and classifier weights. These theoretical results not only offer valuable insights, including explicit feature regularization and rescaled learning rates for enhancing practical training with the unhinged loss, but also extend their applicability to other loss functions. Finally, we empirically demonstrate these theoretical results and insights through extensive experiments. Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
J. Mach. Learn. Res. | 4 |
| 2023 | Self-Supervised Arbitrary-Scale Implicit Point Clouds UpsamplingabstractPoint clouds upsampling (PCU), which aims to generate dense and uniform point clouds from the captured sparse input of 3D sensor such as LiDAR, is a practical yet challenging task. It has potential applications in many real-world scenarios, such as autonomous driving, robotics, AR/VR, etc. Deep neural network based methods achieve remarkable success in PCU. However, most existing deep PCU methods either take the end-to-end supervised training, where large amounts of pairs of sparse input and dense ground-truth are required to serve as the supervision; or treat up-scaling of different factors as independent tasks, where multiple networks are required for different scaling factors, leading to significantly increased model complexity and training time. In this article, we propose a novel method that achieves self-supervised and magnification-flexible PCU simultaneously. No longer explicitly learning the mapping between sparse and dense point clouds, we formulate PCU as the task of seeking nearest projection points on the implicit surface for seed points. We then define two implicit neural functions to estimate projection direction and distance respectively, which can be trained by the pretext learning tasks. Moreover, the projection rectification strategy is tailored to remove outliers so as to keep the shape of object clear and sharp. Experimental results demonstrate that our self-supervised learning based scheme achieves competitive or even better performance than state-of-the-art supervised methods. Wenbo Zhao 0004, Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Asymmetric Loss Functions for Noise-Tolerant Learning: Theory and ApplicationsabstractSupervised deep learning has achieved tremendous success in many computer vision tasks, which however is prone to overfit noisy labels. To mitigate the undesirable influence of noisy labels, robust loss functions offer a feasible approach to achieve noise-tolerant learning. In this work, we systematically study the problem of noise-tolerant learning with respect to both classification and regression. Specifically, we propose a new class of loss function, namelyasymmetric loss functions(ALFs), which are tailored to satisfy the Bayes-optimal condition and thus are robust to noisy labels. For classification, we investigate general theoretical properties of ALFs on categorical noisy labels, and introduce the asymmetry ratio to measure the asymmetry of a loss function. We extend several commonly-used loss functions, and establish the necessary and sufficient conditions to make them asymmetric and thus noise-tolerant. For regression, we extend the concept of noise-tolerant learning for image restoration with continuous noisy labels. We theoretically prove that$\ell _{p}$loss ($p>0$) is noise-tolerant for targets with the additive white Gaussian noise. For targets with general noise, we introduce two losses as surrogates of$\ell _{0}$loss that seeks the mode when clean pixels keep dominant. Experimental results demonstrate that ALFs can achieve better or comparative performance compared with the state-of-the-arts. The source code of our method is available at:https://github.com/hitcszx/ALFs. Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Unsupervised Deep Exemplar Colorization via Pyramid Dual Non-Local AttentionabstractExemplar-based colorization is a challenging task, which attempts to add colors to the target grayscale image with the aid of a reference color image, so as to keep the target semantic content while with the reference color style. In order to achieve visually plausible chromatic results, it is important to sufficiently exploit the global color style and the semantic color information of the reference color image. However, existing methods are either clumsy in exploiting the semantic color information, or lack of the dedicated fusion mechanism to decorate the target grayscale image with the reference semantic color information. Besides, these methods usually use a single-stage encoder-decoder architecture, which results in the loss of spatial details. To remedy these problems, we propose an effective exemplar colorization strategy based on pyramid dual non-local attention network to exploit the long-range dependency as well as multi-scale correlation. Specifically, two symmetrical branches of pyramid non-local attention block are tailored to achieve alignments from the target feature to the reference feature and from the reference feature to the target feature respectively. The bidirectional non-local fusion strategy is further applied to get a sufficient fusion feature that achieves full semantic consistency between multi-modal information. To train the network, we propose an unsupervised learning manner, which employs the hybrid supervision including the pseudo paired supervision from the reference color images and unpaired supervision from both the target grayscale and reference color images. Extensive experimental results are provided to demonstrate that our method achieves better photo-realistic colorization performance than the state-of-the-art methods. Deming Zhai, Xianming Liu 0005, Junjun Jiang, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural PhenomenonabstractEstimating the risk level of adversarial examples is essential for safely deploying machine learning models in the real world. One popular approach for physical-world attacks is to adopt the “sticker-pasting” strategy, which however suffers from some limitations, including difficulties in access to the target or printing by valid colors. A new type of non-invasive attacks emerged recently, which attempt to cast perturbation onto the target by optics based tools, such as laser beam and projector. However, the added optical patterns are artificial but not natural. Thus, they are still conspicuous and attention-grabbed, and can be easily noticed by humans. In this paper, we study a new type of optical adversarial examples, in which the perturbations are generated by a very common natural phenomenon, shadow, to achieve naturalistic and stealthy physical-world adversarial attack under the black-box setting. We extensively evaluate the effectiveness of this new attack on both simulated and real-world environments. Experimental results on traffic sign recognition demonstrate that our algorithm can generate adversarial examples effectively, reaching 98.23% and 90.47% success rates on LISA and GTSRB test sets respectively, while continuously misleading a moving camera over 95% of the time in real-world scenarios. We also offer discussions about the limitations and the defense mechanism of this attack11Our code is available at https://github.com/hncszyq/ShadowAttack. Yiqi Zhong, Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
CVPR | 3 |
| 2022 | Learning Towards The Largest Margins
Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICLR | 3 |
| 2022 | Prototype-Anchored Learning for Learning with Imperfect AnnotationsabstractThe success of deep neural networks greatly relies on the availability of large amounts of high-quality annotated data, which however are difficult or expensive to obtain. The resulting labels may be class imbalanced, noisy or human biased. It is challenging to learn unbiased classification models from imperfectly annotated datasets, on which we usually suffer from overfitting or underfitting. In this work, we thoroughly investigate the popular softmax loss and margin-based loss, and offer a feasible approach to tighten the generalization error bound by maximizing the minimal sample margin. We further derive the optimality condition for this purpose, which indicates how the class prototypes should be anchored. Motivated by theoretical analysis, we propose a simple yet effective method, namely prototype-anchored learning (PAL), which can be easily incorporated into various learning-based classification schemes to handle imperfect annotation. We verify the effectiveness of PAL on class-imbalanced learning and noise-tolerant learning by extensive experiments on synthetic and real-world datasets. Xianming Liu 0005, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICML | 3 |
| 2022 | ChebyLighter: Optimal Curve Estimation for Low-light Image EnhancementabstractLow-light enhancement aims to recover a high contrast normal light image from a low-light image with bad exposure and low contrast. Inspired by curve adjustment in photo editing software and Chebyshev approximation, this paper presents a novel model for brightening low-light images. The proposed model, ChebyLighter, learns to estimate pixel-wise adjustment curves for a low-light image recurrently to reconstruct an enhanced output. In ChebyLighter, Chebyshev image series are first generated. Then pixel-wise coefficient matrices are estimated with Triple Coefficient Estimation (TCE) modules and the final enhanced image is recurrently reconstructed by Chebyshev Attention Weighted Summation (CAWS). The TCE module is specifically designed based on dual attention mechanism with three necessary inputs. Our method can achieve ideal performance because adjustment curves can be obtained with numerical approximation by our model. With extensive quantitative and qualitative experiments on diverse test images, we demonstrate that the proposed method performs favorably against state-of-the-art low-light image enhancement algorithms. Jinwang Pan, Deming Zhai, Yuanchao Bai, Junjun Jiang, Debin Zhao, Xianming Liu 0005 |
ACM Multimedia | 2 |
| 2022 | Hybrid Conditional Deep Inverse Tone MappingabstractEmerging modern displays are capable to render ultra-high definition (UHD) media contents with high dynamic range (HDR) and wide color gamut (WCG). Although more and more native contents as such have been getting produced, the total amount is still in severe lack. Considering the massive amount of legacy contents with standard dynamic range (SDR) which may be exploitable, the urgent demand for proper conversion techniques thus springs up. In this paper, we try to tackle the conversion task from SDR to HDR-WCG for media contents and consumer displays. We propose a deep learning based SDR-to-HDR solution, Hybrid Conditional Deep Inverse Tone Mapping (HyCondITM), which is an end-to-end trainable framework including global transform, local adjustment, and detail refinement in a single unified pipeline. We present a hybrid condition network that can simultaneously extract both global and local priors for guidance to achieve scene-adaptive and spatially-variant manipulations. Experiments show that our method achieves state-of-the-art performance in both quantitative comparisons and visual quality, out-performing the previous methods. Tong Shao, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Multimedia | 2 |
| 2022 | Fully Unsupervised Person Re-Identification via Selective Contrastive LearningabstractPerson re-identification (ReID) aims at searching the same identity person among images captured by various cameras. Existing fully supervised person ReID methods usually suffer from poor generalization capability caused by domain gaps. Unsupervised person ReID has attracted a lot of attention recently, because it works without intensive manual annotation and thus shows great potential in adapting to new conditions. Representation learning plays a critical role in unsupervised person ReID. In this work, we propose a novel selective contrastive learning framework for fully unsupervised feature learning. Specifically, different from traditional contrastive learning strategies, we propose to use multiple positives and adaptively selected negatives for defining the contrastive loss, enabling to learn a feature embedding model with stronger identity discriminative representation. Moreover, we propose to jointly leverage global and local features to construct three dynamic memory banks, among which the global and local ones are used for pairwise similarity computation and the mixture memory bank are used for contrastive loss definition. Experimental results demonstrate the superiority of our method in unsupervised person ReID compared with the state of the art. Our code is available at https://github.com/pangbo1997/Unsup_ReID.git . Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Rectified Meta-learning from Noisy Labels for Robust Image-based Plant Disease ClassificationabstractPlant diseases serve as one of main threats to food security and crop production. It is thus valuable to exploit recent advances of artificial intelligence to assist plant disease diagnosis. One popular approach is to transform this problem as a leaf image classification task, which can be then addressed by the powerful convolutional neural networks (CNNs). However, the performance of CNN-based classification approach depends on a large amount of high-quality manually labeled training data, which inevitably introduce noise on labels in practice, leading to model overfitting and performance degradation. To overcome this problem, we propose a novel framework that incorporates rectified meta-learning module into common CNN paradigm to train a noise-robust deep network without using extra supervision information. The proposed method enjoys the following merits: (i) A rectified meta-learning is designed to pay more attention to unbiased samples, leading to accelerated convergence and improved classification accuracy. (ii) Our method is free on assumption of label noise distribution, which works well on various kinds of noise. (iii) Our method serves as a plug-and-play module, which can be embedded into any deep models optimized by gradient descent-based method. Extensive experiments are conducted to demonstrate the superior performance of our algorithm over the state-of-the-arts. Deming Zhai, Ruifeng Shi, Junjun Jiang, Xianming Liu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Learning with Noisy Labels via Sparse RegularizationabstractLearning with noisy labels is an important and challenging task for training accurate deep neural networks. Some commonly-used loss functions, such as Cross Entropy (CE), suffer from severe overfitting to noisy labels. Robust loss functions that satisfy the symmetric condition were tailored to remedy this problem, which however encounter the underfitting effect. In this paper, we theoretically prove that any loss can be made robust to noisy labels by restricting the network output to the set of permutations over a fixed vector. When the fixed vector is one-hot, we only need to constrain the output to be one-hot, which however produces zero gradients almost everywhere and thus makes gradient-based optimization difficult. In this work, we introduce the sparse regularization strategy to approximate the one-hot constraint, which is composed of network output sharpening operation that enforces the output distribution of a net-work to be sharp and the ℓp-norm (p ≤ 1) regularization that promotes the network output to be sparse. This simple approach guarantees the robustness of arbitrary loss functions while not hindering the fitting ability. Experimental results demonstrate that our method can significantly improve the performance of commonly-used loss functions in the presence of noisy labels and class imbalance, and out-perform the state-of-the-art methods. The code is available at https://github.com/hitcszx/lnl_sr. Xianming Liu 0005, Chenyang Wang 0002, Deming Zhai, Junjun Jiang, Xiangyang Ji |
ICCV | 4 |
| 2021 | Target-guided Adaptive Base Class Reweighting for Few-Shot LearningabstractFor few-shot learning, minimizing the empirical risk cannot reach the optimal hypothesis from image to its label due to the effect of overfitting. Therefore, most of the existing work leverages a set of base classes with sufficient labeled samples to pre-train a general encoder for feature representation, which is then applied for all few-shot classification tasks without considering the uniqueness of the target task. We suppose that different base classes help solve a target task in varying degrees, and some classes even introduce a negative effect. To this end, we propose a Target-guided Base Class Reweighting (TBR) approach, which uses a reweighting-in-the-loop optimization algorithm to assign a set of weights for base classes adaptively given a target task. Specifically, TBR learns the parameter of the encoder via minimizing weighted empirical risk on base class data, then optimizes the weights according to the the encoder's performance on support set of the target task. Such an alternating optimization procedure brings reweighting into the loop which makes the encoder more sensitive to the novel classes of the target task. Extensive experiments demonstrate that the proposed method can improve the performance of model-based approaches on two few-shot classification benchmarks. Jiliang Yan, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Multimedia | 2 |
| 2020 | Parsing Map Guided Multi-Scale Attention Network For Face HallucinationabstractFace hallucination that aims to transform a low-resolution (LR) face image to a high-resolution (HR) one is an active domain-specific image super-resolution problem. The performance of existing methods is usually not satisfactory, especially when the upscaling factor is large, such as 8×. In this paper, we propose an effective two- step face hallucination method based on a deep neural network with multi-scale channel and spatial attention mechanism. Specifically, we develop a ParsingNet to extract the prior knowledge of an input LR face, which is then fed into a carefully designed FishSRNet to recover the target HR face. Experimental results demonstrate that our method outperforms the state-of-the-arts in terms of quantitative metrics and visual quality. Chenyang Wang 0002, Zhiwei Zhong 0001, Junjun Jiang, Deming Zhai, Xianming Liu 0005 |
ICASSP | 4 |
| 2020 | ADRN: Attention-Based Deep Residual Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is of crucial importance for many subsequent applications, such as HSI classification and interpretation. In this paper, we propose an attention-based deep residual network to directly learn a mapping from noisy HSI to the clean one. To jointly utilize the spatial-spectral information, the current band and its K adjacent bands are simultaneously exploited as the input. Then, we adopt convolution layer with different filter sizes to fuse the multi-scale feature, and use shortcut connection to incorporate the multi-level information for better noise removal. In addition, the channel attention mechanism is employed to make the network concentrate on the most relevant auxiliary information and features that are beneficial to the de-noising process best. To ease the training procedure, we reconstruct the output through a residual mode rather than a straightforward prediction. Experimental results demonstrate that our proposed ADRN scheme outperforms the state-of-the-art methods both in quantitative and visual evaluations. Yongsen Zhao, Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ICASSP | 2 |
| 2020 | Semi-Supervised Graph Convolutional Hashing Network For Large-Scale Cross-Modal RetrievalabstractCross-modal retrieval aims to provide flexible retrieval results across different types of multimedia data. To confront with scalability issue, binary codes learning (a.k.a. hash technique) is advocated since it permits exact top-K retrieval with sub-linear time complexity. In this paper, we propose a new method called Semi-supervised Graph Convolutional Hashing network (SGCH), which tries to learn a common hamming space by preserving both intra-modality and intermodality similarities via an end-to-end neural network. On one hand, graph convolutional network is utilized to explore high-order intra-modality similarity, and simultaneously propagate the semantic information from labeled samples to unlabeled data. On the other hand, a siamese network is connected to project the learnt features into a common hamming space. To bridge the inter-modality gap, adversarial loss which aims to learn modality-independent features by confusing a modality classifier is incorporated into the overall loss function. Experimental evaluations on cross-media retrieval tasks demonstrate that SGCH performs competitively against the state-of-the-art methods. Zhanjian Shen, Deming Zhai, Xianming Liu 0005, Junjun Jiang |
ICIP | 2 |
| 2020 | Single Image Deraining via Scale-space Invariant Attention Neural NetworkabstractImage enhancement from degradation of rainy artifacts plays a critical role in outdoor visual computing systems. In this paper, we tackle the notion of scale that deals with visual changes in appearance of rain steaks with respect to the camera. Specifically, we revisit multi-scale representation by scale-space theory, and propose to represent the multi-scale correlation in convolutional feature domain, which is more compact and robust than that in pixel domain. Moreover, to improve the modeling ability of the network, we do not treat the extracted multi-scale features equally, but design a novel scale-space invariant attention mechanism to help the network focus on parts of the features. In this way, we summarize the most activated presence of feature maps as the salient features. Extensive experiments results on synthetic and real rainy scenes demonstrate the superior performance of our scheme over the state-of-the-arts. The source code of our method can be found in: https://github.com/pangbo1997/RainRemoval. Deming Zhai, Junjun Jiang, Xianming Liu 0005 |
ACM Multimedia | 2 |
| 2020 | Contrast Enhancement via Dual Graph Total Variation-Based Image DecompositionabstractImages captured in low lighting environment suffer from both low luminance contrast and noise corruption. However, most existing contrast enhancement algorithms only consider contrast boosting, which tends to reveal or amplify noise that is originally not visible in the dark areas. In this paper, we propose a joint contrast enhancement and denoising algorithm, which is based on structure/texture layer decomposition via minimization of dual forms of graph total variation (GTV). Specifically, the structure layer is expected to be generally smoothing but with sharp edges at the foreground background boundaries, for which we propose a quadratic form of GTV (QGTV) as the prior that promotes signal smoothness along graph structure. For the texture layer, a re-weighted GTV (RGTV) is tailored to noise removal while preserving true image details. We provide theoretical analysis about the filtering behavior of these two priors. Furthermore, a boost factor is derived per patch via optimal contrast-tone mapping to improve the overall brightness level of the patch. Finally, an optimization objective function is formulated, which casts image decomposition, brightness boosting, and noise reduction into a unified optimization framework. We further propose a fast approach to efficiently solve the optimization and provide analysis about the convergency. The experimental results show that the proposed method outperforms the state-of-the-art works in subjective, objective, and statistical quality evaluation. Xianming Liu 0005, Deming Zhai, Yuanchao Bai, Xiangyang Ji, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Color-Guided Depth Image Recovery With Adaptive Data Fidelity and Transferred Graph Laplacian RegularizationabstractDepth images play an important role and are prevalently used in many computer vision and computational imaging tasks. However, due to the limitation of active sensing technology, the captured depth images in practice usually suffer from low resolution and noise, which prevents its further applications. To remedy this problem, in this paper, we first propose an adaptive data fidelity formulation to optimally generate each depth pixel from a mixture probability distribution, characterizing the similarity both in the depth map and the corresponding high-resolution guided color image. The proposed method is able to fit the distribution of the input depth signal as an optimization problem by maximizing the mixture probability. Furthermore, to promote the piecewise property that depth images exhibit, we propose a transferred graph Laplacian model as a regularization term, which is general and able to handle various depth recovery tasks such as super-resolution and denoising well. Specifically, each pixel within the recovered depth image is represented as a vertex in a graph with weights in connected edges representing the similarity between vertices. By minimizing the squared variations of the image signal, the task of depth image recovery can be converted to the problem of graph-based image filtering. Since the proposed graph Laplacian regularization model is able to fully exploit a priori information about the depth image, a much more accurate and robust estimation of the underlying depth can be obtained. Extensive experiment evaluations verify that the proposed method obtains recovered depth with higher quality in terms of both objective and subjective criteria, compared with most of the state-of-the-art methods. Yongbing Zhang 0002, Yihui Feng, Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Haoqian Wang, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Depth Restoration From RGB-D Data via Joint Adaptive Regularization and Thresholding on ManifoldsabstractIn this paper, we propose a novel depth restoration algorithm from RGB-D data through combining characteristics of local and non-local manifolds, which provide low-dimensional parameterizations of the local and non-local geometry of depth maps. Specifically, on the one hand, a local manifold model is defined to favor local neighboring relationship of pixels in depth, according to which, manifold regularization is introduced to promote smoothing along the manifold structure. On the other hand, the non-local characteristics of the patch-based manifold can be used to build highly data-adaptive orthogonal bases to extract elongated image patterns, accounting for self-similar structures in the manifold. We further define a manifold thresholding operator in 3D adaptive orthogonal spectral bases-eigenvectors of the discrete Laplacian of local and non-local manifolds-to retain only low graph frequencies for depth maps restoration. Finally, we propose a unified alternating direction method of multipliers optimization framework, which elegantly casts the adaptive manifold regularization and thresholding jointly to regularize the inverse problem of depth maps recovery. Experimental results demonstrate that our method achieves superior performance compared with the state-of-the-art works with respect to both objective and subjective quality evaluations. Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Depth Super-Resolution via Joint Color-Guided Internal and External RegularizationsabstractDepth information is being widely used in many real-world applications. However, due to the limitation of depth sensing technology, the captured depth map in practice usually has much lower resolution than that of color image counterpart. In this paper, we propose to combine the internal smoothness prior and external gradient consistency constraint in graph domain for depth super-resolution. On one hand, a new graph Laplacian regularizer is proposed to preserve the inherent piecewise smooth characteristic of depth, which has desirable filtering properties. A specific weight matrix of the respect graph is defined to make full use of information of both depth and the corresponding guidance image. On the other hand, inspired by an observation that the gradient of depth is small except at edge separating regions, we introduce a graph gradient consistency constraint to enforce that the graph gradient of depth is close to the thresholded gradient of guidance. We reinterpret the gradient thresholding model as variational optimization with sparsity constraint. In this way, we remedy the problem of structure discrepancy between depth and guidance. Finally, the internal and external regularizations are casted into a unified optimization framework, which can be efficiently addressed by ADMM. Experimental results demonstrate that our method outperforms the state-of-the-art with respect to both objective and subjective quality evaluations. Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Noise-Aware Super-Resolution of Depth Maps Via Graph-Based Plug-And-Play FrameworkabstractDepth information is being widely used in many real-world tasks, such as 3DTV, 3D scene reconstruction, multi-view rendering, etc. However, the captured depth maps in practice usually suffer from quality degradations, including low-resolution and noise corruption, which limit their further applications. Noise-aware super-resolution of depth maps is a challenging task and has received increasingly more attention in recent years. In this paper, we propose a novel method based on the plug-and-play scheme, which casts two powerful graph-based tools-the graph Laplacian regularizer and 3D graph Fourier transform-into a unified ADMM optimization framework. It can be performed in an iterative manner with easily treatable convex optimization sub-problems. Experiments results demonstrate that our method achieves superior performance compared with the state-of-the-art works with respect to both objective and subjective quality evaluations. Deming Zhai, Debin Zhao |
ICIP | 2 |
| 2018 | Robust Contrast Enhancement via Graph-Based Cartoon-Texture DecompositionabstractIn this paper, we propose a robust contrast enhancement algorithm based on cartoon and texture layer decomposition. Specifically, the cartoon layer is expected to be generally smoothing but with sharp edges at the foreground and background boundaries, for which we propose a quadratic form of graph total variation (GTV) as the prior to promote signal smoothness along graph structure. For the texture layer, a re-weighted GTV is tailored to remove noises while preserving true image details. Finally, an optimization objective function is formulated, which casts image decomposition, contrast enhancement and noise reduction into a unified framework. We propose an efficient algorithm to solve it. Experimental results show that our generated images outperform state-of-the-art schemes noticeably in subjective quality evaluation. Deming Zhai, Xianming Lu, Xiangyang Ji, Yuanchao Bai, Debin Zhao, Wen Gao 0001 |
ICME | 1 |
| 2018 | Adaptive Screen Content Image Enhancement Strategy using Layer-based SegmentationabstractThe ubiquitous screen content images (SCIs) play a significant role in various scenarios currently. However, most SCIs captured by consumer devices are frequently corrupted with distortions, especially contrast distortion. Unlike the natural images, SCIs are composed of text, graphics and natural scene pictures so that traditional image enhancement methods are not suitable for these compound images. Therefore, we innovatively proposed an adaptive strategy for enhancing SCIs in this paper. Firstly, we devised a segmentation method to divide SCI into text and pictorial regions. Next, the famous guided image filter (GIF) with big and small kernel sizes served as unsharpness masking for processing different regions adaptively. For verifying performance, the proposed method was tested on recently prevalent SCI datasets including SIQAD, and Webpage Dataset. Experimental results indicate that the proposed approach outperforms state-of-the-art methods in most SCIs with flat background. Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet, Xianming Liu 0005, Deming Zhai, Xiao Gu 0001 |
ISCAS | 6 |
| 2018 | Color-Guided Depth Map Super-Resolution via Joint Graph Laplacian and Gradient Consistency RegularizationabstractDepth information is being widely used in many real-world applications. However, due to the limitation of depth sensing technology, the captured depth map in practice usually has much lower resolution than that of color image counterpart. In this paper, we propose to joint exploit the internal smoothness prior and external gradient consistency constraint in graph domain for depth super-resolution. On one hand, a new graph Laplacian regularizer is proposed to the preserve the inherent piecewise smooth characteristic of depth, which has desirable filtering properties. On the other hand, inspired by an observation that the gradient of depth is zero except at edge separating regions, we introduce a graph gradient consistency constraint to enforce that the graph gradient of depth is close to the thresholded gradient of guidance. Finally, the internal and external regularizations are casted into a unified optimization framework, which can be efficiently addressed by ADMM. Experiments results demonstrate that our method outperforms the state-of-the-art with respect to both objective and subjective quality evaluations. Deming Zhai, Debin Zhao |
MMSP | 2 |
| 2018 | Weakly supervised semantic segmentation based on EM algorithm with localization clues
Yang Liu 0006, GuoJun Liu, Deming Zhai, Maozu Guo 0001 |
Neurocomputing | 4 |
| 2018 | Parametric local multiview hamming distance metric learning
Deming Zhai, Xianming Liu 0005, Hong Chang 0001, Yi Zhen, Xilin Chen 0001, Maozu Guo 0001, Wen Gao 0001 |
Pattern Recognit. | 1 |
| 2018 | Supervised Distributed Hashing for Large-Scale Multimedia RetrievalabstractRecent years have witnessed the growing popularity of hashing for large-scale multimedia retrieval. Extensive hashing methods have been designed for data stored in a single machine, that is, centralized hashing . In many real-world applications, however, the large-scale data are often distributed across different locations, servers, or sites. Although hashing for distributed data can be implemented by assembling all distributed data together as a whole dataset in theory, it usually leads to prohibitive computation, communication, and storage costs in practice. Up to now, only a few methods were tailored for distributed hashing, which are all unsupervised approaches. In this paper, we propose an efficient and effective method called supervised distributed hashing (SupDisH), which learns discriminative hash functions by leveraging the semantic label information in a distributed manner. Specifically, we cast the distributed hashing problem into the framework of classification, where the learned binary codes are expected to be distinct enough for semantic retrieval. By introducing auxiliary variables, the distributed model is then separated into a set of decentralized subproblems with consistency constraints, which can be solved in parallel on each vertex of the distributed network. As such, we can obtain high-quality distinctive unbiased binary codes and consistent hash functions with low computational complexity, which facilitate tackling large-scale multimedia retrieval tasks involving distributed datasets. Experimental evaluations on three large-scale datasets show that SupDisH is competitive to centralized hashing methods and outperforms the state-of-the-art unsupervised distributed method significantly. Deming Zhai, Xianming Liu 0005, Xiangyang Ji, Debin Zhao, Shin'ichi Satoh 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Sparsity-Based Image Error Concealment via Adaptive Dual Dictionary Learning and RegularizationabstractIn this paper, we propose a novel sparsity-based image error concealment (EC) algorithm through adaptive dual dictionary learning and regularization. We define two feature spaces: the observed space and the latent space, corresponding to the available regions and the missing regions of image under test, respectively. We learn adaptive and complete dictionaries individually for each space, where the training data are collected via an adaptive template matching mechanism. Based on the piecewise stationarity of natural images, a local correlation model is learned to bridge the sparse representations of the aforementioned dual spaces, allowing us to transfer the knowledge of the available regions to the missing regions for EC purpose. Eventually, the EC task is formulated as a unified optimization problem, where the sparsity of both spaces and the learned correlation model are incorporated. Experimental results show that the proposed method outperforms the state-of-the-art techniques in terms of both objective and perceptual metrics. Xianming Liu 0005, Deming Zhai, Jiantao Zhou 0001, Shiqi Wang 0001, Debin Zhao, Huijun Gao |
IEEE Trans. Image Process. | 2 |
| 2016 | Compressive Sampling-Based Image Coding for Resource-Deficient Visual CommunicationabstractIn this paper, a new compressive sampling-based image coding scheme is developed to achieve competitive coding efficiency at lower encoder computational complexity, while supporting error resilience. This technique is particularly suitable for visual communication with resource-deficient devices. At the encoder, compact image representation is produced, which is a polyphase down-sampled version of the input image; but the conventional low-pass filter prior to down-sampling is replaced by a local random binary convolution kernel. The pixels of the resulting down-sampled pre-filtered image are local random measurements and placed in the original spatial configuration. The advantages of the local random measurements are two folds: 1) preserve high-frequency image features that are otherwise discarded by low-pass filtering and 2) remain a conventional image and can therefore be coded by any standardized codec to remove the statistical redundancy of larger scales. Moreover, measurements generated by different kernels can be considered as the multiple descriptions of the original image and therefore the proposed scheme has the advantage of multiple description coding. At the decoder, a unified sparsity-based soft-decoding technique is developed to recover the original image from received measurements in a framework of compressive sensing. Experimental results demonstrate that the proposed scheme is competitive compared with existing methods, with a unique strength of recovering fine details and sharp edges at low bit-rates. Xianming Liu 0005, Deming Zhai, Jiantao Zhou 0001, Xinfeng Zhang 0001, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Sparsity-based joint gaze correction and face beautification for conferencing videoabstractA well-known problem in video conferencing is gaze mismatch. Instead of relying exclusively on online captured data for rendering, a recent work first trains offline dictionaries using a large image database of movie and TV stars to learn "beautiful" features. During real-time conferencing, one can then simultaneously correct gaze and beautify the subject's facial components in single images by seeking sparse linear combination of pre-trained dictionary atoms for face reconstruction. Extending on this work, we focus on joint gaze correction / face beautification for video. First, we define a large search space invariant to scale, shift and rotation for facial feature beautification based on SIFT. We then address two practical issues unique to video: i) how beautified results can be temporally consistent across group of pictures (GOP), and ii) how blinking eyes can be beautified even though the training database contains only open-eye facial images. Experimental results show that our method achieves the desired temporal consistency, and the blinking process is smooth and natural. Gene Cheung, Deming Zhai, Debin Zhao |
VCIP | 3 |
| 2015 | Instance-specific canonical correlation analysis
Deming Zhai, Yu Zhang 0006, Dit-Yan Yeung, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001 |
Neurocomputing | 1 |
| 2014 | Joint gaze-correction and beautification of DIBR-synthesized human face via dual sparse codingabstractGaze mismatch is a common problem in video conferencing, where the viewpoint captured by a camera (usually located above or below a display monitor) is not aligned with the gaze direction of the human subject, who typically looks at his counterpart in the center of the screen. This means that the two parties cannot converse eye-to-eye, hampering the quality of visual communication. One conventional approach to the gaze mismatch problem is to synthesize a gaze-corrected face image as viewed from center of the screen via depth-image-based rendering (DIBR), assuming texture and depth maps are available at the camera-captured viewpoint(s). Due to self-occlusion, however, there will be missing pixels in the DIBR-synthesized view image that require satisfactory filling. In this paper, we propose to jointly solve the hole-filling problem and the face beautification problem (subtle modifications of facial features to enhance attractiveness of the rendered face) via a unified dual sparse coding framework. Specifically, we first train two dictionaries separately: one for face images of the intended conference subject, one for images of “beautiful” human faces. During synthesis, we simultaneously seek two code vectors - one is sparse in the first dictionary and explains the available DIBR-synthesized pixels, the other is sparse in the second dictionary and matches well with the first vector up to a restricted linear transform. This ensures a good match with the intended target face, while increasing proximity to “beautiful” facial features to improve attractiveness. Experimental results show naturally rendered human faces with noticeably improved attractiveness. Gene Cheung, Deming Zhai, Debin Zhao, Hiroshi Sankoh, Sei Naito |
ICIP | 3 |
| 2014 | Progressive Image Denoising Through Hybrid Graph Laplacian Regularization: A Unified FrameworkabstractRecovering images from corrupted observations is necessary for many real-world applications. In this paper, we propose a unified framework to perform progressive image recovery based on hybrid graph Laplacian regularized regression. We first construct a multiscale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned, which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space. In this procedure, the intrinsic manifold structure is explicitly considered using both measured and unmeasured samples, and the nonlocal self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of the images. On the other hand, between two successive scales, the proposed model is extended to a projected high-dimensional feature space through explicit kernel mapping to describe the interscale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. In this way, the proposed algorithm gradually recovers more and more image details and edges, which could not been recovered in previous scale. We test our algorithm on one typical image recovery task: impulse noise removal. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art algorithms. Xianming Liu 0005, Deming Zhai, Debin Zhao, Guangtao Zhai, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Image Super-Resolution via Hierarchical and Collaborative Sparse RepresentationabstractIn this paper, we propose an efficient image super-resolution algorithm based on hierarchical and collaborative sparse representation (HCSR). Motivated by the observation that natural images typically exhibit multi-modal statistics, we propose a hierarchical sparse coding model which includes two layers: the first layer encodes individual patches, and the second layer jointly encodes the set of patches that belong to the same homogeneous subset of image space. We further present a simple alternative to achieve such target by identifying optimal sparse representation that is adaptive to specific statistics of images. Specially, we cluster images from the offline training set into regions of similar geometric structure, and model each region (cluster) by learning adaptive bases describing the patches within that cluster using principal component analysis (PCA). This cluster-specific dictionary is then exploited to optimally estimate the underlying HR pixel values using the idea of collaborative sparse coding, in which the similarity between patches in the same cluster is further considered. It conceptually and computationally remedies the limitation of many existing algorithms based on standard sparse coding, in which patches are independently encoded. Experimental results demonstrate the proposed method appears to be competitive with state-of-the-art algorithms. Xianming Liu 0005, Deming Zhai, Debin Zhao, Wen Gao 0001 |
DCC | 2 |
| 2013 | Progressive Image Restoration through Hybrid Graph Laplacian RegularizationabstractIn this paper, we propose a unified framework to perform progressive image restoration based on hybrid graph Laplacian regularized regression. We first construct a multi-scale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space by exploring non-local self-similarity. In this procedure, the intrinsic manifold structure is considered by using both measured and unmeasured samples. On the other hand, between two scales, the proposed model is extended to the parametric manner through explicit kernel mapping to model the inter-scale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art image restoration algorithms. Deming Zhai, Xianming Liu 0005, Debin Zhao, Hong Chang 0001, Wen Gao 0001 |
DCC | 1 |
| 2013 | Instance-specific canonical correlation analysis for pose alignmentabstractCanonical correlation analysis (CCA) based methods achieve great success for pose alignment. However, CCA has limitations as a linear and global algorithm. Although some variants have been proposed to overcome the limitations, neither of them achieves locality and nonlinearity at the same time. In this paper, we propose a novel algorithm called Instance-Specific Canonical Correlation Analysis (ISCCA), which approximates the nonlinear data by computing the instance specific projections along the smooth curve of the manifold. Based on the framework of least squares regression, CCA is extended to the instance-specific case which obtains a set of locally-linear smooth but globally-nonlinear transformations. The optimization problem is proved to be convex and could be solved efficiently by alternating optimization. And the globally optimal solutions could be achieved with theoretical guarantee. Experimental results for pose alignment demonstrate the effectiveness of our proposed method. Deming Zhai, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001 |
ICIP | 1 |
| 2013 | Parametric Local Multimodal Hashing for Cross-View Similarity Search
Deming Zhai, Hong Chang 0001, Yi Zhen, Xianming Liu 0005, Xilin Chen 0001, Wen Gao 0001 |
IJCAI | 1 |
| 2012 | Multi-scale Spatial Error Concealment via Hybrid Bayesian RegressionabstractIn this paper, we propose a novel multi-scale spatial error concealment algorithm to combine the modeling strengthes of the parametric and nonparametric Bayesian regression. We progressively recover missing blocks in the scale space from coarse to fine so that the sharp edges and texture in the finest scale can be eventually recovered. On one hand, in each scale, the nonparametric part of the methodology is used to exploit the intra-scale correlation, which relies on the data itself to dictate the structure of the model. In this procedure, the non-local self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of images. On the other hand, the parametric part is used to explicitly model the inter-scale correlation, in which the local structure regularity is thoroughly explored to recover the sharp edges and major texture features of images. It is not respected if only the nonparametric modeling is considering. We achieve the best of both worlds within a multi-scale framework. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art error concealment algorithms. Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Ruiqin Xiong, Wen Gao 0001 |
DCC | 2 |
| 2012 | Web image interpolation via weighted total least squares regressionabstractAlthough ordinary least squares (OLS) regression achieves great success in clean image interpolation, its effectiveness is questionable in the scenario of web images which are usually compressed beforehand. The inherent flaw of OLS is that it is asymmetric, the perturbation is only confined on the right side of the linear system. It is not reasonable for web images. Considering the drawback of OLS, in this paper, we propose an efficient web image interpolation algorithm based on total least squares (TLS) regression. In the proposed method, small perturbations are allowed in both side of the system, which are optimized by TLS in a patch-based manner. Furthermore, we develop a weighted version of TLS to consider contribution diversity of different samples and patches in model estimation, which can efficiently remove the influence of outliers in regression. Experimental results on benchmark test images demonstrate the efficiency of our method. Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Wen Gao 0001 |
ICASSP | 2 |
| 2012 | Multiview Metric Learning with Global Consistency and Local SmoothnessabstractIn many real-world applications, the same object may have different observations (or descriptions) from multiview observation spaces, which are highly related but sometimes look different from each other. Conventional metric-learning methods achieve satisfactory performance on distance metric computation of data in a single-view observation space, but fail to handle well data sampled from multiview observation spaces, especially those with highly nonlinear structure. To tackle this problem, we propose a new method calledMultiview Metric Learning with Global consistency and Local smoothness(MVML-GL) under a semisupervised learning setting, which jointly considers global consistency and local smoothness. The basic idea is to reveal the shared latent feature space of the multiview observations by embodying global consistency constraints and preserving local geometric structures. Specifically, this framework is composed of two main steps. In the first step, we seek a global consistent shared latent feature space, which not only preserves the local geometric structure in each space but also makes those labeled corresponding instances as close as possible. In the second step, the explicit mapping functions between the input spaces and the shared latent space are learned via regularized locally linear regression. Furthermore, these two steps both can be solved by convex optimizations in closed form. Experimental results with application to manifold alignment on real-world datasets of pose and facial expression demonstrate the effectiveness of the proposed method. Deming Zhai, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Side information extrapolation with temporal and spatial consistencyabstractIn this paper, we present an efficient side information extrapolation scheme with temporal and spatial consistency for low delay Wyner-Ziv video coding. Our method is based on the regularized local linear regression (RLLR) model, in which each pixel in SI is approximated as a linear weighted combination of samples within a local temporal neighborhood. The optimal model parameters are estimated by projecting the transformation function onto the temporal training samples to exploit motion-related dependency. During this procedure, moving weights are incorporated into the objective function to express the relative importance of training samples in estimating parameters of the model. Furthermore, spatial correlation is explored by imposing an additional local smoothness penalty, which does good to estimate the occluded regions and complex motion regions. The learned function is smooth and locally linear, and can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results demonstrate that the RLLR method achieves very competitive SI extrapolation performance compared with the state-of-the-art methods. Xianming Liu 0005, Deming Zhai, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ISCAS | 2 |
| 2010 | Manifold Alignment via Corresponding ProjectionsabstractIn this paper, we propose a novel manifold alignment method by learning the underlying common manifold with supervision of corresponding data pairs from different observation sets. Different from the previous algorithms of semi-supervised manifold alignment, our method learns the explicit corresponding projections from each original observation space to the common embedding space everywhere. Benefiting from this property, our method could process new test data directly rather than re-alignment. Furthermore, our approach doesn’t have any assumption on the data structures, thus it could handle more complex cases and get better results compared with previous work. In the proposed algorithm, manifold alignment is formulated as a minimization problem with proper constraints, which could be solved in an analytical manner with closed-form solution. Experimental results on pose manifold alignment of different objects and faces demonstrate the effectiveness of our proposed method. Deming Zhai, Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
BMVC | 1 |
| 2009 | Semi-Supervised Discriminant Analysis via Spectral TransductionabstractLinear Discriminant Analysis (LDA) is a popular method for dimensionality reduction and classification. In real-world applications when there is no sufficient labeled data, LDA suffers from serious performance drop or even fails to work. In this paper, we propose a novel method called Spectral Transduction Semi-Supervised Discriminant Analysis (STSDA), which can alleviate such problem by utilizing both labeled and unlabeled data. Our method takes into consideration both label augmenting and local structure preserving. First, we formulate label transduction with labeled and unlabeled data as a constrained convex optimization problem and solve it efficiently with a closed-form solution by using orthogonal projector matrices. Then, unlabeled data with reliable class estimations are selected with a balanced strategy to augment the original labeled data set. At last, LDA with manifold regularization is performed. Experimental results on face recognition demonstrate the effectiveness of our proposed method. Deming Zhai, Hong Chang 0001, Bo Li 0086, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
BMVC | 1 |