VLDB 2026 Research / reviewers in the wild / expert
Xiaojie Guo 0001
dblp:43/8066-1
· DBLP profile ↗
102ranked-venue papers
19as first author
41since 2021 · last 2026
0000-0002-0326-8382ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 15 first-author · 26 since 2021Artificial intelligence and machine learning · 54 · 12 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semi-Supervised Semantic Segmentation via Derivative Label PropagationabstractSemi-supervised semantic segmentation, which leverages a limited set of labeled images, helps to relieve the heavy annotation burden. While pseudo-labeling strategies yield promising results, there is still room for enhancing the reliability of pseudo-labels. Hence, we develop a semi-supervised framework, namely DerProp, equipped with a novel derivative label propagation to rectify imperfect pseudo-labels. Our label propagation method imposes discrete derivative operations on pixel-wise feature vectors as additional regularization, thereby generating strictly regularized similarity metrics. Doing so effectively alleviates the ill-posed problem that identical similarities correspond to different features, through constraining the solution space. Extensive experiments are conducted to verify the rationality of our design, and demonstrate our superiority over other methods. Yuanbin Fu, Xiaojie Guo 0001 |
AAAI | 2 |
| 2026 | Practical Video Object Detection via Feature Selection and Aggregation
Yuheng Shi, Xiaojie Guo 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Unmasking the Tiny: Foreground probing for small object detection
Hainuo Wang, Xiaojie Guo 0001 |
Image Vis. Comput. | 4 |
| 2026 | Fine-grained normal estimation via feature-augmented knowledge transfer
Meng Wang 0054, Wenjing Dai, Xiaojie Guo 0001 |
Pattern Recognit. | 3 |
| 2025 | Differential Alignment for Domain Adaptive Object DetectionabstractDomain adaptive object detection (DAOD) aims to generalize an object detector trained on labeled source-domain data to a target domain without annotations, the core principle of which is source-target feature alignment. Typically, existing approaches employ adversarial learning to align the distributions of the source and target domains as a whole, barely considering the varying significance of distinct regions, say instances under different circumstances and foreground vs background areas, during feature alignment. To overcome the shortcoming, we investigate a differential feature alignment strategy. Specifically, a prediction-discrepancy feedback instance alignment module (dubbed PDFA) is designed to adaptively assign higher weights to instances of higher teacher-student detection discrepancy, effectively handling heavier domain-specific information. Additionally, an uncertainty-based foreground-oriented image alignment module (UFOA) is proposed to explicitly guide the model to focus more on regions of interest. Extensive experiments on widely-used DAOD datasets together with ablation studies are conducted to demonstrate the efficacy of our proposed method and reveal its superiority over other SOTA alternatives. Xinyu He 0004, Xiaojie Guo 0001 |
AAAI | 3 |
| 2025 | Reversible Decoupling Network for Single Image Reflection RemovalabstractRecent deep-learning-based approaches to single-image reflection removal have shown promising advances, primarily for two reasons: 1) the utilization of recognition-pretrained features as inputs, and 2) the design of dual-stream interaction networks. However, according to the Information Bottleneck principle, high-level semantic clues tend to be compressed or discarded during layer-by-layer propagation. Additionally, interactions in dual-stream networks follow a fixed pattern across different layers, limiting overall performance. To address these limitations, we propose a novel architecture called Reversible Decoupling Network (RDNet), which employs a reversible encoder to secure valuable information while flexibly decoupling transmission- and reflection-relevant features during the forward pass. Furthermore, we customize a transmission-rate-aware prompt generator to dynamically calibrate features, further boosting performance. Extensive experiments demonstrate the superiority of RDNet over existing SOTA methods on five widely-adopted benchmark datasets. Our code will be made publicly available. Mingjia Li 0001, Xiaojie Guo 0001 |
CVPR | 4 |
| 2025 | ShadowHack: Hacking Shadows via Luminance-Color Divide and ConquerabstractShadows introduce challenges such as reduced brightness, texture deterioration, and color distortion in images, complicating a holistic solution. This study presents \textbf{ShadowHack}, a divide-and-conquer strategy that tackles these complexities by decomposing the original task into luminance recovery and color remedy. To brighten shadow regions and repair the corrupted textures in the luminance space, we customize LRNet, a U-shaped network with a rectified attention module, to enhance information interaction and recalibrate contaminated attention maps. With luminance recovered, CRNet then leverages cross-attention mechanisms to revive vibrant colors, producing visually compelling results. Extensive experiments on multiple datasets are conducted to demonstrate the superiority of ShadowHack over existing state-of-the-art solutions both quantitatively and qualitatively, highlighting the effectiveness of our design. Our code will be made publicly available. Mingjia Li 0001, Xiaojie Guo 0001 |
ICCV | 3 |
| 2025 | Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation DecodersabstractThe introduction of generative models has significantly advanced image super-resolution (SR) in handling real-world degradations. However, they often incur fidelity-related issues, particularly distorting textual structures.
In this paper, we introduce a novel diffusion-based SR framework, namely TADiSR, which integrates text-aware attention and joint segmentation decoders to recover not only natural details but also the structural fidelity of text regions in degraded real-world images. Moreover, we propose a complete pipeline for synthesizing high-quality images with fine-grained full-image text masks, combining realistic foreground text regions with detailed background content. Extensive experiments demonstrate that our approach substantially enhances text legibility in super-resolved images, achieving state-of-the-art performance across multiple evaluation metrics and exhibiting strong generalization to real-world scenarios. Our code is available at [here](https://github.com/mingcv/TADiSR). Linlong Fan, Yiyan Luo, Yuhang Yu, Xiaojie Guo 0001, Qingnan Fan |
NeurIPS | 5 |
| 2025 | MODEM: A Morton-Order Degradation Estimation Mechanism for Adverse Weather Image RecoveryabstractRestoring images degraded by adverse weather remains a significant challenge due to the highly non-uniform and spatially heterogeneous nature of weather-induced artifacts, \emph{e.g.}, fine-grained rain streaks versus widespread haze. Accurately estimating the underlying degradation can intuitively provide restoration models with more targeted and effective guidance, enabling adaptive processing strategies. To this end, we propose a Morton-Order Degradation Estimation Mechanism (MODEM) for adverse weather image restoration. Central to MODEM is the Morton-Order 2D-Selective-Scan Module (MOS2D), which integrates Morton-coded spatial ordering with selective state-space models to capture long-range dependencies while preserving local structural coherence. Complementing MOS2D, we introduce a Dual Degradation Estimation Module (DDEM) that disentangles and estimates both global and local degradation priors. These priors dynamically condition the MOS2D modules, facilitating adaptive and context-aware restoration. Extensive experiments and ablation studies demonstrate that MODEM achieves state-of-the-art results across multiple benchmarks and weather types, highlighting its effectiveness in modeling complex degradation dynamics. Our code will be released soon. Hainuo Wang, Xiaojie Guo 0001 |
NeurIPS | 3 |
| 2024 | FNFORMER: A Transformer-Based Face Normal EstimatorabstractFace normal estimation is a crucial step in the development of 3D facial applications, particularly for face modeling and relighting. U-shaped networks are widely used for the task and have witnessed remarkable success. However, CNN-based methods often suffer from unsatisfied generalization ability to out-of-distribution/unseen data, because they do not adequately model long-range dependencies. To address this limitation, Transformer-based approaches have been developed, which benefit from the global self-attention mechanism. Nevertheless, merely using them to learn face normal may lead to limited localization abilities due to insufficient low-level details. In this work, we customize a hybrid model called FNFormer that combines Transformer and CNN to achieve accurate face normal estimation. The proposed model encodes tokenized image patches from CNN feature maps as input to extract global context features using Transformer blocks. Additionally, it extracts detailed local spatial information from a U-shaped CNN. Both the CNN and Transformer features are then integrated for further learning, enabling the network to take both the local and global information into account effectively. Extensive experimental results demonstrate that our proposed FNFormer achieves state-of-the-art performance on various datasets. Our code is available at https://github.com/AutoHDR/FNFormer. Meng Wang 0054, Xiaojie Guo 0001, Jiawan Zhang |
ICME | 2 |
| 2024 | Semi-supervised Camouflaged Object Detection from Noisy DataabstractMost of previous camouflaged object detection methods heavily lean upon large-scale manually-labeled training samples, which are notoriously difficult to obtain. Even worse, the reliability of labels is compromised by the inherent challenges in accurately annotating concealed targets that exhibit high similarities with their surroundings. To overcome these shortcomings, this paper develops the first semi-supervised camouflaged object detection framework, which requires merely a small amount of samples even having noisy/incorrect annotations. Specifically, on the one hand, we introduce an innovative pixel-level loss re-weighting technique to reduce possible negative impacts from imperfect labels, through a window-based voting strategy. On the other hand, we take advantages of ensemble learning to explore robust features against noises/outliers, thereby generating relatively reliable pseudo labels for unlabelled images. Extensive experimental results on four benchmark datasets have been conducted. Yuanbin Fu, Houlei Lv, Xiaojie Guo 0001 |
ACM Multimedia | 4 |
| 2024 | Regional Attention For Shadow RemovalabstractShadow, as a natural consequence of light interacting with objects, plays a crucial role in shaping the aesthetics of an image, which however also impairs the content visibility and overall visual quality. Recent shadow removal approaches employ the mechanism of attention, due to its effectiveness, as a key component. However, they often suffer from two issues including large model size and high computational complexity for practical use. To address these shortcomings, this work devises a lightweight yet accurate shadow removal framework. First, we analyze the characteristics of the shadow removal task to seek the key information required for reconstructing shadow regions and designing a novel regional attention mechanism to effectively capture such information. Then, we customize a Regional Attention Shadow Removal Model (RASM, in short), which leverages non-shadow areas to assist in restoring shadow ones. Unlike existing attention-based models, our regional attention strategy allows each shadow region to interact more rationally with its surrounding non-shadow areas, for seeking the regional contextual correlation between shadow and non-shadow areas. Extensive experiments are conducted to demonstrate that our proposed method delivers superior performance over other state-of-the-art models in terms of accuracy and efficiency, making it appealing for practical applications. Our code can be found at https://github.com/CalcuLuUus/RASM. Hengxing Liu, Mingjia Li 0001, Xiaojie Guo 0001 |
ACM Multimedia | 3 |
| 2024 | Single Image Reflection Separation via Dual-Stream Interactive TransformersabstractDespite satisfactory results on ``easy'' cases of single image reflection separation, prior dual-stream methods still suffer from considerable performance degradation when facing complex ones, i.e, the transmission layer is densely entangled with the reflection having a wide distribution of spatial intensity. The main reasons come from the lack of concern on the feature correlation during interaction, and the limited receptive field. To remedy these deficiencies, this paper presents a Dual-Stream Interactive Transformer (DSIT) design. Specifically, we devise a dual-attention interactive structure that embraces a dual-stream self-attention and a layer-aware dual-stream cross-attention mechanism to simultaneously capture intra-layer and inter-layer feature correlations. Meanwhile, the introduction of attention mechanisms can also mitigate the receptive field limitation. We modulate single-stream pre-trained Transformer embeddings with dual-stream convolutional features through cross-architecture interactions to provide richer semantic priors, thereby further relieving the ill-posedness of the problem. Extensive experimental results reveal the merits of the proposed DSIT over other state-of-the-art alternatives. Our code is publicly available at https://github.com/mingcv/DSIT. Hainuo Wang, Xiaojie Guo 0001 |
NeurIPS | 3 |
| 2024 | A Face Forgery Video Detection Model Based on Knowledge DistillationabstractWith the rapid evolution of artificial intelligence (AI), face forgery videos have proliferated, posing significant societal challenges. Traditional detection methods struggle with poor generalization and cross-database accuracy, unable to address subtle features and variations in face images across scales and compression levels. This paper reviews current face forgery detection methods, identifying key limitations. It introduces a novel model enhancing features through knowledge distillation, optimizing generalization and robustness via a unique loss function and temperature adjustment strategy. Additionally, a Discrete Cosine Transform with multi-scale and multi-compression capabilities (DCTMS) is integrated, enriching texture and detail capture. Experimental results on deepfake datasets demonstrate the efficacy of the proposed methods, achieving high detection accuracy and robustness across diverse scenarios, including cross-database experiments. This study contributes valuable insights and techniques to advance the field of face forgery detection, addressing risks associated with manipulated video content. Haobo Liang, Yingxiong Leng, Jinman Luo, Xiaojie Guo 0001 |
SNPD | 5 |
| 2024 | Depth-Aware Unpaired Video DehazingabstractThis paper investigates a novel unpaired video dehazing framework, which can be a good candidate in practice by relieving pressure from collecting paired data. In such a paradigm, two key issues including 1) temporal consistency uninvolved in single image dehazing, and 2) better dehazing ability need to be considered for satisfied performance. To handle the mentioned problems, we alternatively resort to introducing depth information to construct additional regularization and supervision. Specifically, we attempt to synthesize realistic motions with depth information to improve the effectiveness and applicability of traditional temporal losses, and thus better regularizing the spatiotemporal consistency. Moreover, the depth information is also considered in terms of adversarial learning. For haze removal, the depth information guides the local discriminator to focus on regions where haze residuals are more likely to exist. The dehazing performance is consequently improved by more pertinent guidance from our depth-aware local discriminator. Extensive experiments are conducted to validate our effectiveness and superiority over other competitors. To the best of our knowledge, this study is the initial foray into the task of unpaired video dehazing. Our code is available at https://github.com/YaN9-Y/DUVD. Yang Yang 0062, Chunle Guo, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | DRLIE: Flexible Low-Light Image Enhancement via Disentangled RepresentationsabstractLow-light image enhancement (LIME) aims to convert images with unsatisfied lighting into desired ones. Different from existing methods that manipulate illumination in uncontrollable manners, we propose a flexible framework to take user-specified guide images as references to improve the practicability. To achieve the goal, this article models an image as the combination of two components, that is, content and exposure attribute, from an information decoupling perspective. Specifically, we first adopt a content encoder and an attribute encoder to disentangle the two components. Then, we combine the scene content information of the low-light image with the exposure attribute of the guide image to reconstruct the enhanced image through a generator. Extensive experiments on public datasets demonstrate the superiority of our approach over state-of-the-art alternatives. Particularly, the proposed method allows users to enhance images according to their preferences, by providing specific guide images. Our source code and the pretrained model are available at https://github.com/Linfeng-Tang/DRLIE. Linfeng Tang, Jiayi Ma 0001, Hao Zhang 0073, Xiaojie Guo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Adaptive Texture Filtering for Single-Domain Generalized SegmentationabstractDomain generalization in semantic segmentation aims to alleviate the performance degradation on unseen domains through learning domain-invariant features. Existing methods diversify images in the source domain by adding complex or even abnormal textures to reduce the sensitivity to domain-specific features. However, these approaches depends heavily on the richness of the texture bank and training them can be time-consuming. In contrast to importing textures arbitrarily or augmenting styles randomly, we focus on the single source domain itself to achieve the generalization. In this paper, we present a novel adaptive texture filtering mechanism to suppress the influence of texture without using augmentation, thus eliminating the interference of domain-specific features. Further, we design a hierarchical guidance generalization network equipped with structure-guided enhancement modules, which purpose to learn the domain-invariant generalized knowledge. Extensive experiments together with ablation studies on widely-used datasets are conducted to verify the effectiveness of the proposed model, and reveal its superiority over other state-of-the-art alternatives. Mingjia Li 0001, Yaxing Wang, Chuan-Xian Ren, Xiaojie Guo 0001 |
AAAI | 5 |
| 2023 | YOLOV: Making Still Image Object Detectors Great at Video Object DetectionabstractVideo object detection (VID) is challenging because of the high variation of object appearance as well as the diverse deterioration in some frames. On the positive side, the detection in a certain frame of a video, compared with that in a still image, can draw support from other frames. Hence, how to aggregate features across different frames is pivotal to VID problem. Most of existing aggregation algorithms are customized for two-stage detectors. However, these detectors are usually computationally expensive due to their two-stage nature. This work proposes a simple yet effective strategy to address the above concerns, which costs marginal overheads with significant gains in accuracy. Concretely, different from traditional two-stage pipeline, we select important regions after the one-stage detection to avoid processing massive low-quality candidates. Besides, we evaluate the relationship between a target frame and reference frames to guide the aggregation. We conduct extensive experiments and ablation studies to verify the efficacy of our design, and reveal its superiority over other state-of-the-art VID approaches in both effectiveness and efficiency. Our YOLOX-based model can achieve promising performance (e.g., 87.5% AP50 at over 30 FPS on the ImageNet VID dataset on a single 2080Ti GPU), making it attractive for large-scale or real-time applications. The implementation is simple, we have made the demo codes and models available at https://github.com/YuHengsss/YOLOV. Yuheng Shi, Naiyan Wang, Xiaojie Guo 0001 |
AAAI | 3 |
| 2023 | Single Image Reflection Separation via Component SynergyabstractThe reflection superposition phenomenon is complex and widely distributed in the real world, which derives various simplified linear and nonlinear formulations of the problem. In this paper, based on the investigation of the weaknesses of existing models, we propose a more general form of the superposition model by introducing a learnable residue term, which can effectively capture residual information during decomposition, guiding the separated layers to be complete. In order to fully capitalize on its advantages, we further design the network structure elaborately, including a novel dual-stream interaction mechanism and a powerful decomposition network with a semantic pyramid encoder. Extensive experiments and ablation studies are conducted to verify our superiority over state-of-the-art approaches on multiple real-world benchmark datasets. Our code is publicly available at https://github.com/mingcv/DSRNet. Xiaojie Guo 0001 |
ICCV | 2 |
| 2023 | Practical Edge Detection via Robust Collaborative LearningabstractEdge detection, as a core component in a wide range of vision-oriented tasks, is to identify object boundaries and prominent edges in natural images. An edge detector is desired to be both efficient and accurate for practical use. To achieve the goal, two key issues should be concerned: 1) How to liberate deep edge models from inefficient pre-trained backbones that are leveraged by most existing deep learning methods, for saving the computational cost and cutting the model size; and 2) How to mitigate the negative influence from noisy or even wrong labels in training data, which widely exist in edge detection due to the subjectivity and ambiguity of annotators, for the robustness and accuracy. In this paper, we attempt to simultaneously address the above problems via developing a collaborative learning based model, termed PEdger. The principle behind our PEdger is that, the information learned from different training moments and heterogeneous (recurrent and non recurrent in this work) architectures, can be assembled to explore robust knowledge against noisy annotations, even without the help of pre-training on extra data. Extensive ablation studies together with quantitative and qualitative experimental comparisons on the BSDS500 and NYUD datasets are conducted to verify the effectiveness of our design, and demonstrate its superiority over other competitors in terms of accuracy, speed, and model size. Yuanbin Fu, Xiaojie Guo 0001 |
ACM Multimedia | 2 |
| 2023 | Hierarchical image peeling: A flexible scale-space filtering framework
Yuanbin Fu, Jiayi Ma 0001, Xiaojie Guo 0001 |
Comput. Vis. Image Underst. | 3 |
| 2023 | Low-light Image Enhancement via Breaking Down the Darkness
Xiaojie Guo 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | SPN2D-GAN: Semantic Prior Based Night-to-Day Image-to-Image TranslationabstractExisting image-to-image translation approaches can deal with simple scenes or styles effectively, such as summer-to-winter, horses-to-zebra, and photo-to-map. Although a great progress has been made by GAN-based methods recently, the performance of night-to-day (N2D) translation remains unsatisfactory due to imbalanced/poor visibility, and thus leading to translation ambiguity. To improve the quality of N2D translation, we propose an unpaired translation scheme based on a semantic prior generator, namely SPN2D-GAN, in a weakly- supervised manner with consideration of both image and semantic information. Specifically, we design a novel N2D generator, which can adopt the semantic information of images as prior knowledge to generate more reasonable and realistic results. Also, we suggest adjusting the brightness of nighttime images to boost the visibility, so that the generator can better extract content information. Moreover, the proposed SPN2D-GAN translates images by enforcing the distribution of daytime images in both image and semantic domains on final outputs. Besides, the cycle consistency is employed to preserve the fidelity between translations from two directions. Extensive experimental results are provided to reveal the effectiveness of our design, and demonstrate its superior performance over other state-of-the-art N2D translation approaches both quantitatively and qualitatively. Xiaopeng Li 0009, Xiaojie Guo 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Learning Generalized Knowledge From a Single Domain on Urban-Scene SegmentationabstractDeep neural networks have made significant progress in various tasks under the assumption of the same distribution between training and testing data. However, the obtained domain-specific knowledge often suffers from performance degradation when facing out-of-distribution data. Towards addressing the degradation, a critical requirement of such networks is the generalization capability to unseen domains, which is the goal of domain generalization (DG). This paper attempts to learn generalized knowledge from a single synthetic domain and then apply it to real and unknown scenarios. Specifically, we propose a contour-aware instance normalization module to effectively learn domain-invariant features via a novel weight-updating strategy, which can largely exploit the generalized information from the observed data. In addition, a category-level contrastive learning mechanism is proposed through understanding the semantic discrepancy and relevance among samples to mitigate the interference of domain-specific features on classification. Extensive experiments together with ablation studies on widely-adopted datasets are conducted to demonstrate the effectiveness of our design and show the superiority of our method over other state-of-the-art schemes on the task of urban-scene segmentation. Mingjia Li 0001, Xiaopeng Li 0009, Xiaojie Guo 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Self-augmented Unpaired Image Dehazing via Density and Depth DecompositionabstractTo overcome the overfitting issue of dehazing models trained on synthetic hazy-clean image pairs, many recent methods attempted to improve models' generalization ability by training on unpaired data. Most of them simply formulate dehazing and rehazing cycles, yet ignore the physical properties of the real-world hazy environment, i.e. the haze varies with density and depth. In this paper, we propose a self-augmented image dehazing framework, termed D4 (Dehazing via Decomposing transmission map into Density and Depth) for haze generation and removal. Instead of merely estimating transmission maps or clean content, the proposed framework focuses on exploring scattering coefficient and depth information contained in hazy and clean images. With estimated scene depth, our method is capable of re-rendering hazy images with different thick-nesses which further benefits the training of the dehazing network. It is worth noting that the whole training process needs only unpaired hazy and clean images, yet succeeded in recovering the scattering coefficient, depth map and clean content from a single hazy image. Comprehensive experiments demonstrate our method outperforms state-of-the-art unpaired dehazing methods with much fewer parameters and FLOPs. Our code is available at https://github.com/YaN9-Y/D4. Risheng Liu, Lin Zhang 0014, Xiaojie Guo 0001, Dacheng Tao |
CVPR | 5 |
| 2022 | Low-Light Image Enhancement via Feature RestorationabstractBesides poor visibility, under-exposed images often suffer from severe noise and color distortion. Most existing Retinex-based methods deal with the noise and color distortion via some careful designs to denoising and/or color correction. In this paper, we propose a simple yet effective network from the perspective of feature map restoration to mitigate such issues without constructing any explicit modules. More concretely, we build an encoder-decoder network to reconstruct images, while a feature restoration subnet is introduced to transform the features of low-light images to those of corresponding clear ones. The enhanced images are consequently acquired through assembling the restored features by the decoder, in which, the noise and possible color distortion can be greatly remedied. Extensive experiments on widely-used datasets are conducted to validate the superiority of our design over other state-of-the-art alternatives both quantitatively and qualitatively. Our code is available at https://github.com/YaN9-Y/FRLIE. Yang Yang 0062, Xiaojie Guo 0001 |
ICASSP | 3 |
| 2022 | Feature-Guided Blind Face Restoration with GAN PriorabstractBlind face restoration (BFR) aims to restore high-quality face images from inputs with complex degradation, which is key to extensive applications. Existing methods usually learn a black-box mapping to achieve the goal, which however often produce over-smoothed results. This work proposes a novel feature-guided framework via leveraging prior from a pre-trained Generative Adversarial Network (GAN) model to recover reasonable textures. Furthermore, we design a feature fusion module to guide the generator with low-level spatial content information in the degraded input, for the sake of holding the consistency on both facial structure and background details. The proposed method can be seamlessly integrated with a learned GAN through a simple yet effective principle to recover realistic results under complex degradation circumstances. Extensive comparisons demonstrate the superiority of our strategy over other state-of-the-art methods in terms of restoration quality and training cost. Zhengzhang Hou, Liang Li 0039, Xiaojie Guo 0001 |
ICME | 3 |
| 2022 | N2D-GAN: A Night-to-Day Image-to-Image TranslatorabstractExisting image-to-image translation methods can effectively deal with simple scenes or styles, such as horses-to-zebra, cat-to-dog, and summer-to-winter. However, the performance of night-to-day (N2D) translation remains unsatisfied due to imbalanced/poor visibility and thus translation ambiguity, although some progress has been made by GAN-based methods recently. This paper proposes a CycleGAN-based N2D translation scheme, namely N2D-GAN, in a weakly-supervised manner with consideration of both image and semantic information. Specifically, we first adjust the brightness of night-time images to boost the visibility, so that the generator can better extract content information. Then, the generator processes the translation by enforcing results to follow the distribution of daytime images in both image and semantic domains. Besides, the cycle consistency is introduced to preserve the fidelity between translations from two directions. Experimental results demonstrate that our strategy outperforms other state-of-the-art N2D methods both quantitatively and qualitatively. Xiaopeng Li 0009, Xiaojie Guo 0001, Jiawan Zhang |
ICME | 2 |
| 2022 | Synthetic-to-Real Generalization for Semantic SegmentationabstractThe discrepancy between synthetic and real data is crucial to the performance of domain generalization for semantic segmentation. Since real data is not always accessible, a popular line of approaches is to enhance the diversity of synthetic data via either complex adversarial generation or unstable stylization. However, the internal structure of the synthetic image is often neglected. To largely explore useful information in synthetic data, we observe that, although objects of the same category have different texture patterns between domains, their shapes are quite similar. Based on this observation, we argue that focusing on structural information and alleviating texture dependence are effective ways to improve generalization capability. In this work, we propose an end-to-end network, which explicitly constrains the network to learn shapes and spatial knowledge, and implicitly relieves the texture reliance of the network. Extensive experiments verify the effectiveness of our proposed method and demonstrate its clear advantages over other competitors. Liang Li 0039, Xiaojie Guo 0001 |
ICME | 3 |
| 2022 | Face Inverse Rendering from Single Images in the WildabstractFace inverse rendering, an important and challenging task in computer vision and computer graphics, attempts to decompose face image into shape, reflectance, and illuminance. This problem becomes fundamentally difficult under non-laboratory conditions without controlled illumination. Though recent works have produced compelling results, most of these techniques rely on multiple lighting images captured under contronlled lighting by complex equipment, such as Light Stage, which is not flexible and applicable to common users. In this paper, we propose a novel face inverse rendering framework, which neither relies on complex devices nor labeled training data. Instead, it learns reflectance, shape, and illuminance from its physical constraints. Extensive experiments on both synthetic and real image datasets demonstrate consistently superior performance of the proposed method. Our code will be made publicly available. Meng Wang 0054, Wenjing Dai, Xiaojie Guo 0001, Jiawan Zhang |
ICME | 3 |
| 2022 | Multi-scale Spatial Representation Learning via Recursive Hermite Polynomial NetworksabstractMulti-scale representation learning aims to leverage diverse features from different layers of Convolutional Neural Networks (CNNs) for boosting the feature robustness to scale variance. For dense prediction tasks, two key properties should be satisfied: the high spatial variance across convolutional layers, and the sub-scale granularity inside a convolutional layer for fine-grained features. To pursue the two properties, this paper proposes Recursive Hermite Polynomial Networks (RHP-Nets for short). The proposed RHP-Nets consist of two major components: 1) a dilated convolution to maintain the spatial resolution across layers, and 2) a family of Hermite polynomials over a subset of dilated grids, which recursively constructs sub-scale representations to avoid the artifacts caused by naively applying the dilation convolution. The resultant sub-scale granular features are fused via trainable Hermite coefficients to form the multi-resolution representations that can be fed into the next deeper layer, and thus allowing feature interchanging at all levels. Extensive experiments are conducted to demonstrate the efficacy of our design, and reveal its superiority over state-of-the-art alternatives on a variety of image recognition tasks. Besides, introspective studies are provided to further understand the properties of our method. Yuanbo Lin Wu, Deyin Liu, Xiaojie Guo 0001, Richang Hong |
IJCAI | 3 |
| 2022 | Deep Flexible Structure Preserving Image SmoothingabstractStructure preserving image smoothing is fundamental to numerous multimedia, computer vision, and graphics tasks. This paper develops a deep network in the light of flexibility in controlling, structure preservation in smoothing, and efficiency. Following the principle of divide-and-rule, we decouple the original problem into two specific functionalities, i.e., controllable guidance prediction and image smoothing conditioned on the predicted guidance. Concretely, for flexibly adjusting the strength of smoothness, we customize a two-branch module equipped with a sluice mechanism, which enables altering the strength during inference in a fixed range from 0 (fully smoothing) to 1 (non-smoothing). Moreover, we build a UNet-in-UNet structure with carefully designed loss terms to seek visually pleasant smoothing results without paired data involved for training. As a consequence, our method can produce promising smoothing results with structures well-preserved at arbitrary levels through a compact model with 0.6M parameters, making it attractive for practical use. Quantitative and qualitative experiments are provided to reveal the efficacy of our design, and demonstrate its superiority over other competitors. The code can be found at https://github.com/lime-j/DeepFSPIS. Mingjia Li 0001, Yuanbin Fu, Xiaojie Guo 0001 |
ACM Multimedia | 4 |
| 2022 | Towards High-Fidelity Face Normal EstimationabstractWhile existing face normal estimation methods have produced promising results on small datasets, they often suffer from severe performance degradation on diverse in-the-wild face images, especially for the high-fidelity face normal estimation. Training a high-fidelity face normal estimation model with generalization capability requires a large amount of training data with face normal ground truth. Since collecting such high-fidelity database is difficult in practice, which prevents current methods from recovering face normal with fine-grained geometric details. To mitigate this issue, we propose a coarse-to-fine framework to estimate face normal from an in-the-wild image with only a coarse exemplar reference. Specifically, we first train a model using limited training data to exploit the coarse normal of a real face image. Then, we leverage the estimated coarse normal as an exemplar and devise an exemplar-based normal estimation network to explore robust mapping from the input face image to the fine-grained normal. In this manner, our method can largely alleviate the negative impact caused by lacking training data, and focus on exploring the high-fidelity normal contained in natural images. Extensive experiments and ablation studies are conducted to demonstrate the efficacy of our design, and reveal its superiority over state-of-the-art methods in terms of both training data requirement and recovery quality of fine-grained face normal. Our code is available at \urlhttps://github.com/AutoHDR/HFFNE. Meng Wang 0054, Xiaojie Guo 0001, Jiawan Zhang |
ACM Multimedia | 3 |
| 2022 | U2Fusion: A Unified Unsupervised Image Fusion NetworkabstractThis study proposes a novel unified and unsupervised end-to-end image fusion network, termed as U2Fusion, which is capable of solving different fusion problems, including multi-modal, multi-exposure, and multi-focus cases. Using feature extraction and information measurement, U2Fusion automatically estimates the importance of corresponding source images and comes up with adaptive information preservation degrees. Hence, different fusion tasks are unified in the same framework. Based on the adaptive degrees, a network is trained to preserve the adaptive similarity between the fusion result and source images. Therefore, the stumbling blocks in applying deep learning for image fusion, e.g., the requirement of ground-truth and specifically designed metrics, are greatly mitigated. By avoiding the loss of previous fusion capabilities when training a single model for different tasks sequentially, we obtain a unified model that is applicable to multiple fusion tasks. Moreover, a new aligned infrared and visible image dataset, RoadScene (available at https://github.com/hanna-xu/RoadScene), is released to provide a new option for benchmark evaluation. Qualitative and quantitative experimental results on three typical image fusion tasks validate the effectiveness and universality of U2Fusion. Our code is publicly available at https://github.com/hanna-xu/U2Fusion. Han Xu 0001, Jiayi Ma 0001, Junjun Jiang, Xiaojie Guo 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Face Inverse Rendering via Hierarchical DecouplingabstractPrevious face inverse rendering methods often require synthetic data with ground truth and/or professional equipment like a lighting stage. However, a model trained on synthetic data or using pre-defined lighting priors is typically unable to generalize well for real-world situations, due to the gap between synthetic data/lighting priors and real data. Furthermore, for common users, the professional equipment and skill make the task expensive and complex. In this paper, we propose a deep learning framework to disentangle face images in the wild into their corresponding albedo, normal, and lighting components. Specifically, a decomposition network is built with a hierarchical subdivision strategy, which takes image pairs captured from arbitrary viewpoints as input. In this way, our approach can greatly mitigate the pressure from data preparation, and significantly broaden the applicability of face inverse rendering. Extensive experiments are conducted to demonstrate the efficacy of our design, and show its superior performance in face relighting over other state-of-the-art alternatives. Our code is available at https://github.com/AutoHDR/HD-Net.git. Meng Wang 0054, Xiaojie Guo 0001, Wenjing Dai, Jiawan Zhang |
IEEE Trans. Image Process. | 2 |
| 2021 | Deep Supervised Image RetargetingabstractRecent learning-based image retargeting methods have achieved significant improvement. However, two main is-sues remain in this challenging task: (i) it is difficult to build ground truth datasets for supervised learning; (ii) most methods are based on a certain operator, not suitable for various images with different target sizes. In this paper, for the first time, we address these issues by providing a deep supervised image retargeting solution. We introduce a new dataset1of 6, 576 pairs generated by multiple operators using Image Re-targeting Quality Assessment (IRQA) algorithm. We then develop a mult-operator image retargeting model named MR-GAN, which learns the deformation process of retargeted images using multiple methods and conducts retargeting operations in feature space. Experimental results validate the effectiveness as well as its superiority against state-of-the-art alternatives of the proposed approach. Yijing Mei, Xiaojie Guo 0001, Di Sun 0001, Gang Pan 0002, Jiawan Zhang |
ICME | 2 |
| 2021 | Image Demoiréing with a Dual-Domain Distilling NetworkabstractDue to slight discrepancy of spatial frequency between the camera sensor array and sub-pixel layout of LCD monitor, moiré pattern artifacts appear in various shapes and colours which seriously degrade the quality of captured images. It is challenging yet practically crucial to remove moiré artifacts from a single camera-captured screen image. In this paper, we propose a dual-domain distilling network (3DNet for short) to tackle this problem in an end-to-end manner. The 3DNet consists of a dual-branch student network (a.k.a. demoiréing network), and two teacher networks. The two branches of student network exploit knowledge in both spatial-domain and frequency-domain for the sake of removing moiré artifacts, based on the observation that rich image details can be discovered in the frequency-domain while structure information can be well kept in the spatial-domain. The demoiréing process of two branches is supervised by the knowledge distilled from two teacher networks trained for reconstructing clear images in the spatial and frequency domains respectively. Comprehensive experimental results are conducted to demonstrate the efficacy of our design, and reveal its superiority over state-of-the-art alternatives. Qiaoyu Tian, Liang Li 0039, Xiaojie Guo 0001 |
ICME | 4 |
| 2021 | Trash or Treasure? An Interactive Dual-Stream Strategy for Single Image Reflection SeparationabstractSingle image reflection separation (SIRS), as a representative blind source separation task, aims to recover two layers, $\textit{i.e.}$, transmission and reflection, from one mixed observation, which is challenging due to the highly ill-posed nature. Existing deep learning based solutions typically restore the target layers individually, or with some concerns at the end of the output, barely taking into account the interaction across the two streams/branches. In order to utilize information more efficiently, this work presents a general yet simple interactive strategy, namely $\textit{your trash is my treasure}$ (YTMT), for constructing dual-stream decomposition networks. To be specific, we explicitly enforce the two streams to communicate with each other block-wisely. Inspired by the additive property between the two components, the interactive path can be easily built via transferring, instead of discarding, deactivated information by the ReLU rectifier from one stream to the other. Both ablation studies and experimental results on widely-used SIRS datasets are conducted to demonstrate the efficacy of YTMT, and reveal its superiority over other state-of-the-art alternatives. The implementation is quite simple and our code is publicly available at https://github.com/mingcv/YTMT-Strategy. Xiaojie Guo 0001 |
NeurIPS | 2 |
| 2021 | Beyond Brightening Low-light Images
Xiaojie Guo 0001, Jiayi Ma 0001, Wei Liu 0005, Jiawan Zhang |
Int. J. Comput. Vis. | 2 |
| 2021 | Bilateral attention decoder: A lightweight decoder for real-time semantic segmentation
Chengli Peng, Tian Tian 0006, Chen Chen 0001, Xiaojie Guo 0001, Jiayi Ma 0001 |
Neural Networks | 4 |
| 2021 | SDPNet: A Deep Network for Pan-Sharpening With Enhanced Information RepresentationabstractIn this article, we propose a surface- and deep-level constraint-based pan-sharpening network, termed SDPNet, to address the pan-sharpening problem. Focusing on the two primary goals of pan-sharpening, i.e., spatial and spectral information preservations, we first design two encoder-decoder networks to extract deep-level features from two types of source images, in addition to surface-level characteristics, as the enhanced information representation. The unique feature maps that characterize the unique information in source images can be obtained through the deep-level feature extraction. We further design a pan-sharpening network with densely connected blocks to strengthen feature propagation and reduce parameter number, where the unique feature maps are utilized to efficiently constrain the similarity between the pan-sharpened result and the ground truth, thus avoiding information distortion. Both qualitative and quantitative comparisons on the reduced-resolution and full-resolution source images demonstrate the advantages of our method over state-of-the-art methods. Our code is publicly available at https://github.com/hanna-xu/SDPNet. Han Xu 0001, Jiayi Ma 0001, Hao Zhang 0073, Junjun Jiang, Xiaojie Guo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | FusionDN: A Unified Densely Connected Network for Image FusionabstractIn this paper, we present a new unsupervised and unified densely connected network for different types of image fusion tasks, termed as FusionDN. In our method, the densely connected network is trained to generate the fused image conditioned on source images. Meanwhile, a weight block is applied to obtain two data-driven weights as the retention degrees of features in different source images, which are the measurement of the quality and the amount of information in them. Losses of similarities based on these weights are applied for unsupervised learning. In addition, we obtain a single model applicable to multiple fusion tasks by applying elastic weight consolidation to avoid forgetting what has been learned from previous tasks when training multiple tasks sequentially, rather than train individual models for every fusion task or jointly train tasks roughly. Qualitative and quantitative results demonstrate the advantages of FusionDN compared with state-of-the-art methods in different fusion tasks. Han Xu 0001, Jiayi Ma 0001, Zhuliang Le, Junjun Jiang, Xiaojie Guo 0001 |
AAAI | 5 |
| 2020 | Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and IntensityabstractIn this paper, we propose a fast unified image fusion network based on proportional maintenance of gradient and intensity (PMGI), which can end-to-end realize a variety of image fusion tasks, including infrared and visible image fusion, multi-exposure image fusion, medical image fusion, multi-focus image fusion and pan-sharpening. We unify the image fusion problem into the texture and intensity proportional maintenance problem of the source images. On the one hand, the network is divided into gradient path and intensity path for information extraction. We perform feature reuse in the same path to avoid loss of information due to convolution. At the same time, we introduce the pathwise transfer block to exchange information between different paths, which can not only pre-fuse the gradient information and intensity information, but also enhance the information to be processed later. On the other hand, we define a uniform form of loss function based on these two kinds of information, which can adapt to different fusion tasks. Experiments on publicly available datasets demonstrate the superiority of our PMGI over the state-of-the-art in terms of both visual effect and quantitative metric in a variety of fusion tasks. In addition, our method is faster compared with the state-of-the-art. Hao Zhang 0073, Han Xu 0001, Xiaojie Guo 0001, Jiayi Ma 0001 |
AAAI | 4 |
| 2020 | Generative Landmark Guided Face Inpainting
Yang Yang 0062, Xiaojie Guo 0001 |
PRCV (1) | 2 |
| 2020 | Mutually Guided Image FilteringabstractFiltering images is required by numerous multimedia, computer vision and graphics tasks. Despite diverse goals of different tasks, making effective rules is key to the filtering performance. Linear translation-invariant filters with manually designed kernels have been widely used. However, their performance suffers from content-blindness. To mitigate the content-blindness, a family of filters, called joint/guided filters, have attracted a great amount of attention from the community. The main drawback of most joint/guided filters comes from the ignorance of structural inconsistency between the reference and target signals like color, infrared, and depth images captured under different conditions. Simply adopting such guidelines very likely leads to unsatisfactory results. To address the above issues, this paper designs a simple yet effective filter, named mutually guided image filter (muGIF), which jointly preserves mutual structures, avoids misleading from inconsistent structures and smooths flat regions. The proposed muGIF is very flexible, which can work in various modes including dynamic only (self-guided), static/dynamic (reference-guided) and dynamic/dynamic (mutually guided) modes. Although the objective of muGIF is in nature non-convex, by subtly decomposing the objective, we can solve it effectively and efficiently. The advantages of muGIF in effectiveness and flexibility are demonstrated over other state-of-the-art alternatives on a variety of applications. Our code is publicly available at https://sites.google.com/view/xjguo/mugif. Xiaojie Guo 0001, Yu Li 0003, Jiayi Ma 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Robust Feature Matching Using Spatial Clustering With Heavy OutliersabstractThis paper focuses on removing mismatches from given putative feature matches created typically based on descriptor similarity. To achieve this goal, existing attempts usually involve estimating the image transformation under a geometrical constraint, where a pre-defined transformation model is demanded. This severely limits the applicability, as the transformation could vary with different data and is complex and hard to model in many real-world tasks. From a novel perspective, this paper casts the feature matching into a spatial clustering problem with outliers. The main idea is to adaptively cluster the putative matches into several motion consistent clusters together with an outlier/mismatch cluster. To implement the spatial clustering, we customize the classic density based spatial clustering method of applications with noise (DBSCAN) in the context of feature matching, which enables our approach to achieve quasi-linear time complexity. We also design an iterative clustering strategy to promote the matching performance in case of severely degraded data. Extensive experiments on several datasets involving different types of image transformations demonstrate the superiority of our approach over state-of-the-art alternatives. Our approach is also applied to near-duplicate image retrieval and co-segmentation and achieves promising performance. Xingyu Jiang 0005, Jiayi Ma 0001, Junjun Jiang, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Single Image Deraining: A Comprehensive Benchmark AnalysisabstractWe present a comprehensive study and evaluation of existing single image deraining algorithms, using a new large-scale benchmark consisting of both synthetic and real-world rainy images.This dataset highlights diverse data sources and image contents, and is divided into three subsets (rain streak, rain drop, rain and mist), each serving different training or evaluation purposes. We further provide a rich variety of criteria for dehazing algorithm evaluation, ranging from full-reference metrics, to no-reference metrics, to subjective evaluation and the novel task-driven evaluation. Experiments on the dataset shed light on the comparisons and limitations of state-of-the-art deraining algorithms, and suggest promising future directions. Siyuan Li 0001, Iago Breno Araujo, Wenqi Ren, Zhangyang Wang, Eric K. Tokuda, Roberto Hirata Jr., Roberto Marcondes Cesar Junior, Jiawan Zhang, Xiaojie Guo 0001, Xiaochun Cao |
CVPR | 9 |
| 2019 | Kindling the Darkness: A Practical Low-light Image EnhancerabstractImages captured under low-light conditions often suffer from (partially) poor visibility. Besides unsatisfactory lightings, multiple types of degradations, such as noise and color distortion due to the limited quality of cameras, hide in the dark. In other words, solely turning up the brightness of dark regions will inevitably amplify hidden artifacts. This work builds a simple yet effective network for Kindling the Darkness (denoted as KinD), which, inspired by Retinex theory, decomposes images into two components. One component (illumination) is responsible for light adjustment, while the other (reflectance) for degradation removal. In such a way, the original space is decoupled into two smaller subspaces, expecting to be better regularized/learned. It is worth to note that our network is trained with paired images shot under different exposure conditions, instead of using any ground-truth reflectance and illumination information. Extensive experiments are conducted to demonstrate the efficacy of our design and its superiority over state-of-the-art alternatives. Our KinD is robust against severe visual defects, and user-friendly to arbitrarily adjust light levels. In addition, our model spends less than 50ms to process an image in VGA resolution on a 2080Ti GPU. All the above merits make our KinD attractive for practical use. Jiawan Zhang, Xiaojie Guo 0001 |
ACM Multimedia | 3 |
| 2019 | Single image rain removal via a deep decomposition-composition network
Siyuan Li 0001, Wenqi Ren, Jiawan Zhang, Jinke Yu, Xiaojie Guo 0001 |
Comput. Vis. Image Underst. | 5 |
| 2019 | Locality Preserving Matching
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Xiaojie Guo 0001 |
Int. J. Comput. Vis. | 5 |
| 2019 | Bag similarity network for deep multi-instance learning
Xinggang Wang, Yongluan Yan, Peng Tang 0005, Wenyu Liu 0001, Xiaojie Guo 0001 |
Inf. Sci. | 5 |
| 2019 | Multi-view subspace clustering with intactness-aware similarity
Xiaobo Wang 0001, Zhen Lei 0001, Xiaojie Guo 0001, Changqing Zhang 0002, Hailin Shi, Stan Z. Li |
Pattern Recognit. | 3 |
| 2019 | LMR: Learning a Two-Class Classifier for Mismatch RemovalabstractFeature matching, which refers to establishing reliable correspondence between two sets of features, is a critical prerequisite in a wide spectrum of vision-based tasks. Existing attempts typically involve the mismatch removal from a set of putative matches based on estimating the underlying image transformation. However, the transformation could vary with different data. Thus, a pre-defined transformation model is often demanded, which severely limits the applicability. From a novel perspective, this paper casts the mismatch removal into a two-class classification problem, learning a general classifier to determine the correctness of an arbitrary putative match, termed as Learning for Mismatch Removal (LMR). The classifier is trained based on a general match representation associated with each putative match through exploiting the consensus of local neighborhood structures based on a multiple K -nearest neighbors strategy. With only ten training image pairs involving about 8000 putative matches, the learned classifier can generate promising matching results in linearithmic time complexity on arbitrary testing data. The generality and robustness of our approach are verified under several representative supervised learning techniques as well as on different training and testing data. Extensive experiments on feature matching, visual homing, and near-duplicate image retrieval are conducted to reveal the superiority of our LMR over the state-of-the-art competitors. Jiayi Ma 0001, Xingyu Jiang 0005, Junjun Jiang, Ji Zhao 0001, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Deep Multi-instance Learning with Dynamic PoolingabstractEnd-to-end optimization of multi-instance learning (MIL) using neural networks is an important problem with many applications, in which a core issue is how to design a permutation-invariant pooling function without losing much instance-level information. Inspired by the dynamic routing in recent capsule networks, we propose a novel dynamic pooling function for MIL. It is an adaptive scheme for both key instance selection and modeling the contextual information among instances in a bag. The dynamic pooling iteratively updates the instance contribution to its bag. It is permutation-invariant and can interpret instance-to-bag relationship. The proposed dynamic pooling based multi-instance neural network has been validated on many MIL tasks and outperforms other MIL methods. Yongluan Yan, Xinggang Wang, Xiaojie Guo 0001, Jiemin Fang, Wenyu Liu 0001, Junzhou Huang |
ACML | 3 |
| 2018 | Structure-Texture Decomposition via Joint Structure Discovery and Texture SmoothingabstractStructure-texture decomposition from an image (a.k.a. structure-preserving image smoothing) is important for a variety of multimedia, computer vision and graphics tasks. Its performance heavily depends on the precision of indicating where are structural edges to maintain and where are textures to remove. An intuitive thought for constructing indication is to directly execute edge detection on the input image, which however would suffer from rich textures. Feeding inaccurate or erroneous indications into the smoother is at high risk of generating unsatisfactory results. It is almost sure that edge detectors can do a better job on inputs with textures removed. The above two components, say the smoother and the indicator, turn out to be in a chicken-egg situation. To address this issue, we propose a method to jointly detect structural edges and remove textures, by iteratively smoothing the input based on the edges detected from the previous smoothed result and refining the edges based on the newly processed image. Experiments on a number of challenging cases are conducted to show that the edge detection task and the smoothing task can benefit from each other, and reveal the superiority of our method over other state-of-the-art alternatives. Our code is publicly available at https://sites.google.com/view/xjguo/sdts. Xiaojie Guo 0001, Siyuan Li 0001, Liang Li 0039, Jiawan Zhang |
ICME | 1 |
| 2018 | Soft Clustering Guided Image SmoothingabstractImage smoothing, which aims to remove unwanted textures and preserve desired structures, plays an important role in many multimedia and computer vision tasks. The key to image smoothing, despite different applications, is to distinguish the structures from the textures. This paper presents a novel image smoothing method, following the principle that, for a certain pixel, its neighbors in both space and intensity should contribute more on smoothing, while the distant ones be insulated for avoiding over-smoothing. Intuitively, clustering is a good candidate to achieve the goal. However, due to rich textures and clutters within images, simply performing the clustering on the input likely obtains inaccurate results, and thus leads to unsatisfied smoothing results. In addition, for our task, using traditional hard clustering techniques is at high risk of generating staircase artifacts. For addressing these issues, an algorithm is customized, which on the one hand adopts the soft clustering to more faithfully assign pixels, on the other hand iterates the soft clustering and smoothing, expecting to improve each other. Experiments on several challenging images are provided to show the efficacy of our method, and its superiority over other prevailing approaches. Liang Li 0039, Xiaojie Guo 0001, Wei Feng 0005, Jiawan Zhang |
ICME | 2 |
| 2018 | Co-Referenced Subspace ClusteringabstractSubspace clustering refers to the problem of grouping data into their underlying groups. To address this task, spectral clustering based technique is arguably one of the most popular approaches, and its performance largely depends on the constructed similarity. However, most existing works merely employ the primary representation (e.g., sparse or low-rank representation) as the similarity. In this paper, we propose to explore a high-level co-referenced similarity by employing the Hilbert-Schmidt Independence Criterion (HSIC). Moreover, geometry interpretation of the advantage of our co-referenced similarity is provided. Representation-induced kernels such as Mahalanobis metric, can also be easily embedded into the formulation. Extensive experiments on both synthetic and real-world data are conducted to show the superiority of the proposed method over the state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Hailin Shi, Xiaojie Guo 0001, Xiangyu Zhu 0001, Stan Z. Li |
ICME | 4 |
| 2018 | Co-Saliency Detection via Hierarchical Consistency MeasureabstractCo-saliency detection is a newly emerging research topic in multimedia and computer vision, the goal of which is to extract common salient objects from multiple images. Effectively seeking the global consistency among multiple images is critical to the performance. To achieve the goal, this paper designs a novel model with consideration of a hierarchical consistency measure. Different from most existing co-saliency methods that only exploit common features (such as color and texture), this paper further utilizes the shape of object as another cue to evaluate the consistency among common salient objects. More specifically, for each involved image, an intra-image saliency map is firstly generated via a single image saliency detection algorithm. Having the intra-image map constructed, the consistency metrics at object level and superpixel level are designed to measure the corresponding relationship among multiple images and obtain the inter saliency result by considering multiple visual attention features and multiple constrains. Finally, the intra-image and inter-image saliency maps are fused to produce the final map. Experiments on benchmark datasets are conducted to demonstrate the effectiveness of our method, and reveal its advances over other state-of-the-art alternatives. Liang Li 0039, Runmin Cong, Xiaojie Guo 0001, Jiawan Zhang |
ICME | 4 |
| 2018 | Visual Homing via Guided Locality Preserving MatchingabstractThis study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoramic images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from hundreds of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost true matches without sacrifice in accuracy. To apply our GLPM to the visual homing problem, we develop a method for dense motion flow estimation from sparse feature matches based on Tikhonov regularization. Moreover, the focus-of-contraction/focus-of-expansion is derived to determine homing directions. The effectiveness of our method is demonstrated on a panoramic database in both feature matching and visual homing. Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Yu Zhou 0016, Zheng Wang 0007, Xiaojie Guo 0001 |
ICRA | 7 |
| 2018 | Ensemble Soft-Margin Softmax Loss for Image ClassificationabstractSoftmax loss is arguably one of the most popular losses to train CNN models for image classification. However, recent works have exposed its limitation on feature discriminability. This paper casts a new viewpoint on the weakness of softmax loss. On the one hand, the CNN features learned using the softmax loss are often inadequately discriminative. We hence introduce a soft-margin softmax function to explicitly encourage the discrmination between different classes. On the other hand, the learned classifier of softmax loss is weak. We propose to assemble multiple these weak classifiers to a strong one, inspired by the recognition that the diversity among weak classifiers is critical to a good ensemble. To achieve the diversity, we adopt the Hilbert-Schmidt Independence Criterion (HSIC). Considering these two aspects in one framework, we design a novel loss, named as Ensemble Soft-Margin Softmax (EM-Softmax). Extensive experiments on benchmark datasets are conducted to show the superiority of our design over the baseline softmax loss and several state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Si Liu 0001, Xiaojie Guo 0001, Stan Z. Li |
IJCAI | 5 |
| 2018 | DAAL: Deep activation-based attribute learning for action recognition in depth videos
Chenyang Zhang 0001, Yingli Tian, Xiaojie Guo 0001, Jingen Liu |
Comput. Vis. Image Underst. | 3 |
| 2018 | Dependence-Aware Feature Coding for Person Re-IdentificationabstractIn this letter, we focus on how to boost the performance of person re-identification by exploring the discriminative information among person pairs. A novel dependence-aware feature coding framework is proposed for this task. Specifically, we employ the Hilbert–Schmidt independence criterion as the discriminative term, which is to explore the dependence between different kinds of person pairs, i.e., the same person pairs should be dependence maximized, while the different ones should be dependence minimized. Theoretical discussion and analysis on the convexity of the proposed constraint, as well as the convergence of our algorithm, are provided. Experimental results on two benchmark datasets have demonstrated the advantages of our method over the state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Shengcai Liao, Xiaojie Guo 0001, Yang Yang 0062, Stan Z. Li |
IEEE Signal Process. Lett. | 4 |
| 2018 | Guided Locality Preserving Feature Matching for Remote Sensing Image RegistrationabstractFeature matching, which refers to establishing reliable correspondences between two sets of feature points, is a critical prerequisite in feature-based image registration. This paper proposes a simple yet surprisingly effective approach, termed as guided locality preserving matching, for robust feature matching of remote sensing images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost the true matches without sacrifice in accuracy. Experiments on various real remote sensing image pairs demonstrate the generality of our method for handling both rigid and nonrigid image deformations, and it is more than two orders of magnitude faster than the state-of-the-art methods with better accuracy, making it practical for real-time applications. Jiayi Ma 0001, Junjun Jiang, Huabing Zhou, Ji Zhao 0001, Xiaojie Guo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Low-Rank Matrix Recovery Via Robust Outlier EstimationabstractIn practice, high-dimensional data are typically sampled from low-dimensional subspaces, but with intrusion of outliers and/or noises. Recovering the underlying structure and the pollution from the observations is of utmost importance to understanding the data. Besides properly modeling the subspace structure, how to handle the pollution is a core question regarding the recovery quality, the main origins of which include small dense noises and gross sparse outliers. Compared with the small noises, the outliers more likely ruin the recovery, as their arbitrary magnitudes can dominate the fidelity, and thus lead to misleading/erroneous results. Concerning the above, this paper concentrates on robust outlier estimate for low rank matrix recovery, termed as ROUTE. The principle is to classify each entry as an outlier or an inlier (with confidence). We formulate the outlier screening and the recovery into a unified framework. To seek the optimal solution to the problem, we first introduce a block coordinate descent based optimizer (ROUTE-BCD), then customize an alternating direction method of multipliers based one (ROUTE-ADMM). Through analyzing theoretical properties and practical behaviors, ROUTE-ADMM shows its superiority over ROUTE-BCD in terms of computational complexity, initialization insensitivity and recovery accuracy. Extensive experiments on both synthetic and real data are conducted to show the efficacy of our strategy and reveal its significant improvement over other state-of-the-art alternatives. Our code is publicly available at https://sites.google.com/view/xjguo/route. Xiaojie Guo 0001, Zhouchen Lin |
IEEE Trans. Image Process. | 1 |
| 2017 | Exclusivity-Consistency Regularized Multi-view Subspace ClusteringabstractMulti-view subspace clustering aims to partition a set of multi-source data into their underlying groups. To boost the performance of multi-view clustering, numerous subspace learning algorithms have been developed in recent years, but with rare exploitation of the representation complementarity between different views as well as the indicator consistency among the representations, let alone considering them simultaneously. In this paper, we propose a novel multi-view subspace clustering model that attempts to harness the complementary information between different representations by introducing a novel position-aware exclusivity term. Meanwhile, a consistency term is employed to make these complementary representations to further have a common indicator. We formulate the above concerns into a unified optimization framework. Experimental results on several benchmark datasets are conducted to reveal the effectiveness of our algorithm over other state-of-the-arts. Xiaobo Wang 0001, Xiaojie Guo 0001, Zhen Lei 0001, Changqing Zhang 0002, Stan Z. Li |
CVPR | 2 |
| 2017 | Efficient low rank matrix approximation via orthogonality pursuit and ℓ2 regularizationabstractLow rank matrix approximation, in the presence of missing data and outliers, has previously shown its significance as a theoretic foundation in a wide spectrum of tabulated information processing applications. To fit low rank models, minimizing the nuclear norm of matrices is a popular scheme, the computational load of which, however, is heavy. While bilinear factorization can largely mitigate the computational complexity. Unfortunately, without a known or precisely estimated target rank, this strategy often performs vulnerably when the given data is dirty. This paper attempts to simultaneously achieve the computational efficiency as well as the robustness to mild rank initialization and gross corruptions. Moreover, several Augmented Lagrange Multiplier based solvers and a heuristic rank estimator are customized to seek the optimal solution. Theoretical analysis on convergence and complexity, and experiments on both synthetic and real data are provided to reveal the efficacy of our method and show its superiority over the state-of-the-art alternatives. Siyuan Li 0001, Jiawan Zhang, Xiaojie Guo 0001 |
ICME | 3 |
| 2017 | ROUTE: Robust Outlier Estimation for Low Rank Matrix RecoveryabstractIn practice, even very high-dimensional data are typically sampled from low-dimensional subspaces but with intrusion of outliers and/or noises. Recovering the underlying structure and the pollution from the observations is key to understanding and processing such data. Besides properly modeling the low-rank structure of subspace, how to handle the pollution, is core regarding the performance of recovery. Often, the observed data is posed as a superimposition of the clean data and residual, while the residual can be roughly divided into two groups, including small dense noises and gross sparse outliers. Compared with small noises, outliers more likely ruin the recovery, as they can be arbitrarily large. By considering the above, this paper designs a method for recovering the low rank matrix with robust outlier estimation, termed as ROUTE, in a unified manner. Theoretical analysis on convergence and optimality, and experimental results on both synthetic and real data are provided to demonstrate the efficacy of our proposed method and show its superiority over other state-of-the-arts. Xiaojie Guo 0001, Zhouchen Lin |
IJCAI | 1 |
| 2017 | Exclusivity Regularized Machine: A New Ensemble SVM ClassifierabstractThe diversity of base learners is of utmost importance to a good ensemble. This paper defines a novel measurement of diversity, termed as exclusivity. With the designed exclusivity, we further propose an ensemble SVM classifier, namely Exclusivity Regularized Machine (ExRM), to jointly suppress the training error of ensemble and enhance the diversity between bases. Moreover, an Augmented Lagrange Multiplier based algorithm is customized to effectively and efficiently seek the optimal solution of ExRM. Theoretical analysis on convergence, global optimality and linear complexity of the proposed algorithm, as well as experiments are provided to reveal the efficacy of our method and show its superiority over state-of-the-arts in terms of accuracy and efficiency. Xiaojie Guo 0001, Xiaobo Wang 0001, Haibin Ling |
IJCAI | 1 |
| 2017 | Mutually Guided Image FilteringabstractImage filtering is helpful to numerous multimedia, computer vision and graphics tasks. Linear translation-invariant filters with manually designed kernels have been widely used. However, their performance suffers from the content-blindness, say identically treating noises, textures and structures. To mitigate the content-blindness, a family of filters, called joint/guided filters, has attracted much attention from the community, the principle of which is transferring the structure in the reference image to the target one. The main drawback of most joint/guided filters comes from the ignorance of structural inconsistency between the reference and target signals that can be like color, infrared and depth images captured under different conditions. Simply adopting such guidances very likely leads to unsatisfactory results. To address the above issues, this paper designs a simple yet effective filter, named as mutually guided image filter (muGIF), which jointly preserves mutual structures, avoids misleading from inconsistent structures and smooths flat regions. The proposed muGIF is very flexible, which can perform in one of dynamic only (self-guided), static/dynamic and dynamic/dynamic modes. Although the objective of muGIF is in nature non-convex, by subtly decomposing the objective, we can solve it effectively and efficiently. The advantages of muGIF in terms of effectiveness and flexibility are demonstrated over other state-of-the-art alternatives on a variety of applications. Xiaojie Guo 0001, Yu Li 0003, Jiayi Ma 0001 |
ACM Multimedia | 1 |
| 2017 | LIME: Low-Light Image Enhancement via Illumination Map EstimationabstractWhen one captures images in low-light conditions, the images often suffer from low visibility. Besides degrading the visual aesthetics of images, this poor quality may also significantly degenerate the performance of many computer vision and multimedia algorithms that are primarily designed for high-quality inputs. In this paper, we propose a simple yet effective low-light image enhancement (LIME) method. More concretely, the illumination of each pixel is first estimated individually by finding the maximum value in R, G, and B channels. Furthermore, we refine the initial illumination map by imposing a structure prior on it, as the final illumination map. Having the well-constructed illumination map, the enhancement can be achieved accordingly. Experiments on a number of challenging low-light images are present to reveal the efficacy of our LIME and show its superiority over several state-of-the-arts in terms of enhancement quality and efficiency. Xiaojie Guo 0001, Yu Li 0003, Haibin Ling |
IEEE Trans. Image Process. | 1 |
| 2017 | Single Image Rain Streak Decomposition Using Layer PriorsabstractRain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision and image analysis systems. Rain streak removal algorithms try to recover a rain streak free background scene. In this paper, we address the problem of rain streak removal from a single image by formulating it as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either sparse dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, their performance can often be unsatisfactory, for they tend to either over-smooth the background images or generate -images that still contain noticeable rain streaks. To address the problems, we propose a method that imposes priors for both the background and rain streak layers. These priors are based on Gaussian mixture models learned on small patches that can accommodate a variety of background appearances as well as the appearance of the rain streaks. Moreover, we introduce a structure residue recovery step to further separate the background residues and improve the decomposition quality. Quantitative evaluation shows our method outperforms existing methods by a large margin. We overview our method and demonstrate its effectiveness over prior work on a number of examples. Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown |
IEEE Trans. Image Process. | 3 |
| 2016 | Rain Streak Removal Using Layer PriorsabstractThis paper addresses the problem of rain streak removal from a single image. Rain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision algorithms. Rain streak removal can be formulated as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, they tend to leave too many rain streaks in the background image or over-smooth the background image. In this paper, we propose an effective method that uses simple patch-based priors for both the background and rain layers. These priors are based on Gaussian mixture models and can accommodate multiple orientations and scales of the rain streaks. This simple approach removes rain streaks better than the existing methods qualitatively and quantitatively. We overview our method and demonstrate its effectiveness over prior work on a number of examples. Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown |
CVPR | 3 |
| 2016 | Visual data deblocking using structural layer priorsabstractThe blocking artifact frequently appears in compressed real-world images or video sequences, especially coded at low bit rates, which is visually annoying and likely hurts the performance of many computer vision algorithms. A compressed frame can be viewed as the superimposition of an intrinsic layer and an artifact one. Recovering the two layers from such frames seems to be a severely ill-posed problem since the number of unknowns to recover is twice as many as the given measurements. In this paper, we propose a simple and robust method to separate these two layers, which exploits structural layer priors including the gradient sparsity of the intrinsic layer, and the independence of the gradient fields of the two layers. A novel Augmented Lagrangian Multiplier based algorithm is designed to efficiently and effectively solve the recovery problem. Experimental results demonstrate the efficacy of our method. Siyuan Li 0001, Jiawan Zhang, Xiaojie Guo 0001 |
ICME | 3 |
| 2016 | LIME: A Method for Low-light IMage EnhancementabstractWhen one captures images in low-light conditions, the images often suffer from low visibility. This poor quality may significantly degrade the performance of many computer vision and multimedia algorithms that are primarily designed for high-quality inputs. In this paper, we propose a very simple and effective method, named as LIME, to enhance low-light images. More concretely, the illumination of each pixel is first estimated individually by finding the maximum value in R, G and B channels. Further, we refine the initial illumination map by imposing a structure prior on it, as the final illumination map. Having the well-constructed illumination map, the enhancement can be achieved accordingly. Experiments on a number of challenging real-world low-light images are present to reveal the efficacy of our LIME and show its superiority over several state-of-the-arts. Xiaojie Guo 0001 |
ACM Multimedia | 1 |
| 2016 | High Capacity Reversible Data Hiding in Encrypted Images by Patch-Level Sparse RepresentationabstractReversible data hiding in encrypted images has attracted considerable attention from the communities of privacy security and protection. The success of the previous methods in this area has shown that a superior performance can be achieved by exploiting the redundancy within the image. Specifically, because the pixels in the local structures (like patches or regions) have a strong similarity, they can be heavily compressed, thus resulting in a large hiding room. In this paper, to better explore the correlation between neighbor pixels, we propose to consider the patch-level sparse representation when hiding the secret data. The widely used sparse coding technique has demonstrated that a patch can be linearly represented by some atoms in an over-complete dictionary. As the sparse coding is an approximation solution, the leading residual errors are encoded and self-embedded within the cover image. Furthermore, the learned dictionary is also embedded into the encrypted image. Thanks to the powerful representation of sparse coding, a large vacated room can be achieved, and thus the data hider can embed more secret messages in the encrypted image. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art methods in terms of the embedding rate and the image quality. Xiaochun Cao, Xingxing Wei 0001, Dan Meng 0002, Xiaojie Guo 0001 |
IEEE Trans. Cybern. | 5 |
| 2016 | Total Variation Regularized RPCA for Irregularly Moving Object Detection Under Dynamic BackgroundabstractMoving object detection is one of the most fundamental tasks in computer vision. Many classic and contemporary algorithms work well under the assumption that backgrounds are stationary and movements are continuous, but degrade sharply when they are used in a real detection system, mainly due to: 1) the dynamic background (e.g., swaying trees, water ripples and fountains in real scenarios, as well as raindrops and snowflakes in bad weather) and 2) the irregular object movement (like lingering objects). This paper presents a unified framework for addressing the difficulties mentioned above, especially the one caused by irregular object movement. This framework separates dynamic background from moving objects using the spatial continuity of foreground, and detects lingering objects using the temporal continuity of foreground. The proposed framework assumes that the dynamic background is sparser than the moving foreground that has smooth boundary and trajectory. We regard the observed video as being made up of the sum of a low-rank static background, a sparse and smooth foreground, and a sparser dynamic background. To deal with this decomposition, i.e., a constrained minimization problem, the augmented Lagrangian multiplier method is employed with the help of the alternating direction minimizing strategy. Extensive experiments on both simulated and real data demonstrate that our method significantly outperforms the state-of-the-art approaches, especially for the cases with dynamic backgrounds and discontinuous movements. Xiaochun Cao, Liang Yang 0002, Xiaojie Guo 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | Image Deblurring via Enhanced Low-Rank PriorabstractLow-rank matrix approximation has been successfully applied to numerous vision problems in recent years. In this paper, we propose a novel low-rank prior for blind image deblurring. Our key observation is that directly applying a simple low-rank model to a blurry input image significantly reduces the blur even without using any kernel information, while preserving important edge information. The same model can be used to reduce blur in the gradient map of a blurry input. Based on these properties, we introduce an enhanced prior for image deblurring by combining the low rank prior of similar patches from both the blurry image and its gradient map. We employ a weighted nuclear norm minimization method to further enhance the effectiveness of low-rank prior for image deblurring, by retaining the dominant edges and eliminating fine texture and slight edges in intermediate images, allowing for better kernel estimation. In addition, we evaluate the proposed enhanced low-rank prior for both the uniform and the non-uniform deblurring. Quantitative and qualitative experimental evaluations demonstrate that the proposed algorithm performs favorably against the state-of-the-art deblurring methods. Wenqi Ren, Xiaochun Cao, Jinshan Pan, Xiaojie Guo 0001, Wangmeng Zuo, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Saliency-Aware Nonparametric Foreground Annotation Based on Weakly Labeled DataabstractIn this paper, we focus on annotating the foreground of an image. More precisely, we predict both image-level labels (category labels) and object-level labels (locations) for objects within a target image in a unified framework. Traditional learning-based image annotation approaches are cumbersome, because they need to establish complex mathematical models and be frequently updated as the scale of training data varies considerably. Thus, we advocate the nonparametric method, which has shown potential in numerous applications and turned out to be attractive thanks to its advantages, i.e., lightweight training load and scalability. In particular, we exploit the salient object windows to describe images, which is beneficial to image retrieval and, thus, the subsequent image-level annotation and localization tasks. Our method, namely, saliency-aware nonparametric foreground annotation, is practical to alleviate the full label requirement of training data, and effectively addresses the problem of foreground annotation. The proposed method only relies on retrieval results from the image database, while pretrained object detectors are no longer necessary. Experimental results on the challenging PASCAL VOC 2007 and PASCAL VOC 2008 demonstrate the advance of our method. Xiaochun Cao, Changqing Zhang 0002, Huazhu Fu, Xiaojie Guo 0001, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Adaptively Unified Semi-Supervised Dictionary Learning with Active PointsabstractSemi-supervised dictionary learning aims to construct a dictionary by utilizing both labeled and unlabeled data. To enhance the discriminative capability of the learned dictionary, numerous discriminative terms have been proposed by evaluating either the prediction loss or the class separation criterion on the coding vectors of labeled data, but with rare consideration of the power of the coding vectors corresponding to unlabeled data. In this paper, we present a novel semi-supervised dictionary learning method, which uses the informative coding vectors of both labeled and unlabeled data, and adaptively emphasizes the high confidence coding vectors of unlabeled data to enhance the dictionary discriminative capability simultaneously. By doing so, we integrate the discrimination of dictionary, the induction of classifier to new testing data and the transduction of labels to unlabeled data into a unified framework. To solve the proposed problem, an effective iterative algorithm is designed. Experimental results on a series of benchmark databases show that our method outperforms other state-of-the-art dictionary learning methods in most cases. Xiaobo Wang 0001, Xiaojie Guo 0001, Stan Z. Li |
ICCV | 2 |
| 2015 | A new encryption scheme for surveillance videos
Xiaochun Cao, Meili Ma, Xiaojie Guo 0001, Dongdai Lin |
Frontiers Comput. Sci. | 3 |
| 2015 | Scene Text Deblurring Using Text-Specific Multiscale DictionariesabstractTexts in natural scenes carry critical semantic clues for understanding images. When capturing natural scene images, especially by handheld cameras, a common artifact, i.e., blur, frequently happens. To improve the visual quality of such images, deblurring techniques are desired, which also play an important role in character recognition and image understanding. In this paper, we study the problem of recovering the clear scene text by exploiting the text field characteristics. A series of text-specific multiscale dictionaries (TMD) and a natural scene dictionary is learned for separately modeling the priors on the text and nontext fields. The TMD-based text field reconstruction helps to deal with the different scales of strings in a blurry image effectively. Furthermore, an adaptive version of nonuniform deblurring method is proposed to efficiently solve the real-world spatially varying problem. Dictionary learning allows more flexible modeling with respect to the text field property, and the combination with the nonuniform method is more appropriate in real situations where blur kernel sizes are depth dependent. Experimental results show that the proposed method achieves the deblurring results with better visual quality than the state-of-the-art methods. Xiaochun Cao, Wenqi Ren, Wangmeng Zuo, Xiaojie Guo 0001, Hassan Foroosh |
IEEE Trans. Image Process. | 4 |
| 2015 | SLED: Semantic Label Embedding Dictionary Representation for Multilabel Image AnnotationabstractMost existing methods on weakly supervised image annotation rely on jointly unsupervised feature representation, the components of which are not directly correlated with specific labels. In practical cases, however, there is a big gap between the training and the testing data, say the label combination of the testing data is not always consistent with that of the training. To bridge the gap, this paper presents a semantic label embedding dictionary representation that not only achieves the discriminative feature representation for each label in the image, but also mines the semantic relevance between co-occurrence labels for context information. More specifically, to enhance the discriminative representation of labels, the training data is first divided into a set of overlapped groups by graph shift based on the exclusive label graph. Afterward, given a group of exclusive labels, we try to learn multiple label-specific dictionaries to explicitly decorrelate the feature representation of each label. A joint optimization approach is proposed according to the Fisher discrimination criterion for seeking its solution. Then, to discover the context information hidden in the co-occurrence labels, we explore the semantic relationship between visual words in dictionaries and labels in a multitask learning way with respect to the reconstruction coefficients of the training data. In the annotation stage, with the discriminative dictionaries and exclusive label groups as well as a group sparsity constraint, the reconstruction coefficients of a test image can be easily obtained. Finally, we introduce a label propagation scheme to compute the score of each label for the test image based on its reconstruction coefficients. Experimental results on three challenging data sets demonstrate that our proposed method leads to significant performance gains over existing methods. Xiaochun Cao, Hua Zhang 0008, Xiaojie Guo 0001, Si Liu 0001, Dan Meng 0002 |
IEEE Trans. Image Process. | 3 |
| 2014 | Robust Separation of Reflection from Multiple ImagesabstractWhen one records a video/image sequence through a transparent medium (e.g. glass), the image is often a superposition of a transmitted layer (scene behind the medium) and a reflected layer. Recovering the two layers from such images seems to be a highly ill-posed problem since the number of unknowns to recover is twice as many as the given measurements. In this paper, we propose a robust method to separate these two layers from multiple images, which exploits the correlation of the transmitted layer across multiple images, and the sparsity and independence of the gradient fields of the two layers. A novel Augmented Lagrangian Multiplier based algorithm is designed to efficiently and effectively solve the decomposition problem. The experimental results on both simulated and real data demonstrate the superior performance of the proposed method over the state of the arts, in terms of accuracy and simplicity. Xiaojie Guo 0001, Xiaochun Cao, Yi Ma 0001 |
CVPR | 1 |
| 2014 | Image Retrieval and Ranking via Consistently Reconstructing Multi-attribute Queries
Xiaochun Cao, Hua Zhang 0008, Xiaojie Guo 0001, Si Liu 0001, Xiaowu Chen 0001 |
ECCV (1) | 3 |
| 2014 | Robust Foreground Detection Using Smoothness and Arbitrariness Constraints
Xiaojie Guo 0001, Xinggang Wang, Liang Yang 0002, Xiaochun Cao, Yi Ma 0001 |
ECCV (7) | 1 |
| 2014 | Output Feature Augmented LassoabstractLasso simultaneously conducts variable selection and supervised regression. In this paper, we extend Lasso to multiple output prediction, which belongs to the categories of structured learning. Though structured learning makes use of both input and output simultaneously, the joint feature mapping in current framework of structured learning is usually application-specific. As a result, ad hoc heuristics have to be employed to design different joint feature mapping functions for different applications, which results in the lackness of generalization ability for multiple output prediction. To address this limitation, in this paper, we propose to augment Lasso with output by decoupling the joint feature mapping function of traditional structured learning. The contribution of this paper is three-fold: 1) The augmented Lasso conducts regression and variable selection on both the input and output features, and thus the learned model could fit an output with both the selected input variables and the other correlated outputs. 2) To be more general, we set up nonlinear dependencies among output variables by generalized Lasso. 3) Moreover, the Augmented Lagrangian Method (ALM) with Alternating Direction Minimizing (ADM) strategy is used to find the optimal model parameters. The extensive experimental results demonstrate the effectiveness of the proposed method. Changqing Zhang 0002, Yahong Han, Xiaojie Guo 0001, Xiaochun Cao |
ICDM | 3 |
| 2014 | Speeding uplow rank matrix recovery for foreground separation in surveillance videosabstractVideo surveillance currently is one of the most active research topics in public safety and security, in which foreground extraction is important and fundamental for further processing, such as target tracking, activity recognition, and behavior prediction. In this paper, by assuming the background is highly correlated across different frames, we propose to separate foregrounds via speeded up low rank matrix recovery. The proposed method first shrinks the scale of data to roughly catch outliers (foregrounds) of the background. Based on the outliers, we design a sampling strategy that selects a number of frames to construct the low rank model of background. According to the constructed background model, our method further recovers both the background and the foreground for the rest frames in a reconstruction manner. Experimental results on both simulated and real data demonstrate the clear advantage of our approach compared to the state of the arts, in terms of accuracy and efficiency. Xiaojie Guo 0001, Xiaochun Cao |
ICME | 1 |
| 2014 | Augmented Image Retrieval using Multi-order Object Layout with AttributesabstractIn image retrieval, users' search intention is usually specified by textual queries, exemplar images, concept maps, and even sketches, which can only express the search intention partially. These query strategies lack the abilities to indicate the Regions Of Interests (ROIs) and represent the spatial or semantic correlations among the ROIs, which results in the so-called semantic gap between users' search intention and images' low-level visual content. In this paper, we propose a novel image search method, which allows the users to indicate any number of Regions Of Interest (ROIs) within the query as well as utilize various semantic concepts and spatial relations to search images. Specifically, we firstly propose a structured descriptor to jointly represent the categories, attributes, and spatial relations among objects. Then, based on the defined descriptor, our method ranks the images in the database according to the matching scores w.r.t. the category, attribute, and spatial relations. We conduct the experiments on the aPascal and aYahoo datasets, and experimental results show the advantage of the proposed method compared to the state of the arts. Xiaochun Cao, Xingxing Wei 0001, Xiaojie Guo 0001, Yahong Han, Jinhui Tang 0001 |
ACM Multimedia | 3 |
| 2014 | Beautifying Fisheye Images using Orientation and Shape CuesabstractFisheye images, due to their wide range of vision, become more and more popular in our daily life. However, the fisheye images usually suffer from misalignment that reduces their visual pleasure. In this paper, we develop a computational method for enhancing the aesthetics of such images by exploiting the orientation and shape cues. More specifically, the orientation cue is based on the observation that cameras are often oriented when taking photos, so that their upvectors are parallel to vertical linear structures in the scene. While the shape one refers to that after repositing the fisheye image, the circular shape should be preserved. By employing these two rules as our basic aesthetic guidelines, our method can correct the rotation angle between the camera coordinate and the world coordinate to make the virtual camera oriented, and complete the missing part. Experimental results on a number of challenging indoor and outdoor fisheye images show the effectiveness of our approach, and demonstrate the superior aesthetics of the proposed method compared to the state-of-the-arts. Xiaobo Wang 0001, Xiaochun Cao, Xiaojie Guo 0001, Zhanjie Song |
ACM Multimedia | 3 |
| 2014 | Video color conceptualization using optimization
Xiaochun Cao, Xiaojie Guo 0001, Yiu-Ming Cheung |
Sci. China Inf. Sci. | 3 |
| 2013 | Video Editing with Temporal, Spatial and Appearance ConsistencyabstractGiven an area of interest in a video sequence, one may want to manipulate or edit the area, e.g. remove occlusions from or replace with an advertisement on it. Such a task involves three main challenges including temporal consistency, spatial pose, and visual realism. The proposed method effectively seeks an optimal solution to simultaneously deal with temporal alignment, pose rectification, as well as precise recovery of the occlusion. To make our method applicable to long video sequences, we propose a batch alignment method for automatically aligning and rectifying a small number of initial frames, and then show how to align the remaining frames incrementally to the aligned base images. From the error residual of the robust alignment process, we automatically construct a trimap of the region for each frame, which is used as the input to alpha matting methods to extract the occluding foreground. Experimental results on both simulated and real data demonstrate the accurate and robust performance of our method. Xiaojie Guo 0001, Xiaochun Cao, Xiaowu Chen 0001, Yi Ma 0001 |
CVPR | 1 |
| 2013 | SYM-FISH: A Symmetry-Aware Flip Invariant Sketch Histogram Shape DescriptorabstractRecently, studies on sketch, such as sketch retrieval and sketch classification, have received more attention in the computer vision community. One of its most fundamental and essential problems is how to more effectively describe a sketch image. Many existing descriptors, such as shape context, have achieved great success. In this paper, we propose a new descriptor, namely Symmetric-aware Flip Invariant Sketch Histogram (SYM-FISH) to refine the shape context feature. Its extraction process includes three steps. First the Flip Invariant Sketch Histogram (FISH) descriptor is extracted on the input image, which is a flip-invariant version of the shape context feature. Then we explore the symmetry character of the image by calculating the kurtosis coefficient. Finally, the SYM-FISH is generated by constructing a symmetry table. The new SYM-FISH descriptor supplements the original shape context by encoding the symmetric information, which is a pervasive characteristic of natural scene and objects. We evaluate the efficacy of the novel descriptor in two applications, i.e., sketch retrieval and sketch classification. Extensive experiments on three datasets well demonstrate the effectiveness and robustness of the proposed SYM-FISH descriptor. Xiaochun Cao, Hua Zhang 0008, Si Liu 0001, Xiaojie Guo 0001, Liang Lin 0004 |
ICCV | 4 |
| 2013 | Horizon matters: Image re-targeting using horizon cuesabstractIn this work, we propose an effective method for improving the aesthetics of a given image based on the horizon information extracted from the image. We employ the vanishing line detection method to determine the horizon and then adjust the angle of the scene to level the horizon. The image is thus retargeted to highlight the visual appearance. In addition, our method can also be applied on the task of distinguishing the image acquisition methods. For example, hand held phones typically don't have perfect level horizons. Experimental results demonstrate the performance of our proposed method. Xiaochun Cao, Siyuan Li 0001, Xiaojie Guo 0001 |
ICME | 4 |
| 2013 | Motion matters: a novel framework for compressing surveillance videosabstractCurrently, video surveillance plays a very important role in the fields of public safety and security. For storing the videos that usually contain extremely long sequences, it requires huge space. Video compression techniques can be used to release the storage load to some extent, such as H.264/AVC. However, the existing codecs are not sufficiently effective and efficient for encoding surveillance videos as they do not specifically consider the characteristic of surveillance videos, i.e. the background of surveillance video has intensive redundancy. This paper introduces a novel framework for compressing such videos. We first train a background dictionary based on a small number of observed frames. With the trained background dictionary, we then separate every frame into the background and motion (foreground), and store the compressed motion together with the reconstruction coefficient of the background corresponding to the background dictionary. The decoding is carried out on the encoded frame in an inverse procedure. The experimental results on extensive surveillance videos demonstrate that our proposed method significantly reduces the size of videos while gains much higher PSNR compared to the state of the art codecs. Xiaojie Guo 0001, Siyuan Li 0001, Xiaochun Cao |
ACM Multimedia | 1 |
| 2012 | Motion saliency detection using low-rank and sparse decompositionabstractMotion saliency detection has an important impact on further video processing tasks, such as video segmentation, object recognition and adaptive compression. Different to image saliency, in videos, moving regions (objects) catch human beings' attention much easier than static ones. Based on this observation, we propose a novel method of motion saliency detection, which makes use of the low-rank and sparse decomposition on video slices along X-T and Y-T planes to achieve the goal, i.e. separating foreground moving objects from backgrounds. In addition, we adopt the spatial information to preserve the completeness of the detected motion objects. In virtue of adaptive threshold selection and efficient noise elimination, the proposed approach is suitable for different video scenes, and robust to low resolution and noisy cases. The experiments demonstrate the performance of our method compared with the state-of-the-art. Yawen Xue, Xiaojie Guo 0001, Xiaochun Cao |
ICASSP | 2 |
| 2012 | MIFT: A framework for feature descriptors to be mirror reflection invariant
Xiaojie Guo 0001, Xiaochun Cao |
Image Vis. Comput. | 1 |
| 2012 | Good match exploration using triangle constraint
Xiaojie Guo 0001, Xiaochun Cao |
Pattern Recognit. Lett. | 1 |
| 2011 | Identifying Image Composites Through Shadow Matte ConsistencyabstractIn this paper, we propose a framework for detecting tampered digital images based on photometric consistency of illumination in shadows. In particular, we formulate color characteristics of shadows measured by the shadow matte value. The shadow boundaries and the penumbra shadow region in an image are first extracted. Then a simple and efficient method is used to estimate shadow matte values of shadows. Our approach efficiently extracts these constraints from a single view of a target scene and makes use of them for the digital forgery detection. Experimental results on both simulated photos and visually plausible real images demonstrate the effectiveness of the proposed method. Qiguang Liu, Xiaochun Cao, Xiaojie Guo 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2010 | FIND: A Neat Flip Invariant DescriptorabstractIn this paper, we introduce a novel Flip Invariant Descriptor (FIND). FIND improves the degenerated performance resulted from image flips and reduces both space and time costs. Flip invariance of FIND enables the intractable flip detection to be achieved easily, instead of duplicately implementing the procedure. To alleviate the pressure brought by the increasing scale of image and video data, FIND utilizes a concise structure with less storage space. Comparing to SIFT, FIND reduces 35.94% length for a descriptor. We compare FIND against SIFT with respect to accuracy, speed and space cost. An application to image search over a database of 3.27 million descriptors is also shown. Xiaojie Guo 0001, Xiaochun Cao |
ICPR | 1 |
| 2010 | Triangle-Constraint for Finding More Good FeaturesabstractWe present a novel method for finding more good feature pairs between two sets of features. We first select matched features by Bi-matching method as seed points, then organize these seed points by adopting the Delaunay triangulation algorithm. Finally, we use Triangle-Constraint (T-C) to increase both number of correct matches and matching score (the ratio between number of correct matches and total number of matches). Delaunay triangulation algorithm.The experimental evaluation shows that our method is robust to most of geometric and photometric transformations including rotation, scale change, blur, viewpoint change, JPEG compression and illumination change, and significantly improves both number of correct matches and matching score. Xiaojie Guo 0001, Xiaochun Cao |
ICPR | 1 |
| 2010 | Water Reflection Detection Using a Flip Invariant Shape DetectorabstractWater reflection detection is a tough task in computer vision, since the reflection is distorted by ripples irregularly. This paper proposes an effective method to detect water reflections. We introduce a descriptor that is not only invariant to scales, rotations and affine transformations, but also tolerant to the flip transformation and even non-rigid distortions, such as ripple effects. We analyze the structure of our descriptor and show how it outperforms the existing mirror feature descriptors in the context of water reflection. The experimental results demonstrate that our method is able to detect the water reflections. Hua Zhang 0008, Xiaojie Guo 0001, Xiaochun Cao |
ICPR | 2 |
| 2009 | MIFT: A Mirror Reflection Invariant Feature Descriptor
Xiaojie Guo 0001, Xiaochun Cao, Jiawan Zhang |
ACCV (2) | 1 |