EDBT 2026 Demo / reviewers in the wild / expert
Yi Xu 0001
dblp:14/5580-1
· DBLP profile ↗
81ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0001-6508-4469ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 25 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Photorealistic Style Transfer with Multimodal Guidance and Robustness to Content Images in Arbitrary StylesabstractExisting photorealistic style transfer methods are broadly categorized into two groups: image-guided and text-guided approaches. The image-guided paradigm requires a reference style image, which is superior when the target style is difficult to define precisely with text. Unfortunately, such references are not always available in practical scenarios. In contrast, the text-guided paradigm offers greater flexibility. However, existing text-guided methods often fail to preserve the details of content images or perform poorly when content images deviate from normal style. In this paper, we present a novel multimodal-guided photorealistic style transfer framework, supporting flexible switching and fusion of both modalities while ensuring robust performance across content images in arbitrary styles. Specifically, we adopt a two-stage pipeline. First, the Style Removal Module removes the original style from the content image. Then, the Style Injection Module applies stylization based on the style guidance (image, text, or their fusion). To make the text-guided branch compatible with this pipeline, we propose the Image-Assisted Textual Style Injection (IATSI) strategy. Additionally, we design a Dual-Residual Adaptive MLP (DRA-MLP), which exhibits strong color mapping capability and avoids spatial distortions. Extensive experiments show that our method achieves state-of-the-art (SOTA) performance in both image-guided and text-guided settings. Moreover, we innovatively implement multimodal fusion-guided photorealistic style transfer, achieving promising results. Ruikai Zhou, Yi Xu 0001 |
WACV | 3 |
| 2026 | Self-distilled learning of adaptive interval 3D lookup tables on real-time image enhancement
Ruikai Zhou, Canqian Yang, Meiguang Jin, Xu Jia 0012, Ying Chen 0011, Yi Xu 0001 |
Pattern Recognit. | 7 |
| 2026 | Text-Driven Relation Manipulation of Diffusion ImageryabstractText-guided image manipulation has recently attracted significant attention. Prevailing algorithms predominantly focus on modifying the appearances of existing instances, such as texture and attribute editing, while they often fail to address the interactions between different instances or achieve fundamental structural changes, such as multi-object editing. This paper introduces a novel text-guided manipulation task named "relation manipulation", aimed at fundamentally altering the structure of images. This task is capable of modifying the quantity of instances and, more importantly, enhancing the understanding and editing of interactions among diverse instances. Our approach comprises two main components: relation customization and multi-region guided diffusion. Relation customization fine-tunes specific relationships using a compact dataset of exemplary relations, facilitating nuanced understanding and implementation of instance interactions. Multi-region guided diffusion employs gradient optimization to update the generation process across multiple regions, integrating a fine-grained attention control strategy to minimize regional interference and conflict. Additionally, the demonstrated applications of our method in multi-region inversion underline its potential in practical scenarios, such as relation manipulation of real images and consecutive image manipulation. Compatible with different variants of Stable Diffusion models, our approach seamlessly integrates into the Stable Diffusion WebUI, enabling high-quality image generation and exceptional control over extensive manipulation. This makes it a robust tool for both academic research and creative industries. Code is available at https://github.com/liyiming09/RMD. Peng Zhou 0010, Hongwei Hu, Xiaokang Qin, Jun Sun 0005, Yi Xu 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Position-LoRA: Enhanced Relation Customization through Structural Prior in Initial Latent NoiseabstractRecent advancements in concept customization via diffusion models have significantly enhanced controllability and quality. However, precise relation customization, which controls the position of interactions among multiple instances, remains challenging due to unpredictable initial latent noise. Existing methods primarily rely on conditional prompts and attention control, overlooking the structured potential of initial noise. This paper introduces Position-LoRA, a novel framework leveraging structural prior in initial noise to improve relation customization and layout control. Position-LoRA employs a differential fine-tuning scheme and a latent noise encoder. The guided fine-tuning enhances generation tendencies from structured initial noise, embedding explicit relationship-specific spatial information. The latent noise encoder dynamically manipulates latent noises, enabling precise spatial control and flexibility in relational image generation. Furthermore, a fine-grained guidance and control strategy is employed during generation to enhance the image-text alignment and layout alignment. Experiments demonstrate that Position-LoRA improves stability, controllability, and fidelity in relational image generation with layout control, surpassing existing concept customization and layout-to-image methods in qualitative and quantitative evaluations. Code is available at https://github.com/liyiming09/Position-LoRA. Peng Zhou 0010, Xiaokang Qin, Hongwei Hu, Jun Sun 0005, Yi Xu 0001 |
ACM Multimedia | 6 |
| 2025 | SDRNet: Joint Modeling of Static and Dynamic Representations for Breast Ultrasound DiagnosisabstractBreast ultrasound (BUS) diagnosis requires synergistic analysis of static images (capturing 2D features like microcalcifications and boundaries) and dynamic videos (revealing 3D structural dynamics such as ductal infiltration and fluid mobility), particularly for complex lesions including intraductal carcinomas and phyllodes tumors. Existing deep learning methods predominantly focus on single modalities, limiting diagnostic accuracy. To address this, we propose SDR-Net, a dual-branch framework integrating three innovations: 1) a Medical Prior-Guided Sampling (MPGS) module that automatically selects diagnostically critical video keyframes; 2) cross-modal attention fusion enabling bidirectional alignment of static spatial details with spatiotemporal video patterns; and 3) Conditional Layer Normalization (CLN) that injects static anatomical context into video feature learning, enhancing fine-grained cross-modal alignment. A temporal attention module further prioritizes diagnostically salient video segments. Evaluated on a BUS dataset containing challenging dual-modality cases, SDRNet achieves state-of-the-art performance (92.6% AUC, 87.0% accuracy and 88.2% malignant F1-score), outperforming single-modality models by 2.8 % in AUC and late-fusion baselines by 2.6%. Yi Xu 0001 |
SMC | 4 |
| 2023 | Interventional Bag Multi-Instance Learning On Whole-Slide Pathological ImagesabstractMulti-instance learning (MIL) is an effective paradigm for whole-slide pathological images (WSIs) classification to handle the gigapixel resolution and slide-level label. Prevailing MIL methods primarily focus on improving the feature extractor and aggregator. However, one deficiency of these methods is that the bag contextual prior may trick the model into capturing spurious correlations between bags and labels. This deficiency is a confounder that limits the performance of existing MIL methods. In this paper, we propose a novel scheme, Interventional Bag Multi-Instance Learning (IBMIL), to achieve deconfounded bag-level prediction. Unlike traditional likelihood-based strategies, the proposed scheme is based on the backdoor adjustment to achieve the interventional training, thus is capable of suppressing the bias caused by the bag contextual prior. Note that the principle of IBMIL is orthogonal to existing bag MIL methods. Therefore, IBMIL is able to bring consistent performance boosting to existing schemes, achieving new state-of-the-art performance. Code is available at https://github.com/HHHedo/IBMIL. Tiancheng Lin 0001, Zhimiao Yu, Hongyu Hu, Yi Xu 0001, Chang Wen Chen |
CVPR | 4 |
| 2023 | Background Clustering Pre-Training for Few-Shot SegmentationabstractRecent few-shot segmentation (FSS) methods introduce an extra pre-training stage before meta-training to obtain a stronger backbone, which has become a standard step in few-shot learning. Despite the effectiveness, current pre-training scheme suffers from the merged background problem: only base classes are labelled as foregrounds, making it hard to distinguish between novel classes and actual background. In this paper, we propose a new pre-training scheme for FSS via decoupling the novel classes from background, called Background Clustering Pre-Training (BCPT). Specifically, we adopt online clustering to the pixel embeddings of merged background to explore the underlying semantic structures, bridging the gap between pre-training and adaptation to novel classes. Given the clustering results, we further propose the background mining loss and leverage base classes to guide the clustering process, improving the quality and stability of clustering results. Experiments on PASCAL-5iand COCO-20ishow that BCPT yields advanced performance. Code will be available at https://github.com/Carboxy/BCPT. Zhimiao Yu, Tiancheng Lin 0001, Yi Xu 0001 |
ICIP | 3 |
| 2023 | SLPD: Slide-Level Prototypical Distillation for WSIsabstractImproving the feature representation ability is the foundation of many whole slide pathological image (WSIs) tasks. Recent works have achieved great success in pathological-specific self-supervised learning (SSL). However, most of them only focus on learning patch-level representations, thus there is still a gap between pretext and slide-level downstream tasks, e.g., subtyping, grading and staging. Aiming towards slide-level representations, we propose Slide-Level Prototypical Distillation (SLPD) to explore intra- and inter-slide semantic structures for context modeling on WSIs. Specifically, we iteratively perform intra-slide clustering for the regions (4096 $$\times $$ 4096 patches) within each WSI to yield the prototypes and encourage the region representations to be closer to the assigned prototypes. By representing each slide with its prototypes, we further select similar slides by the set distance of prototypes and assign the regions by cross-slide prototypes for distillation. SLPD achieves state-of-the-art results on multiple slide-level benchmarks and demonstrates that representation learning of semantic structures of slides can make a suitable proxy task for WSI analysis. Code will be available at https://github.com/Carboxy/SLPD . Zhimiao Yu, Tiancheng Lin 0001, Yi Xu 0001 |
MICCAI (1) | 3 |
| 2023 | Relational Contrastive Learning for Scene Text RecognitionabstractContext-aware methods achieved great success in supervised scene text recognition via incorporating semantic priors from words. We argue that such prior contextual information can be interpreted as the relations of textual primitives due to the heterogeneous text and background, which can provide effective self-supervised labels for representation learning. However, textual relations are restricted to the finite size of dataset due to lexical dependencies, which causes the problem of over-fitting and compromises representation robustness. To this end, we propose to enrich the textual relations via rearrangement, hierarchy and interaction, and design a unified framework called RCLSTR: Relational Contrastive Learning for Scene Text Recognition. Based on causality, we theoretically explain that three modules suppress the bias caused by the contextual prior and thus guarantee representation robustness. Experiments on representation quality show that our method outperforms state-of-the-art self-supervised STR methods. Code is available at https://github.com/ThunderVVV/RCLSTR. Jinglei Zhang 0003, Tiancheng Lin 0001, Yi Xu 0001, Kai Chen 0006, Rui Zhang 0052 |
ACM Multimedia | 3 |
| 2023 | O2M-UDA: Unsupervised dynamic domain adaptation for one-to-multiple medical image segmentation
Ziyue Jiang 0004, Yuting He 0001, Xiaomei Zhu, Yi Xu 0001, Yang Chen 0008, Jean-Louis Coatrieux, Shuo Li 0001, Guanyu Yang 0001 |
Knowl. Based Syst. | 6 |
| 2023 | SGCL: Spatial guided contrastive learning on whole-slide pathological images
Tiancheng Lin 0001, Zhimiao Yu, Zengchao Xu, Hongyu Hu, Yi Xu 0001, Chang Wen Chen |
Medical Image Anal. | 5 |
| 2023 | Fast Quaternion Product Units for Learning Disentangled Representations in $\mathbb {SO}(3)$abstractReal-world 3D structured data like point clouds and skeletons often can be represented as data in a 3D rotation group (denoted as [Formula: see text]). However, most existing neural networks are tailored for the data in the euclidean space, which makes the 3D rotation data not closed under their algebraic operations and leads to sub-optimal performance in 3D-related learning tasks. To resolve the issues caused by the above mismatching between data and model, we propose a novel non-real neuron model called quaternion product unit (QPU) to represent data on 3D rotation groups. The proposed QPU leverages quaternion algebra and the law of the 3D rotation group, representing 3D rotation data as quaternions and merging them via a weighted chain of Hamilton products. We demonstrate that the QPU mathematically maintains the [Formula: see text] structure of the 3D rotation data during the inference process and disentangles the 3D representations into "rotation-invariant" features and "rotation-equivariant" features, respectively. Moreover, we design a fast QPU to accelerate the computation of QPU. The fast QPU applies a tree-structured data indexing process, and accordingly, leverages the power of parallel computing, which reduces the computational complexity of QPU in a single thread from O(N) to O(logN). Taking the fast QPU as a basic module, we develop a series of quaternion neural networks (QNNs), including quaternion multi-layer perceptron (QMLP), quaternion message passing (QMP), and so on. In addition, we make the QNNs compatible with conventional real-valued neural networks and applicable for both skeletons and point clouds. Experiments on synthetic and real-world 3D tasks show that the QNNs based on our fast QPUs are superior to state-of-the-art real-valued models, especially in the scenarios requiring the robustness to random rotations. The code of this work is available at https://github.com/SuferQin/Fast-QPU. Shaofei Qin, Hongteng Xu, Yi Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Interventional Multi-Instance Learning with Deconfounded Instance-Level PredictionabstractWhen applying multi-instance learning (MIL) to make predictions for bags of instances, the prediction accuracy of an instance often depends on not only the instance itself but also its context in the corresponding bag. From the viewpoint of causal inference, such bag contextual prior works as a confounder and may result in model robustness and interpretability issues. Focusing on this problem, we propose a novel interventional multi-instance learning (IMIL) framework to achieve deconfounded instance-level prediction. Unlike traditional likelihood-based strategies, we design an Expectation-Maximization (EM) algorithm based on causal intervention, providing a robust instance selection in the training phase and suppressing the bias caused by the bag contextual prior. Experiments on pathological image analysis demonstrate that our IMIL method substantially reduces false positives and outperforms state-of-the-art MIL methods. Tiancheng Lin 0001, Hongteng Xu, Canqian Yang, Yi Xu 0001 |
AAAI | 4 |
| 2022 | Anatomy-Aware Self-Supervised Learning for Aligned Multi-Modal Medical Data
Hongyu Hu, Tiancheng Lin 0001, Yuanfan Guo, Yi Xu 0001 |
BMVC | 6 |
| 2022 | HCSC: Hierarchical Contrastive Selective CodingabstractHierarchical semantic structures naturally exist in an image dataset, in which several semantically relevant image clusters can be further integrated into a larger cluster with coarser-grained semantics. Capturing such structures with image representations can greatly benefit the semantic understanding on various downstream tasks. Existing contrastive representation learning methods lack such an important model capability. In addition, the negative pairs used in these methods are not guaranteed to be semantically distinct, which could further hamper the structural correctness of learned image representations. To tackle these limitations, we propose a novel contrastive learning framework called Hierarchical Contrastive Selective Coding (HCSC). In this framework, a set of hierarchical prototypes are constructed and also dynamically updated to represent the hierarchical semantic structures underlying the data in the latent space. To make image representations better fit such semantic structures, we employ and further improve conventional instance-wise and prototypical contrastive learning via an elaborate pair selection scheme. This scheme seeks to select more diverse positive pairs with similar semantics and more precise negative pairs with truly distinct semantics. On extensive downstream tasks, we verify the state-of-the-art performance of HCSC and also the effectiveness of major model components. We are continually building a comprehensive model zoo (see supplementary material). Our source code and model weights are available at https://github.com/gyfastas/HCSC. Yuanfan Guo, Bingbing Ni, Zhenbang Sun, Yi Xu 0001 |
CVPR | 7 |
| 2022 | AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image EnhancementabstractThe 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhancement but neglect the importance of sampling strategy. They adopt a sub-optimal uniform sampling point allocation, limiting the expressiveness of the learned LUTs since the (tri-)linear interpolation between uniform sampling points in the LUT transform might fail to model local non-linearities of the color transform. Focusing on this problem, we present AdaInt (Adaptive Intervals Learning), a novel mechanism to achieve a more flexible sampling point allocation by adaptively learning the non-uniform sampling intervals in the 3D color space. In this way, a 3D LUT can increase its capability by conducting dense sampling in color ranges requiring highly non-linear transforms and sparse sampling for near-linear transforms. The proposed AdaInt could be implemented as a compact and efficient plug-and-play module for a 3D LUT-based method. To enable the end-to-end learning of AdaInt, we design a novel differentiable operator called AiLUT-Transform (Adaptive Interval LUT Transform) to locate input colors in the non-uniform 3D LUT and provide gradients to the sampling intervals. Experiments demonstrate that methods equipped with AdaInt can achieve state-of-the-art performance on two public benchmark datasets with a negligible overhead increase. Our source code is available at https://github.com/ImCharlesY/AdaInt. Canqian Yang, Meiguang Jin, Xu Jia 0012, Yi Xu 0001, Ying Chen 0011 |
CVPR | 4 |
| 2022 | Posterior Refinement on Metric Matrix Improves Generalization Bound in Metric Learning
Canqian Yang, Yi Xu 0001 |
ECCV (26) | 3 |
| 2022 | SepLUT: Separable Image-Adaptive Lookup Tables for Real-Time Image Enhancement
Canqian Yang, Meiguang Jin, Yi Xu 0001, Rui Zhang 0052, Ying Chen 0011, Huaida Liu |
ECCV (18) | 3 |
| 2022 | Solving The Long-Tailed Problem Via Intra- And Inter-Category BalanceabstractBenchmark datasets for visual recognition assume that data is uniformly distributed, while real-world datasets obey long-tailed distribution. Current approaches handle the long-tailed problem to transform the long-tailed dataset to uniform distribution by re-sampling or re-weighting strategies. These approaches emphasize the tail classes but ignore the hard examples in head classes, which result in performance degradation. In this paper, we propose a novel gradient harmonized mechanism with category-wise adaptive precision to decouple the difficulty and sample size imbalance in the long-tailed problem, which are correspondingly solved via intra- and inter-category balance strategies. Specifically, intra-category balance focuses on the hard examples in each category to optimize the decision boundary, while inter-category balance aims to correct the shift of decision boundary by taking each category as a unit. Extensive experiments demonstrate that the proposed method consistently outperforms other approaches on all the datasets. Renhui Zhang, Tiancheng Lin 0001, Rui Zhang 0052, Yi Xu 0001 |
ICASSP | 4 |
| 2022 | BKC-Net: Bi-Knowledge Contrastive Learning for renal tumor diagnosis on 3D CT images
Jindi Kong, Yuting He 0001, Xiaomei Zhu, Yi Xu 0001, Yang Chen 0008, Jean-Louis Coatrieux, Guanyu Yang 0001 |
Knowl. Based Syst. | 5 |
| 2022 | Fine-Grained Video Captioning via Graph-based Multi-Granularity Interaction LearningabstractLearning to generate continuous linguistic descriptions for multi-subject interactive videos in great details has particular applications in team sports auto-narrative. In contrast to traditional video caption, this task is more challenging as it requires simultaneous modeling of fine-grained individual actions, uncovering of spatio-temporal dependency structures of frequent group interactions, and then accurate mapping of these complex interaction details into long and detailed commentary. To explicitly address these challenges, we propose a novel framework Graph-based Learning for Multi-Granularity Interaction Representation (GLMGIR) for fine-grained team sports auto-narrative task. A multi-granular interaction modeling module is proposed to extract among-subjects' interactive actions in a progressive way for encoding both intra- and inter-team interactions. Based on the above multi-granular representations, a multi-granular attention module is developed to consider action/event descriptions of multiple spatio-temporal resolutions. Both modules are integrated seamlessly and work in a collaborative way to generate the final narrative. In the meantime, to facilitate reproducible research, we collect a new video dataset from YouTube.com called Sports Video Narrative dataset (SVN). It is a novel direction as it contains 6K team sports videos (i.e., NBA basketball games) with 10K ground-truth narratives(e.g., sentences). Furthermore, as previous metrics such as METEOR (i.e., used in coarse-grained video caption task) DO NOT cope with fine-grained sports narrative task well, we hence develop a novel evaluation metric named Fine-grained Captioning Evaluation (FCE), which measures how accurate the generated linguistic description reflects fine-grained action details as well as the overall spatio-temporal interactional structure. Extensive experiments on our SVN dataset have demonstrated the effectiveness of the proposed framework for fine-grained team sports video auto-narrative. Yichao Yan, Ning Zhuang, Bingbing Ni, Jian Zhang 0079, Qi Tian 0001, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2022 | MVSGAN: Spatial-Aware Multi-View CMR Fusion for Accurate 3D Left Ventricular Myocardium SegmentationabstractThe accurate 3D left ventricular (LV) myocardium segmentation in short-axis (SAX) view of cardiac magnetic resonance (CMR) is challenged by the sparse spatial structure of CMR. The strategy of multi-view CMR fusion can provide fine-grained spatial structure for accurate segmentation. However, the large information misalignment and lack of dense 3D CMR as fusion target in multi-view CMR fusion, and the different spatial resolution between the fusion result and the ground truth in segmentation limit the strategy. In this study, we propose a multi-view spatial-aware adversarial network (MVSGAN). It studies the perception of fine-grained cardiac structure for accurate segmentation by the spatialaware multi-view CMR fusion. It consists of three modules: (1) A residual adversarial fusion (RAF) module takes inter-slices deep correlation and anatomical prior to refine the spatial structures by residual supplement and adversarial optimization. (2) A structural perception-aggregation (SPA) module establishes the spatial correlation between the dense cardiac model and sparse label for accurate CMR LV myocardium segmentation. (3) A joint training strategy utilizes the dense SAX volume as explicit and implicit goals to jointly optimize the framework. The experiments are applied on a public dataset and a clinical dataset to evaluate the performance of MVSGAN. The average Dice and Jaccard score of LV myocardium segmentation obtained by MVSGAN are highest among seven existing state-of-the-art methods, which are up to 0.92 and 0.75. It is concluded that the spatial-aware multi-view CMR fusion can provide meaningful spatial correlation for accurate LV myocardium segmentation. Xiaoming Qi, Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Jian Yang 0009, Wangyag Liu, Yinsu Zhu, Yi Xu 0001, Huazhong Shu, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2021 | Self-Supervised Disentangled Embedding For Robust Image Classification
Lanqing Liu, Zhenyu Duan, Guozheng Xu, Yi Xu 0001 |
ICIP | 4 |
| 2021 | Enhanced Breast Lesion Classification via Knowledge Guided Cross-Modal and Semantic Data Augmentation
Kun Chen 0016, Yuanfan Guo, Canqian Yang, Yi Xu 0001, Rui Zhang 0052 |
MICCAI (5) | 4 |
| 2021 | SSLP: Spatial Guided Self-supervised Learning on Pathological Images
Tiancheng Lin 0001, Yi Xu 0001 |
MICCAI (2) | 3 |
| 2021 | Decoupled gradient harmonized detector for partial annotation: Application to signet ring cell detection
Tiancheng Lin 0001, Yuanfan Guo, Canqian Yang, Jiancheng Yang, Yi Xu 0001 |
Neurocomputing | 5 |
| 2021 | Using BI-RADS Stratifications as Auxiliary Information for Breast Masses Classification in Ultrasound ImagesabstractBreast Ultrasound (BUS) imaging has been recognized as an essential imaging modality for breast masses classification in China. Current deep learning (DL) based solutions for BUS classification seek to feed ultrasound (US) images into deep convolutional neural networks (CNNs), to learn a hierarchical combination of features for discriminating malignant and benign masses. One existing problem in current DL-based BUS classification was the lack of spatial and channel-wise features weighting, which inevitably allow interference from redundant features and low sensitivity. In this study, we aim to incorporate the instructive information provided by breast imaging reporting and data system (BI-RADS) within DL-based classification. A novel DL-based BI-RADS Vector-Attention Network (BVA Net) that trains with both texture information and decoded information from BI-RADS stratifications was proposed for the task. Three baseline models, pre-trained DenseNet-121, ResNet-50 and Residual-Attention Network (RA Net) were included for comparison. Experiments were conducted on a large scale private main dataset and two public datasets, UDIAT and BUSI. On the main dataset, BVA Net outperformed other models, in terms of AUC (area under the receiver operating curve, 0.908), ACC (accuracy, 0.865), sensitivity (0.812) and precision (0.795). BVA Net also achieved the high AUC (0.87 and 0.882) and ACC (0.859 and 0.843), on UDIAT and BUSI. Moreover, we proposed a method that integrates both BVA Net binary classification and BI-RADS stratification estimation, called integrated classification. The introduction of integrated classification helped improving the overall sensitivity while maintaining a high specificity. Qinyang Lu, Aijun Yu, Yi Xu 0001, Xiaoling Xia, Yue Sun 0001, Jing Xiao 0006, Lingyun Huang |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | Quaternion Product Units for Deep Learning on 3D Rotation GroupsabstractWe propose a novel quaternion product unit (QPU) to represent data on 3D rotation groups. The QPU leverages quaternion algebra and the law of 3D rotation group, representing 3D rotation data as quaternions and merging them via a weighted chain of Hamilton products. We prove that the representations derived by the proposed QPU can be disentangled into "rotation-invariant" features and "rotation-equivariant" features, respectively, which supports the rationality and the efficiency of the QPU in theory. We design quaternion neural networks based on our QPUs and make our models compatible with existing deep learning models. Experiments on both synthetic and real-world data show that the proposed QPU is beneficial for the learning tasks requiring rotation robustness. Shaofei Qin, Yi Xu 0001, Hongteng Xu |
CVPR | 3 |
| 2020 | MUGGLE: MUlti-Stream Group Gaze Learning and EstimationabstractBeing able to accurately predict the common gaze point of a group of persons is of particular interest to precise marketing and automatic group attention assessment. Group gaze estimation faces challenges including small face/head size and outlier observers. To address these challenges, we proposed a novel framework called Multi-stream Group Gaze Learning and Estimation (MUGGLE). The MUGGLE infrastructure includes two inference streams: 1) a holistic stream which utilizes fused attention map as input to a global deep convolutional structure to explore the global geometric configurations and contexts of interesting persons in the scene; and 2) an aggregative stream which robustly aggregates individual gazes via a recurrent structure (e.g., LSTM) to obtain outlier-tolerant estimation. Both streams are seamlessly integrated via a fusion network. Extensive experiments are performed on a fully annotated group gaze image dataset with 8,000+ images and 100,000+ faces (which is publicly releasable). The results demonstrate the effectiveness of the proposed MUGGLE framework in group gaze estimation. Ning Zhuang, Bingbing Ni, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001, Zefan Li, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Single-Image Rain Removal Via Multi-Scale Cascading Image GenerationabstractA novel single-image rain removal method is proposed based on multi-scale cascading image generation (MSCG). In particular, the proposed method consists of an encoder extracting multi-scale features from images and a decoder generating de-rained images with a cascading mechanism. The encoder ensembles the convolution neural networks using the kernels with different sizes, and integrates their outputs across different scales. The decoder implements a coarse-to-fine image generation framework, adding fine details incrementally to the final de-rained images according to the spatial contextual information on different scales. We test the proposed method on both synthetic and real-world datasets. Experimental results show that the proposed method is robust to the changes of scene, e.g., the viewpoint and the depth, the heaviness of rain, etc., which suppresses the blurring problem of de-rained image and outperforms state-of-the-art methods consistently. Yi Xu 0001, Bingbing Ni, Hongteng Xu |
ICIP | 2 |
| 2019 | Cross Modality Alignment of Medical Volumes using Spatio-Semantic Attentive Cycle-GANabstractLack of large available datasets fully annotated is a fundamental bottleneck in pulmonary nodule detection, especially when the corresponding computed tomography(CT) images obtained are device-dependent. We propose a spatio-semantic attentive CycleGAN (SSA-CycleGAN) capable of aligning modalities, as well as distinguishing nodule vs. non-nodule, which in turn achieves effective data augmentation. Specifically, a novel training loss function is established, providing a constraint for semantic preservation and local fidelity of nodule regions. Extensive experimental results on varied datasets demonstrate the proposed framework achieves significant performance gain on pulmonary nodule detection. Xiaohui Lin 0007, Yi Xu 0001, Bingbing Ni, Xiaokang Yang 0001, Guangyu Tao, Xiaodan Ye |
ICME | 2 |
| 2019 | Probabilistic Radiomics: Ambiguous Diagnosis with Controllable Shape Analysis
Jiancheng Yang, Rongyao Fang, Bingbing Ni, Yi Xu 0001, Linguo Li |
MICCAI (6) | 5 |
| 2019 | Recognition oriented facial image quality assessment via deep convolutional neural network
Ning Zhuang, Cenhui Pan, Bingbing Ni, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
Neurocomputing | 5 |
| 2018 | Crowd Counting via Adversarial Cross-Scale Consistency PursuitabstractCrowd counting or density estimation is a challenging task in computer vision due to large scale variations, perspective distortions and serious occlusions, etc. Existing methods generally suffer from two issues: 1) the model averaging effects in multi-scale CNNs induced by the widely adopted ℓ2regression loss; and 2) inconsistent estimation across different scaled inputs. To explicitly address these issues, we propose a novel crowd counting (density estimation) framework called Adversarial Cross-Scale Consistency Pursuit (ACSCP). On one hand, a U-net structured generation network is designed to generate density map from input patch, and an adversarial loss is directly employed to shrink the solution onto a realistic subspace, thus attenuating the blurry effects of density map estimation. On the other hand, we design a novel scale-consistency regularizer which enforces that the sum up of the crowd counts from local patches (i.e., small scale) is coherent with the overall count of their region union (i.e., large scale). The above losses are integrated via a joint training scheme, so as to help boost density estimation performance by further exploring the collaboration between both objectives. Extensive experiments on four benchmarks have well demonstrated the effectiveness of the proposed innovations as well as the superior performance over prior art. Zan Shen, Yi Xu 0001, Bingbing Ni, Minsi Wang, Jianguo Hu, Xiaokang Yang 0001 |
CVPR | 2 |
| 2018 | Scale-Transferrable Object DetectionabstractScale problem lies in the heart of object detection. In this work, we develop a novel Scale-Transferrable Detection Network (STDN) for detecting multi-scale objects in images. In contrast to previous methods that simply combine object predictions from multiple feature maps from different network depths, the proposed network is equipped with embedded super-resolution layers (named as scale-transfer layer/module in this work) to explicitly explore the interscale consistency nature across multiple detection scales. Scale-transfer module naturally fits the base network with little computational cost. This module is further integrated with a dense convolutional network (DenseNet) to yield a one-stage object detector. We evaluate our proposed architecture on PASCAL VOC 2007 and MS COCO benchmark tasks and STDN obtains significant improvements over the comparable state-of-the-art detection models. Peng Zhou 0010, Bingbing Ni, Cong Geng, Jianguo Hu, Yi Xu 0001 |
CVPR | 5 |
| 2018 | Geometric Constrained Joint Lane Segmentation and Lane Boundary Detection
Yi Xu 0001, Bingbing Ni, Zhenyu Duan |
ECCV (1) | 2 |
| 2018 | Quaternion Convolutional Neural Networks
Yi Xu 0001, Hongteng Xu, Changjian Chen |
ECCV (8) | 2 |
| 2018 | Flexible Network Binarization with Layer-Wise PriorityabstractHow to effectively approximate real-valued parameters with binary codes plays a central role in neural network binarization. In this work, we reveal an important fact that binarizing different layers has a widely varied effect on the compression ratio of network and the loss of performance. Based on this fact, we propose a novel and flexible neural network binarization method by introducing the concept of layer-wise priority which binarizes parameters in inverse order of their layer depth. In each training step, our method selects a specific network layer, minimizes the discrepancy between the original real-valued weights and its binary approximations following block coordinate descent scheme. During the iteration of the above process, it is significant that we can flexibly decide whether to binarize the remaining floating layers or not and explore a trade-off between the loss of performance and the compression ratio of model. The resulting binary network is applied for efficient pedestrian detection. Experimental results on several benchmarks show that under the same compression ratio, the model compressed by our method achieves much lower miss rate and faster detection speed than those obtained by the state-of-the-art neural network binarization method. Yi Xu 0001, Bingbing Ni, Lixue Zhuang, Hongteng Xu |
ICIP | 2 |
| 2017 | Video Segmentation via Multiple Granularity AnalysisabstractWe introduce a Multiple Granularity Analysis framework for video segmentation in a coarse-to-fine manner. We cast video segmentation as a spatio-temporal superpixel labeling problem. Benefited from the bounding volume provided by off-the-shelf object trackers, we estimate the foreground/ background super-pixel labeling using the spatiotemporal multiple instance learning algorithm to obtain coarse foreground/background separation within the volume. We further refine the segmentation mask in the pixel level using the graph-cut model. Extensive experiments on benchmark video datasets demonstrate the superior performance of the proposed video segmentation algorithm. Bingbing Ni, Chao Ma 0004, Yi Xu 0001, Xiaokang Yang 0001 |
CVPR | 4 |
| 2017 | Pedestrian Detection via Bi-directional Multi-scale AnalysisabstractScale analysis plays a vital role in pedestrian detection. Conventional approaches usually directly concatenate multi-scale outputs, which is only capable of modeling first-order dependency among various scales. In contrast, this work proposes a novel scale-context modeling scheme by exploiting the highly nonlinear dependency among scales. The proposed scheme aggregates output response maps from mid-results of convolutional layers via a bi-directional recurrent sub-network. Therefore scale information could flow among different layers and implicit underlying dependency structure information in the scale space would be disclosed, which yields more consistency detection. Experimental results on Caltech Pedestrian detection benchmark demonstrate the superior detection (state-of-the-art miss rate of 8.56%) of the proposed method over prior art. Zhenyu Duan, Jinpeng Lan, Yi Xu 0001, Bingbing Ni, Lixue Zhuang, Xiaokang Yang 0001 |
ACM Multimedia | 3 |
| 2017 | Deep Cross-Modality Alignment for Multi-Shot Person Re-IDentificationabstractMulti-shot person Re-IDentification (Re-ID) has recently received more research attention as its problem setting is more realistic compared to single-shot Re-ID in terms of application. While many large-scale single-shot Re-ID human image datasets have been released, most existing multishot Re-ID video sequence datasets containonly a few (i.e., several hundreds) human instances, which hinders further improvement of multi-shot Re-ID performance. To this end, we propose a deep cross-modality alignment network, which jointly explores both human sequence pairs and image pairs to facilitate training better multi-shot human Re-ID models, i.e., via transferring knowledge from image data to sequence data. To mitigate modality-to-modality mismatch issue, the proposed network is equipped with an image-to-sequence adaption module called cross-modality alignment sub-network, which successfully maps each human image into a pseudo human sequence to facilitate knowledge transferring and joint training. Extensive experimental results on several multi-shot person Re-ID benchmarks demonstrate great performance gain brought up by the proposed network. Zhichao Song, Bingbing Ni, Yichao Yan, Zhe Ren, Yi Xu 0001, Xiaokang Yang 0001 |
ACM Multimedia | 5 |
| 2017 | Enhancing pulmonary nodule detection via cross-modal alignmentabstractLack of large available datasets fully annotated is a fundamental bottleneck in pulmonary nodule detection, especially when the sensing equipment and the corresponding computed tomography (CT) images obtained are device dependent. This work presents a novel cross modal scheme, pursuing modal alignment, to facilitate our aggregate channel detector training. Named as multi-class cycle-consistent adversarial network (CycleGAN), our proposed framework utilizes a generative adversarial model to transfer nodule morphological characteristics from source modal to target modal, and we propose an end to end objective function to unify the transfer and detection procedures. The outputs of the two parts are combined with a dedicated fusion method for final classification. Extensive experimental results on 1948 scans of the private dataset demonstrate the proposed modal transfer method is very effective in data augmentation. Yumeng Zhu, Yi Xu 0001, Bingbing Ni, Xiaokang Yang 0001 |
VCIP | 2 |
| 2016 | Exploiting neural models for no-reference image quality assessmentabstractWe propose an improved algorithm for no-reference image quality assessment (NR-IQA) using the convolutional neural network (CNN) and neural theory based saliency detection. Firstly, we extract non-overlapping patches from the input image. For each patch, we obtain the quality score by CNN network, which consists of seven layers and integrates feature learning and regression into image patch quality estimation. Considering that the patches attracting much attention take significant role in visual perception, an efficient technique based on free energy based neural model is used to detect the saliency map. This saliency map is then applied as a weighting mask to output the quality score of the whole image. Results of experiments show that our algorithm achieves state-of-the-art performance, as compared with the prevailing IQA methods. Cenhui Pan, Yi Xu 0001, Yichao Yan, Ke Gu 0001, Xiaokang Yang 0001 |
VCIP | 2 |
| 2016 | When Correlation Filters Meet Convolutional Neural Networks for Visual TrackingabstractCorrelation filters have been widely applied to visual tracking in recent years as adaptive correlation filters with short-term memory are robust to large appearance changes. However, tracking methods relying on correlation filters are prone to drifting due to noisy updates. Moreover, these methods are unable to recover from tracking failures caused by temporary or persistent heavy occlusions. In this paper, we interpret correlation filters as the counterparts of convolution filters in deep neural networks. Correlation filters encode the holistic template of target appearance, while convolution filters with smaller size encode the part-based template. In the light of this idea, we propose to exploit deep convolutional networks that directly learn mapping as a spatial correlation between two consecutive frames for visual tracking. We show that these deeply learned networks are effective in maintaining the long-term memory of target appearance for handling heavy occlusion or out-of-view. We further take the response maps both from the deep networks and conventional correlation filters into account for precisely locating the target. Experimental results on large-scale benchmark sequences show that the proposed algorithm performs favorably against the state-of-the-art methods. Chao Ma 0004, Yi Xu 0001, Bingbing Ni, Xiaokang Yang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2015 | Dictionary Learning with Mutually Reinforcing Group-Graph StructuresabstractIn this paper, we propose a novel dictionary learning method in the semi-supervised setting by dynamically coupling graph and group structures. To this end, samples are represented by sparse codes inheriting their graph structure while the labeled samples within the same class are represented with group sparsity, sharing the same atoms of the dictionary. Instead of statically combining graph and group structures, we take advantage of them in a mutually reinforcing way — in the dictionary learning phase, we introduce the unlabeled samples into groups by an entropy-based method and then update the corresponding local graph, resulting in a more structured and discriminative dictionary. We analyze the relationship between the two structures and prove the convergence of our proposed method. Focusing on image classification task, we evaluate our approach on several datasets and obtain superior performance compared with the state-of-the-art methods, especially in the case of only a few labeled samples and limited dictionary size. Hongteng Xu, Licheng Yu, Dixin Luo, Hongyuan Zha, Yi Xu 0001 |
AAAI | 5 |
| 2015 | Real time and scene invariant crowd counting: Across a line or inside a regionabstractIn this paper, we propose a blob-based method of crowd counting across a line of interest (LOI), which can be further extended to counting inside a region of interest (ROI). Firstly, we detect moving blobs in which low-level features are extracted and grouped. Since features vary with different walking pace, blob velocity is estimated using optical flow, and principal velocity component is further extracted in case of the interference of local articulated motion. Besides, spatial normalization is implemented to compensate for different image depth and different scenes. Finally, we apply Gaussian Process Regression to model the global linear and local nonlinear relationships between the extracted features and crowd counts. The experimental results demonstrate that the proposed method has good applicability to both LOI and ROI crowd counting. As compared with the state-of-the-art methods, our method achieves higher counting accuracy over two representative datasets and meanwhile processes much faster. Mingjie Deng, Yi Xu 0001, Pufan Jiang, Xiaokang Yang 0001 |
MMSP | 2 |
| 2015 | Vector Sparse Representation of Color Image Using Quaternion Matrix AnalysisabstractTraditional sparse image models treat color image pixel as a scalar, which represents color channels separately or concatenate color channels as a monochrome image. In this paper, we propose a vector sparse representation model for color images using quaternion matrix analysis. As a new tool for color image representation, its potential applications in several image-processing tasks are presented, including color image reconstruction, denoising, inpainting, and super-resolution. The proposed model represents the color image as a quaternion matrix, where a quaternion-based dictionary learning algorithm is presented using the K-quaternion singular value decomposition (QSVD) (generalized K-means clustering for QSVD) method. It conducts the sparse basis selection in quaternion space, which uniformly transforms the channel images to an orthogonal color space. In this new color space, it is significant that the inherent color structures can be completely preserved during vector reconstruction. Moreover, the proposed sparse model is more efficient comparing with the current sparse models for image restoration tasks due to lower redundancy between the atoms of different color channels. The experimental results demonstrate that the proposed sparse image model avoids the hue bias issue successfully and shows its potential as a general and powerful tool in color image analysis and processing domain. Yi Xu 0001, Licheng Yu, Hongteng Xu, Truong Q. Nguyen |
IEEE Trans. Image Process. | 1 |
| 2015 | Single Image Superresolution Based on Gradient Profile SharpnessabstractSingle image superresolution is a classic and active image processing problem, which aims to generate a high-resolution (HR) image from a low-resolution input image. Due to the severely under-determined nature of this problem, an effective image prior is necessary to make the problem solvable, and to improve the quality of generated images. In this paper, a novel image superresolution algorithm is proposed based on gradient profile sharpness (GPS). GPS is an edge sharpness metric, which is extracted from two gradient description models, i.e., a triangle model and a Gaussian mixture model for the description of different kinds of gradient profiles. Then, the transformation relationship of GPSs in different image resolutions is studied statistically, and the parameter of the relationship is estimated automatically. Based on the estimated GPS transformation relationship, two gradient profile transformation models are proposed for two profile description models, which can keep profile shape and profile gradient magnitude sum consistent during profile transformation. Finally, the target gradient field of HR image is generated from the transformed gradient profiles, which is added as the image prior in HR image reconstruction model. Extensive experiments are conducted to evaluate the proposed algorithm in subjective visual effect, objective quality, and computation time. The experimental results demonstrate that the proposed approach can generate superior HR images with better visual quality, lower reconstruction error, and acceptable computation efficiency as compared with state-of-the-art works. Yi Xu 0001, Xiaokang Yang 0001, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2014 | Single color image super-resolution using quaternion-based sparse representationabstractIn current color image super-resolution methods, superresolution based on sparse representation achieves state-of-the-art performance. However, the exploited sparse representation models deal with the color images as independent channel planes. Consequently, these approaches process the color pixels as scalar quantity, lacking of accuracy in describing inter-relationship among color channels. In this paper, we propose a quaternion-based online dictionary learning method and solve color image super-resolution by employing a quaternion-based sparse representation model. This sparse representation model implements color image superresolution in a kind of vectorial reconstruction, effectively accounting for both luminance and chrominance geometry in images. The proposed color image super-resolution method can better describe the inter-channel changes. In the case that changing lighting conditions affect color more than the luminance perception, it can obtain superior performance comparing to the methods based on monochromatic sparse models with 1dB improvement. Mengqi Yu, Yi Xu 0001 |
ICASSP | 2 |
| 2014 | Blind image quality assessment based on a new feature of nature scene statisticsabstractA recently proposed model, known as blind/referenceless image spatial quality evaluator (BRISQUE), achieves the state-of-the-art performance in context of blind image quality assessment (IQA). This model used the predefined generalized Gaussian distribution (GGD) to describe the regularity of natural scene statistics, introducing fitting errors due to variations of image contents. In this paper, a more generalized model is proposed to better characterize the regularity of extensive image contents, which is learned from the concatenated histograms of mean subtracted contrast normalized (MSCN) coefficients and pairwise products of MSCN coefficients of neighbouring pixels. The new feature based on MSCN shows its capability of preserving intrinsic distribution of image statistics. Consequently support vector machine regression (SVR) can map it to more accurate image quality scores. Experimental results show that the proposed approach achieves a slight gain from BRISQUE, which indicates the crafted GGD modelling step in BRISQUE is not essential for final performance. Li Song 0001, Yi Xu 0001, Gengjian Xue, Yi Zhou 0003 |
VCIP | 3 |
| 2014 | HEASK: Robust homography estimation based on appearance similarity and keypoint correspondences
Yi Xu 0001, Xiaokang Yang 0001, Truong Q. Nguyen |
Pattern Recognit. | 2 |
| 2014 | Separation of Weak Reflection from a Single Superimposed ImageabstractIt is an inherently ill-posed problem to separate a single superimposed image into a reflection image and a transmission image. In this letter, a novel algorithm is proposed based on the prior knowledge that edges of weak reflection are always smoother than most edges of observed objects. To filter out the edges of weak reflection, an MRF-EM (Markov Random Field and Expectation Maximization) framework is proposed. In the MRF model, a data energy function is established based on the edge smoothness metric GPS (Gradient Profile Sharpness), and a spatial smoothness energy function is formulated using a weighted Potts model. Moreover, the parameters in the data energy function are updated using the EM algorithm. Experimental results demonstrate that the proposed algorithm can produce superior separation results with less residuals and color distortions compared to state-of-the-art methods. Yi Xu 0001, Xiaokang Yang 0001, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 2 |
| 2013 | Quaternion-based sparse representation of color imageabstractIn this paper, we propose a quaternion-based sparse representation model for color images and its corresponding dictionary learning algorithm. Differing from traditional sparse image models, which represent RGB channels separately or process RGB channels as a concatenated real vector, the proposed model describes the color image as a quaternion vector matrix, where each color pixel is encoded as a quaternion unit and thus the inter-relationship among RGB channels is well preserved. Correspondingly, we propose a quaternion-based dictionary learning algorithm using a socalled K-QSVD method. It conducts the sparse basis selection in quaternion vector space, providing a kind of vectorial representation for the inherent color structures rather than a scalar representation via current sparse image models. The proposed sparse model is validated in the applications of color image denoising and inpainting. The experimental results demonstrate that our sparse image model avoids the hue bias phenomenon successfully and shows its potential as a powerful tool in color image analysis and processing domain. Licheng Yu, Yi Xu 0001, Hongteng Xu |
ICME | 2 |
| 2013 | Separation of weak reflection from a single superimposed image using gradient profile sharpnessabstractIt is a massively ill-posed problem to separate a superimposed image into an object image of our interested object and an interference image of reflection. Previous studies relied on redundant information introduced by multiple exposure or multi-view configurations in the separation. Later some new methods proposed tailor-made constraints to remove reflection in specific conditions for a single superimposed image. However, the separated results of these methods always have a lot of residuals or a few tone distortions. In this paper, we aim to realize a clear separation of weak reflection for a single superimposed image. Since the reflection is weak and always out of focus, the resulted interference image would have a smoother edge map than the object image. We utilize this smoothness constraint to obtain an initial separation by classifying gradients according to GPS (gradient profile sharpness) computation. Then we propose a gradient validation framework to reduce the structural correlation between the object image and the interference image. This framework can well correct the misclassified gradients obtained in the initial separation. The experimental results demonstrate that our method can generate promising separation results with little residual or color distortions. Yi Xu 0001, Xiaokang Yang 0001 |
ISCAS | 2 |
| 2013 | Measuring orderliness based on social force model in collective motionsabstractCollective motions, one of the coordinated behaviors in crowd system, widely exist in nature. Orderliness characterizes how well an individual will move smoothly and consistently with his neighbors in collective motions. It is still an open problem in computer vision. In this paper, we propose an orderliness descriptor based on correlation of interactive social force between individuals. In order to include the force correlation between two individuals in a distance, we propose a Social Force Correlation Propagation algorithm to calculate orderliness of every individual effectively and efficiently. We validate the effectiveness of the proposed orderliness descriptor on synthetic simulation. Experimental results on challenging videos of real scene crowds demonstrate that orderliness descriptor can perceive motion with low smoothness and locate disorder. Yi Xu 0001, Xiaokang Yang 0001 |
VCIP | 2 |
| 2013 | A generalized EMD with body prior for pedestrian identification
Lianyang Ma, Xiaokang Yang 0001, Yi Xu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | A Robust Homography Estimation Method Based on Keypoint Consensus and Appearance SimilarityabstractIn this paper, a robust homography estimation method is proposed to match multiview images in the uncalibrated case. This method formulates a new loss function to verify homography hypothesis, which combines models of key-point consensus and appearance similarity. In the consensus model, Lap lace distribution is exploited to better characterize the imprecision of key points. And in the appearance model, a truncated exponential function is utilized to represent the distribution of image similarity values. With these improvements, our method can be more robust in complex situations when a rather high percentage of key points are ambiguous and unreliable, and can output a more accurate homography that satisfies most pixels' geometric relationship. The experimenttal results highlight the robustness and accuracy of our method in the matching tasks of both synthetic images and real life photos. Yi Xu 0001, Xiaokang Yang 0001 |
ICME | 2 |
| 2012 | Image super-resolution based on a novel edge sharpness prior
Yi Xu 0001, Xiaokang Yang 0001, Kai Chen 0006 |
ICPR | 2 |
| 2012 | Colorization Using Quaternion Algebra with Automatic Scribble Generation
Yi Xu 0001, Lei Deng 0001, Xiaokang Yang 0001 |
MMM | 2 |
| 2012 | Robust object tracking with bidirectional corner matching and trajectory smoothness algorithmabstractThis paper proposes a novel method for robust object tracking. The method consists of three different components: a short term tracker, an object detector, and an online object model. For the short term tracker, we use an advanced Lucas Kanade tracker with bidirectional corner matching to capture object frame by frame. Meanwhile, statistical filtering and matching algorithm combined with haar-like feature random fern play as a detector to extract all possible object candidates in the current frame. Making use of trajectory information, the online object model decides the best target match among the candidates. And the model also trains the random fern feature adaptively online to better guide consecutive tracking. We demonstrate our method is robust to track an object in a long term and under large variations of view angle and lighting conditions. Moreover, our method is efficient to re-detect the object and keep tracking even after it's out of view or recover from heavy occlusion. To achieve state-of-the-art performance, it is highlighted that our method can be extended to multiple objects tracking application. Finally, comparisons with other state-of-the-art trackers are presented to show the robustness of our tracker. Guoshan Wu, Yi Xu 0001, Xiaokang Yang 0001, Ke Gu 0001 |
MMSP | 2 |
| 2011 | Human identification using body prior and generalized EMDabstractThe general configuration of body is a valuable cue for human identification, which is ignored by the existing approaches. In this paper, we present an approach for human identification by using body prior and the generalized Earth Mover's Distance (EMD). The common knowledge that a pedestrian is composed of upper body and the lower one is employed as a body prior. To achieve more robust body segmentation, we pursue their boundary by inducing a logistic probability map, which is approximated based on minimizing its KL divergence to the posterior probability of the observed person image. Furthermore, we generalize EMD by assigning different weights to regions of body, which are learned through logistic regression to boost discriminative power for human identification. The experimental results show that both body prior and the generalized EMD facilitate performance on human identification. Lianyang Ma, Xiaokang Yang 0001, Yi Xu 0001 |
ICIP | 3 |
| 2011 | Hypothesis comparison guided cross validation for unsupervised signer adaptationabstractSigner adaptation is important to sign language recognition systems in that a one-size-fits-all model set can not perform well on all kinds of signers. Supervised signer adaptation must utilize the labeled adaptation data that are collected explicitly. To skip the data collecting process in signer adaptation, we propose an unsupervised adaptation method called hypothesis comparison guided cross validation (HC CV) algorithm. The algorithm not only addresses the problem of overlap between the data set to be labeled and the data set for adaptation, but also employs an additional hypothesis comparison step to decrease the noise rate of the adaptation data set. Experimental results show that the HC CV adaptation algorithm is superior to the CV adaptation algorithm and the conventional self-teaching algorithm. Though the algorithm is proposed for signer adaptation, it can also be applied to speaker adaptation and writer adaptation straightforwardly. Yu Zhou 0015, Xiaokang Yang 0001, Weiyao Lin, Yi Xu 0001, Long Xu 0001 |
ICME | 4 |
| 2011 | Separation of superimposed images with unknown motions using sparsity priorsabstractWhen people take photos through a transparent surface, it is ubiquitous that the images are superimposed with two source images, the transmitted layer and the reflected layer. In this paper, we utilize the sparsity priors over image color distribution and image gradients, and we realize the separation of such images by exploiting the diversity (relative motions of source images) in different snapshots. Our approach can estimate the motion even when layer image is quite faint, extract layer gradients when layer intensity is unchanged, and obtain a reliable separation by considering the gradients as limitations. The effectiveness of our approach is shown in experiments on both synthetic images and real world photos. Xiaokang Yang 0001, Yi Xu 0001 |
VCIP | 3 |
| 2011 | Crowd instability analysis using velocity-field based social force modelabstractThis paper proposes a novel method to locate crowd behavior instability spatio-temporally using a velocity-field based social force model. Considering the impacts of velocity field on interaction force between individuals, we establish an improved social force model by introducing collision probability in view of velocity distribution. As compared with commonly- used social force model, which defines interaction force as a dependent variable of relative geometric (physical) position of the individuals, this improved model can provide a better prediction of interactions using the collision probability in a dynamic crowd. With spatio-temporal instability analysis, we can extract video clips with potential abnormality and as well locate region of interest where abnormality is likely to happen. The experimental results demonstrate that the proposed method can be applied to detection of abnormal events with high accuracy of instability estimation due to the velocity-field based social force model. Yi Xu 0001, Xiaokang Yang 0001 |
VCIP | 2 |
| 2010 | Bayesian error concealment with DCT pyramidabstractIn this paper, the problem of concealing missing image/video blocks is casted into a framework of Bayesian estimation. The conditional expectation of the missing block vector is taken over a pilot vector of correctly decoded pixels near the missing block. Multiple observations of the missing vector and pilot vectors obtained in a nonlocal manner are used to approximate the expectation. We design a multiscale estimation approach with DCT pyramid to improve estimation efficiency. The DC image of the missing block is recovered first, and then more details related to high frequency AC coefficients are recovered successively. The algorithm is found to be quite competitive among state-of-the-art, and more substantial improvement over existing algorithms on image with heavy loss rate/large block size is observed in our experiments. Guangtao Zhai, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001, Yi Xu 0001 |
ICASSP | 5 |
| 2010 | Fire Surveillance Method Based on Quaternionic Wavelet Features
Zhou Yu 0001, Yi Xu 0001, Xiaokang Yang 0001 |
MMM | 2 |
| 2009 | Event recognition with time varying Hidden Markov ModelabstractStandard hidden Markov model (HMM) and the more general dynamic Bayesian network (DBN) models assume stationarity of state transition distribution. However, this assumption does not hold for many real life events of interest. In this paper, we propose a new time sequence model that extends HMM to time varying scenario. The time varying property is realized in our model by explicitly allowing the change of state transition density as the time spent in a particular state passes by. Rather than keeping transition densities at different time spots independent of each other, we exploit their temporal correlation by applying a hierarchical Dirichlet prior. This leads to a more robust time varying model, especially when training data are scarce. We also employ Markov chain Monte Carlo (MCMC) sampling in learning the MAP estimate of time varying parameters, with a transition kernel incorporating linear optimization. The proposed model is applied to recognizing real video events, and is shown to outperform existing HMM-based methods. Ercan E. Kuruoglu, Xiaokang Yang 0001, Yi Xu 0001, Songyu Yu |
ICASSP | 4 |
| 2009 | Learning distance metric for regression by semidefinite programming with application to human age estimationabstractA good distance metric for the input data is crucial in many pattern recognition and machine learning applications. Past studies have demonstrated that learning a metric from labeled samples can significantly improve the performance of classification and clustering algorithms. In this paper, we investigate the problem of learning a distance metric that measures the semantic similarity of input data for regression problems. The particular application we consider is human age estimation. Our guiding principle for learning the distance metric is to preserve the local neighborhoods based on a specially designed distance as well as to maximize the distances between data that are not in the same neighborhood in the semantic space.Without any assumption about the structure and the distribution of the input data, we show that this can be done by using semidefinite programming. Furthermore, the low-level feature space can be mapped to the high-level semantic space by a linear transformation with very low computational cost. Experimental results on the publicly available FG-NET database show that 1) the learned metric correctly discovers the semantic structure of the data even when the amount of training data is small and 2) significant improvement over the traditional Euclidean metric for regression can be obtained using the learned metric. Most importantly, simple regression methods such as k nearest neighbors (kNN), combined with our learned metric, become quite competitive (and sometimes even superior) in terms of accuracy when compared with the state-of-the-art human age estimation approaches. Xiaokang Yang 0001, Yi Xu 0001, Hongyuan Zha |
ACM Multimedia | 3 |
| 2009 | CamShift guided particle filter for visual tracking
Xiaokang Yang 0001, Yi Xu 0001, Songyu Yu |
Pattern Recognit. Lett. | 3 |
| 2008 | Local Quaternionic Gabor Binary Patterns for color face recognitionabstractIn this paper, a novel color face recognition method is proposed based on local binary patterns (LBP) of quaternionic Gabor features (QGF). By introducing quaternion Gabor analysis into image representation, we make full use of the interrelationship among different color channels to enhance the performance of the face recognition system. Moreover, the QGF are used to encode the positions and attributes of the face elements. Non-parametric transformation is then imposed on these QGF using LBP method to obtain the robustness against variations of pose, illumination and facial expressions. Compared with the monochromatic face recognition systems, which nowadays dominate the marketplace and research field, this approach materializes the strong potential use of color face recognition system by establishing invariant quaternion wavelet features of color images. The experimental results on the open face database testify the validity of the proposed method under severe noise corruption and distinct variations of scale, illumination and facial expressions. Wei Lu 0021, Yi Xu 0001, Xiaokang Yang 0001, Li Song 0001 |
ICASSP | 2 |
| 2008 | Color image watermarking using local quaternion Fourier spectral analysisabstractWe propose a watermarking scheme for color images based on local quaternion Fourier spectral analysis (LQFSA).The merits of the proposed scheme include: 1) Quaternion Fourier transform is defined in a 4D vector space and thus provides a larger embedding scope for watermark than conventional monochannel transformation techniques. 2) We improve the imperceptibility of watermark with regard to human color vision properties through LQFSA. 3) We introduce invariant feature transform (IFT) and geometric correction scheme so as to enhance the robustness to extensive attacks, which is another essential factor to evaluate a watermarking scheme. 4) We adopt the nearest-neighborhood search to ensure the correctness of watermark extraction. Extensive experiments on the Stirmark platform validate the aforementioned merits. Yi Xu 0001, Li Song 0001, Xiaokang Yang 0001, Hans Burkhardt |
ICME | 2 |
| 2008 | No-reference noticeable blockiness estimation in images
Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
Signal Process. Image Commun. | 5 |
| 2008 | Efficient Image Deblocking Based on Postfiltering in Shifted WindowsabstractWe propose a simple yet effective deblocking method for JPEG compressed image through postfiltering in shifted windows (PSW) of image blocks. The MSE is compared between the original image block and the image blocks in shifted windows, so as to decide whether these altered blocks are used in the smoothing procedure. Our research indicates that there exists strong correlation between the optimal mean squared error threshold and the image quality factor Q, which is selected in the encoding end and can be computed from the quantization table embedded in the JPEG file. Also we use the standard deviation of each original block to adjust the threshold locally so as to avoid the over-smoothing of image details. With various image and bit-rate conditions, the processed image exhibits both great visual effect improvement and significant peak signal-to-noise ratio gain with fairly low computational complexity. Extensive experiments and comparison with other deblocking methods are conducted to justify the effectiveness of the proposed PSW method in both objective and subjective measures. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2008 | Efficient Deblocking With Coefficient Regularization, Shape-Adaptive Filtering, and Quantization ConstraintabstractWe propose an effective deblocking scheme with extremely low computational complexity. The algorithm involves three parts: local ac coefficient regularization (ACR) of shifted blocks in the discrete cosine transform (DCT) domain, block-wise shape adaptive filtering (BSAF) in the spatial domain, and quantization constraint (QC) in the DCT domain. The DCT domain ACR suppresses the grid noise (blockiness) in monotone areas. The spatial-domain BSAF alleviates the staircase noise along the edge, and the ringing near the edge and the corner outliers. The narrow quantization constraint set is imposed to prevent possible oversmoothing and improve PSNR performance. Extensive simulation results and comparative studies are provided to justify the effectiveness and efficiency of the proposed deblocking algorithm. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2007 | 2D Quaternion Fourier Transform: The Spectrum Properties and its Application in Color Image RegistrationabstractWe first investigate 2D quatemion Fourier transform (QFT) spectrum relationship between an image and its geometrically transformed counterpart from the aspects of gray images and color images respectively, and then propose a 2D QFT-based color image registration approach, which is able to handle large translation, rotation and scaling. Fourier transform (FT) can be utilized in gray image registration but can not process color images naturally. As the extended version of FT in multidimensional signal processing, QFT is capable of processing three color components of color images together as pure quatemions and is able to deal with color image registration. Zheng Lu 0003, Yi Xu 0001, Xiaokang Yang 0001, Li Song 0001, Leonardo Traversoni |
ICME | 2 |
| 2007 | Cooperative Stereo Matching using Quaternion Wavlets and Top-Down SegmentationabstractWe explore the principles of quaternion wavelet construction for achieving multiscale analysis of geometric image features. Then the quaternion wavelets are applied to propose a cooperative stereo matching algorithm using top-down segmentation-based disparity propagation. Without bidirectional matching to remove ambiguous outliers, uniqueness constraint is enforced on cost function by inhibiting the matches along similar sightlines. To produce smooth disparity maps with the discontinuities well-preserved, cost aggregation is performed in segmentation-based local support and high confidence matches serve as heavyweight seeds for disparity propagation in the supports. Compared with the current matching methods based on quatemion wavelets, the main merit of the proposed algorithm is that the matching results are encouraging in extensive comparison data, ranging from calibrated images to uncalibrated images, indoor images to aerial images. Yi Xu 0001, Xiaokang Yang 0001, Peifeng Zhang, Li Song 0001, Leonardo Traversoni |
ICME | 1 |
| 2007 | A Unified Framework for Removing Blocking ArtifactsabstractAl-Fahoum and Reza [1] characterized the blocking artifacts in block based DCT (BDCT) compressed image into five types: grid noise, staircase noise, ringing artifacts, corner outliers and the corruption of edges. Most of the comprehensive deblocking algorithms lack a unified framework, and different artifacts are processed with independent ad hoc schemes. In this paper, we propose a comprehensive postprocessing method for removing all the blocking-related artifacts in the framework of overcomplete wavelet expansion (OWE). We use the wavelet transform modulus maxima extension (WTMME) and angle extracted from the wavelet coefficients of 3-level OWE to represent the image. Both the WTMME and the angle image are reconstructed accordingly using inter-/ intra-band correlation to suppress the influence of the distortions. Simulation and comparative study have demonstrated the effectiveness of the proposed algorithm in terms of both subjective and objective quality of the resultant images. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
ICME | 5 |
| 2007 | Multi-Scale Gabor Phase-Based Stereo Matching using Graph CutsabstractIn this paper, we present a multi-scale Gabor phase-based stereo matching scheme. Unlike the mechanism in the existing phase-based stereo matching methods, where disparity is formulated as the ratio of phase difference between two views to the local frequency at the given position, we set up a robust data measure from multi-scale Gabor phases to greatly alleviate the negative effect of phase singularity. A cost function is then advanced based on this robust data measure. To further improve the accuracy of disparity estimation, we formulate the cost function as three coupled Markov Random Field (MRF) cost terms in frequency domain. To obtain globally optimized disparity map in wide range, graph cut is employed to perform the minimization of the cost function. Compared with the state-of-the-art stereo matching methods, experimental results demonstrate that our approach gets comparable matching performance in indoor scenes and achieves much better results in aerial scenes. Peifeng Zhang, Yi Xu 0001, Xiaokang Yang 0001, Leonardo Traversoni |
ICME | 2 |
| 2007 | Quaternion wavelet phase based stereo matching for uncalibrated images
Jun Zhou 0007, Yi Xu 0001, Xiaokang Yang 0001 |
Pattern Recognit. Lett. | 2 |
| 2006 | Modeling Blocking Visual Sensitivity ProfileabstractBlocking artifact is the most prevailing degradation caused by block-based DCT coding techniques under low bit-rate conditions. To alleviate blockings perceptually, it is desirable to measure the visibility of blocking artifacts. In this paper, we propose an efficient method of estimating the visual sensitivity of blocking artifacts in block-based DCT coding. The differences on block boundaries are measured and transformed into block discontinuity map. We consider the effects of luminance adaptation and texture masking on the blockings and integrate them using nonlinear operator to form an overall masking map. This masking map is then incorporated with the discontinuity map to generate the blocking visual sensitivity map (BVSM). This map can be used to guide perceptual quality assessment, codec parameter optimization, post-processing, etc. We demonstrate the validity of the BVSM through its application in image quality assessment Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Yi Xu 0001 |
ICME | 4 |
| 2006 | GES: a new image quality assessment metric based on energy features in Gabor transform domainabstractWe propose Gabor energy similarity (GES), a new full reference image quality assessment metric based on the measuring of Gabor energy features of images. It has been recognized that: 1) 2D Gabor filters can attain the theoretical conjoint resolution limit in spatial and frequency domain defined by Heisenberg's uncertainty principle; 2) it is widely reported that simple cells in visual cortex can be well modeled by 2D Gabor functions; and 3) the feature of Gabor energy is closely related to the model of complex cells in primary visual cortex that are very sensitive to minor changes in nature scene. Motivated by these facts we attempt to design a new image quality assessment by exploring the similarity in Gabor energy functions between the original and the distorted image. The images are firstly decomposed by a filter bank consists of 48 Gabor filters (6 directions, 4 spatial resolutions, and symmetric or anti-symmetric). Then the local energy feature is extracted by computing the modulus of the responses of symmetric and anti-symmetric kernel filters at each point. We finally propose an image quality metric based on averaged cross correlation between two images. Extensive experimental results are used to justify the effectiveness of the proposed GES. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Susu Yao, Yi Xu 0001 |
ISCAS | 5 |