VLDB 2026 Research / reviewers in the wild / expert
Shikui Wei
dblp:15/2139
· DBLP profile ↗
87ranked-venue papers
8as first author
45since 2021 · last 2026
0000-0003-3803-9763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 5 first-author · 35 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seam-Guided Unsupervised Image Stitching With Parallax-Aware Mask GenerationabstractImage stitching under large parallax remains a challenging task due to the conflict of content alignment and shape preservation. Most methods focus on precisely aligning overlapping regions via spatially varying transformations, often causing unexpected distortions in large-parallax areas. Differently, we aim to produce stitched images that are both visually natural and free of artifacts. To this end, we present a parallax-aware unsupervised warping model for seam-guided image stitching. To preserve natural content, we first design an edge-enhanced mask generation module to distinguish large-parallax regions and suppress excessive deformation around these areas. It is constrained by a comprehensive objective function that integrates masked photometric difference, nontrivial mask learning, and adaptive regularization, simultaneously ensuring mask reliability and alignment robustness. Besides, to eliminate parallax artifacts, we incorporate a seam-guided alignment strategy into our warping network, which iteratively registers local regions with the assistance of optimal seam estimation. Through adaptively finetuning the warping model, we progressively improve the stitching quality with improved seam quality. To facilitate the learning process of perceiving parallax, we construct a new image stitching dataset with larger parallax than that of UDISD, which could benefit the model’s generalization in challenging scenarios. Experiments show our solution not only removes misaligned regions but also maintains shape consistency especially in challenging parallax scenarios. Yuzhu Tao, Lang Nie, Yakun Chang, Shikui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Toward Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-Critical ScenariosabstractAutonomous driving has made significant progress in both academia and industry, including performance improvements in perception tasks and the development of end-to-end autonomous driving systems. However, the safety and robustness assessment of autonomous driving has not received sufficient attention. Current evaluations of autonomous driving are typically conducted in natural driving scenarios. However, accidents often occur in edge cases, also known as safety-critical scenarios. These safety-critical scenarios are difficult to collect, and there is currently no clear definition of what constitutes a safety-critical scenario. In this work, we explore the safety and robustness of autonomous driving in safety-critical scenarios. First, we provide a definition of safety-critical scenarios, including static traffic scenarios such as adversarial attack scenarios and natural distribution shifts, as well as dynamic traffic scenarios such as accident scenarios. Then, we develop an autonomous driving test framework to comprehensively evaluate autonomous driving systems, encompassing not only the assessment of perception modules but also system-level evaluations. Our work systematically constructs a safety verification process for autonomous driving, providing technical support for the industry to establish standardized test framework. Jingzheng Li, Xianglong Liu 0001, Shikui Wei, Yufei Ge, Bing Li 0001, Qing Guo 0005, Xianqi Yang, Yanjun Pu, Qianren Mao, Jiakai Wang |
IEEE Trans. Image Process. | 3 |
| 2026 | Joint Attribute Graph Reasoning and Aggregation for Composed Image Retrieval
Na Cai, Zhedong Zheng, Huaxin Pang, Shikui Wei |
IEEE Trans. Multim. | 6 |
| 2025 | Automatic Spectral Calibration of Hyperspectral Images: Method, Dataset and BenchmarkabstractHyperspectral images (HSI) densely sample the world in both the space and frequency domains and, therefore, are more distinctive than RGB images. Usually, HSI needs to be calibrated to minimize the impact of various illumination conditions. The traditional way to calibrate HSI utilizes a physical reference, which involves manual operations, occlusions, and/or limits camera mobility. These limitations inspire this paper to automatically calibrate HSIs using a learning-based method. Towards this goal, a large-scale HSI calibration dataset, which has 765 high-quality HSI pairs covering diversified natural scenes and illuminations, is created. The dataset is further expanded to 7650 pairs by combining with 10 different physically measured illuminations. A spectral illumination transformer (SIT) together with an illumination attention module is proposed. Extensive benchmarks demonstrate the SoTA performance of the proposed SIT. The benchmarks also indicate that low-light conditions are more challenging than normal conditions. The dataset and codes are available online: https://github.com/duranze/Automatic-spectral-calibration-of-HSI. Zhuoran Du, Shaodi You, Shikui Wei |
CVPR | 4 |
| 2025 | SDHNet: Semantic Disentanglement and Heterogeneous Feature Fusion Network for Composed Image Retrieval
Shikui Wei |
ICIC (16) | 6 |
| 2025 | Self-supervised Image Flicker Removal for Rolling-shutter CamerasabstractAlternating current (AC)-powered artificial lighting systems often introduce high-frequency flickering, which manifests as banding-pattern flickers in images captured by rolling-shutter cameras. Existing methods rely heavily on prior knowledge of lighting systems, synthetic data, or specialized hardware, limiting their practicality in real-world scenarios. To address these limitations, we propose a self-supervised framework for flicker removal using real-world data. Our method is grounded in a theoretical analysis demonstrating that flicker-induced luminance fluctuations follow a zero-mean distribution in the temporal domain, enabling the adoption of a self-supervised strategy. We design a modified U-Net architecture and introduce the Banding Exclusion (BE) loss to suppress residual artifacts along edges while preserving structural details. To support robust training and evaluation, we curate a comprehensive dataset comprising synthetic images generated via a physics-based flicker simulation and 160 real-world RAW sequences captured under flickering illumination. Experimental results demonstrate that our framework outperforms existing methods. Shuoxin Shan, Yakun Chang, Yujia Liu 0005, Renshuai Tao, Shikui Wei, Yao Zhao 0001 |
VCIP | 6 |
| 2025 | Spatio-Temporal Weighted Graph Reason Learning for Multivariate Time-Series Anomaly DetectionabstractConstructing an efficient and deployable anomaly detection system requires achieving high accuracy, low latency, and reliability. Existing methods either spend considerable time extracting rich spatio-temporal features to enhance anomaly detection performance, or blindly integrate multi-source features to boost accuracy, often neglecting the reliability of feature aggregation. The trade-off between the three objectives must be carefully considered when developing the model. To address these challenges, we introduce a novel Spatio-Temporal Weighted Graph Reasoning Learning (STWGRL) framework for multivariate time-series anomaly detection. Specifically, we propose a series-denoising receptance-weighted key value (D-RWKV) module to efficiently capture and model expressive long-term sequence information through a linear scaling mechanism. D-RWKV ensures compatibility by alleviating the memory bottleneck and enabling parallelized training. Furthermore, we design a targeted-awareness graph adaptive aggregation (TaGAA) module to learn the directed graph and adaptively enhance the signal’s intrinsic characteristics. Two graph-constraint losses are employed to strengthen the consistency and sparsity of the learned graphs. Experimental results on multiple benchmark tasks clearly demonstrate the effectiveness of the proposed framework. STWGRL achieves more accurate scores than most baselines, while containing fewer than 10K parameters. Huaqi Zhang, Huaxin Pang, Guandong Gao, Shikui Wei, Yao Zhao 0001 |
IEEE Internet Things J. | 6 |
| 2025 | Adaptive estimation of instance-dependent noise transition matrix for learning with instance-dependent label noise
Yuan Wang 0078, Huaxin Pang, Shikui Wei, Yao Zhao 0001 |
Neural Networks | 4 |
| 2025 | Rethinking erasing strategy on weakly supervised object localization
Yuming Fan, Shikui Wei, Chuangchuang Tan, Dongming Yang, Yao Zhao 0001 |
Signal Process. Image Commun. | 2 |
| 2025 | Unbiased Sample Selection and Label Improvement for Mitigating Noisy Labels in Class-Imbalanced DatasetsabstractReal-world datasets often suffer from both noisy labels and imbalanced class distribution, presenting significant challenges for the effective deployment of deep neural networks (DNNs). Existing studies typically address these challenges separately and struggle to perform effectively when they occur simultaneously. In this paper, we introduce an unbiased Sample Selection method based on the Graph Attention Network (GAT), namely GSS. GSS can effectively divide the training set into clean and noisy subsets while avoiding sample selection bias by analyzing the intrinsic relationships between the training set and a small clean validation set. For the clean subset, we propose an Adaptive Label Refinement (ALR) strategy to improve the reliability of the labels within the clean subset. ALR dynamically integrates the network’s predictions with the given labels, mitigating the adverse impacts of misidentification. For the noisy subset, we introduce a Class-Balanced Pseudo Labeling (CBPL) method. CBPL addresses the cognitive bias in model predictions caused by class imbalance by integrating class distribution information into the pseudo-label generation process, resulting in more accurate pseudo-labels. Comprehensive evaluations on both synthetic and real-world datasets highlight the effectiveness and superiority of our approach, especially in scenarios characterized by noisy labels and imbalanced class distributions. Yuan Wang 0078, Yakun Chang, Yao Zhao 0001, Shikui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Confidence-Driven Unimodal Interference Removal for Enhanced Multimodal Object DetectionabstractMulti-spectral imaging senses objects from different perspectives, exhibiting the advantages of cross-modal collaboration. However, most existing cross-modal detection algorithms focus mainly on the design of fusion mechanisms, neglecting to assess the effectiveness of individual modalities. In fact, if a certain modality fails to provide distinguishable features, it will introduce unimodal interference and weaken the feature representation of dominate modalities. To address this problem, we propose an enhanced multi-spectral object detection algorithm via Confidence-driven unimodal Interference Removal (CIRDet). Specifically, we explicitly decompose unimodal visual contents into cross-modal consensus features and conflict features. For visual contents where both modalities express confidence, we employ an equal weighting fusion strategy to exploit the synergistic effect of modal information. In cases of modal discrepancy, we introduce the global and local feature confidence fusion mechanisms to induce the network to follow the guidance of the dominant modality, thereby removing conflicted interference from inferior modality. By decoupling the features and processing separately, the proposed method prevents the loss of valid information in inferior modality and filters out unimodal interference more accurately. Extensive experiments on three widely-used multi-spectral object detection benchmarks demonstrate our method outperforms state-of-the-arts by a large margin, e.g., with a CNN backbone, CIRDet achieves 4.2 mAP@[0.5:0.95] improvement compared to Transformer-based methods. The code will be released after possible publication. Yu Wang 0006, Shikui Wei, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hierarchical Grafting Network With Structural Alignment for Ultra-High Resolution Image Segmentation
Ting Liu 0012, Shikui Wei, Yanning Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Toward Open-World Domain Adaptation via Iteratively Contrastive Learning and ClusteringabstractThe open-set domain adaptation (DA) aims to address both covariate shift and category shift between a labeled source domain and an unlabeled target domain. Nevertheless, existing open-set DA methods always ignore the demand for discovering novel classes that are not present in the source domain and simply reject them as "unknown" sets without further exploration, which motivates us to understand the unknown sets more specifically. In this article, we present a more challenging open-world DA problem that recognizes seen classes while discovering novel classes in the target domain. To address this problem, we propose a novel framework that converts this problem into a clustering task via contrastive learning to learn pairwise relationships among the instances. More specifically, our method consists of two iterative steps. The semi-supervised clustering step clusters the unlabeled target data and separates it into seen and novel classes. In the contrastive learning step, based on the cluster assignments, we design tailored contrastive losses that learn pairwise relationships to reduce domain discrepancy and discover novel classes. Our method can be optimized as an example of expectation maximization (EM). We establish several baselines by extending related work. Our method obtains the superior performance on five public datasets, benchmarking this challenging setting for future research. Jingzheng Li, Hailong Sun 0001, Jiyi Li, Shikui Wei |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Lyapunov-Stable Deep Equilibrium ModelsabstractDeep equilibrium (DEQ) models have emerged as a promising class of implicit layer models, which abandon traditional depth by solving for the fixed points of a single nonlinear layer. Despite their success, the stability of the fixed points for these models remains poorly understood. By considering DEQ models as nonlinear dynamic systems, we propose a robust DEQ model named LyaDEQ with guaranteed provable stability via Lyapunov theory. The crux of our method is ensuring the Lyapunov stability of the DEQ model's fixed points, which enables the proposed model to resist minor initial perturbations. To avoid poor adversarial defense due to Lyapunov-stable fixed points being located near each other, we orthogonalize the layers after the Lyapunov stability module to separate different fixed points. We evaluate LyaDEQ models under well-known adversarial attacks, and experimental results demonstrate significant improvement in robustness. Furthermore, we show that the LyaDEQ model can be combined with other defense methods, such as adversarial training, to achieve even better adversarial robustness. Haoyu Chu, Shikui Wei, Ting Liu 0012, Yao Zhao 0001, Yuto Miyatake |
AAAI | 2 |
| 2024 | Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningabstractThis research addresses the challenge of developing a universal deepfake detector that can effectively identify unseen deepfake images despite limited training data. Existing frequency-based paradigms have relied on frequency-level artifacts introduced during the up-sampling in GAN pipelines to detect forgeries. However, the rapid advancements in synthesis technology have led to specific artifacts for each generation model. Consequently, these detectors have exhibited a lack of proficiency in learning the frequency domain and tend to overfit to the artifacts present in the training data, leading to suboptimal performance on unseen sources. To address this issue, we introduce a novel frequency-aware approach called FreqNet, centered around frequency domain learning, specifically designed to enhance the generalizability of deepfake detectors. Our method forces the detector to continuously focus on high-frequency information, exploiting high-frequency representation of features across spatial and channel dimensions. Additionally, we incorporate a straightforward frequency domain learning module to learn source-agnostic features. It involves convolutional layers applied to both the phase spectrum and amplitude spectrum between the Fast Fourier Transform (FFT) and Inverse Fast Fourier Transform (iFFT). Extensive experimentation involving 17 GANs demonstrates the effectiveness of our proposed method, showcasing state-of-the-art performance (+9.8\%) while requiring fewer parameters. The code is available at https://github.com/chuangchuangtan/FreqNet-DeepfakeDetection. Chuangchuang Tan, Yao Zhao 0001, Shikui Wei, Guanghua Gu, Ping Liu 0004, Yunchao Wei |
AAAI | 3 |
| 2024 | Learning Invariant Inter-pixel Correlations for Superpixel GenerationabstractDeep superpixel algorithms have made remarkable strides by substituting hand-crafted features with learnable ones. Nevertheless, we observe that existing deep superpixel methods, serving as mid-level representation operations, remain sensitive to the statistical properties (e.g., color distribution, high-level semantics) embedded within the training dataset. Consequently, learnable features exhibit constrained discriminative capability, resulting in unsatisfactory pixel grouping performance, particularly in untrainable application scenarios. To address this issue, we propose the Content Disentangle Superpixel (CDS) algorithm to selectively separate the invariant inter-pixel correlations and statistical properties, i.e., style noise. Specifically, We first construct auxiliary modalities that are homologous to the original RGB image but have substantial stylistic variations. Then, driven by mutual information, we propose the local-grid correlation alignment across modalities to reduce the distribution discrepancy of adaptively selected features and learn invariant inter-pixel correlations. Afterwards, we perform global-style mutual information minimization to enforce the separation of invariant content and train data styles. The experimental results on four benchmark datasets demonstrate the superiority of our approach to existing state-of-the-art methods, regarding boundary adherence, generalization, and efficiency. Code and pre-trained model are available at https://github.com/rookiie/CDSpixel. Shikui Wei, Lixin Liao |
AAAI | 2 |
| 2024 | Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake DetectionabstractRecently, the proliferation of highly realistic synthetic images, facilitated through a variety of GANs and Diffusions, has significantly heightened the susceptibility to misuse. While the primary focus of deepfake detection has traditionally centered on the design of detection algorithms, an investigative inquiry into the generator architectures has remained conspicuously absent in recent years. This paper contributes to this lacuna by rethinking the architectures of CNN-based generator, thereby establishing a generalized representation of synthetic artifacts. Our findings illuminate that the up-sampling operator can, beyond frequency-based artifacts, produce generalized forgery artifacts. In particular, the local interdependence among image pixels caused by upsampling operators is significantly demon-strated in synthetic images generated by GAN or diffusion. Building upon this observation, we introduce the concept of Neighboring Pixel Relationships(NPR) as a means to capture and characterize the generalized structural artifacts stemming from up-sampling operations. A comprehensive analysis is conducted on an open-world dataset, comprising samples generated by 28 distinct generative models. This analysis culminates in the establishment of a novel state-of-the-art performance, showcasing a remarkable 12.8% im-provement over existing methods. The code is available at https://github.com/chuangchuangtan/NPR-DeepfakeDetection. Chuangchuang Tan, Huan Liu 0001, Yao Zhao 0001, Shikui Wei, Guanghua Gu, Ping Liu 0004, Yunchao Wei |
CVPR | 4 |
| 2024 | Multi-granular Semantic Mining for Composed Image RetrievalabstractComposed Image Retrieval(CIR) aims to model users’ query intention with multiple modalities and retrieve the desired images from a large image corpus. The biggest challenge is how to effectively integrate the semantic information between two different modalities. A popular solution is to design attention-based modules to extract the query embedding in a coarse manner, which leads to certain confusion about search intention. To address this problem, we propose a new method for query integration, which is composed of two key modules, i.e., Multi-granular Subspace Fusion (MSF) and Residual Regression (RR) constraint. Specifically, MSF focuses on mining cross-modal semantic dependency between reference image regions and modification text pieces in multi-granular subspaces, which can construct an implicit, holistic semantic relationship in a fine manner. And RR constraint pushes the visual-text semantic alignment under specific supervision. Extensive experiments on three prevalent datasets demonstrate the state-of-the-art performance of our method. Shikui Wei, Gangjian Zhang, Yao Zhao 0001 |
ICME | 2 |
| 2024 | Structure-Preserving Physics-Informed Neural Networks with Energy or Lyapunov Structure
Haoyu Chu, Yuto Miyatake, Wenjun Cui, Shikui Wei, Daisuke Furihata |
IJCAI | 4 |
| 2024 | Unsupervised node representation learning of pure graph via symmetric cumulative sampling strategy
Huaxin Pang, Shikui Wei, Tianzhi Jia, Yao Zhao 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Improving neural ordinary differential equations via knowledge distillationabstractAbstract Neural ordinary differential equations (ODEs) (Neural ODEs) construct the continuous dynamics of hidden units using ODEs specified by a neural network, demonstrating promising results on many tasks. However, Neural ODEs still do not perform well on image recognition tasks. The possible reason is that the one‐hot encoding vector commonly used in Neural ODEs can not provide enough supervised information. A new training based on knowledge distillation is proposed to construct more powerful and robust Neural ODEs fitting image recognition tasks. Specially, the training of Neural ODEs is modelled into a teacher‐student learning process, in which ResNets are proposed as the teacher model to provide richer supervised information. The experimental results show that the new training manner can improve the classification accuracy of Neural ODEs by 5.17%, 24.75%, 7.20%, and 8.99%, on Street View House Numbers, CIFAR10, CIFAR100, and Food‐101, respectively. In addition, the effect of knowledge distillation is also evaluated in Neural ODEs on robustness against adversarial examples. The authors discover that incorporating knowledge distillation, coupled with the increase of the time horizon, can significantly enhance the robustness of Neural ODEs. The performance improvement is analysed from the perspective of the underlying dynamical system. Haoyu Chu, Shikui Wei, Qiming Lu, Yao Zhao 0001 |
IET Comput. Vis. | 2 |
| 2024 | Training Superpixel Network Only OnceabstractAlthough existing deep superpixel methods have significantly improved the performance, they are generally dataset-specific and need re-training for unseen data. This issue results in the failure of deep superpixel methods when facing untrainable application scenarios. In this letter, we propose a novel deep superpixel algorithm to enhance the domain adaptability and generalization performance of deep superpixel methods. Specifically, we propose the domain-free embedding to replace traditional LAB color coding, reducing the network's reliance on training set statistical properties by narrowing the gap between training data and the real-world domain shift. Simultaneously, to prevent performance loss due to color information loss, we introduce the reconstructed contour constraint to directly enhance the boundary-fitting capability of superpixels. Experimental results demonstrate that our superpixel model achieve optimal cross-domain adaptation capability with just one training session, even facing extreme domain shift. Shikui Wei, Yao Zhao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Graph Representation Learning Based on Specific Subgraphs for Biomedical Interaction PredictionabstractDiscovering the novel associations of biomedical entities is of great significance and can facilitate not only the identification of network biomarkers of disease but also the search for putative drug targets.Graph representation learning (GRL) has incredible potential to efficiently predict the interactions from biomedical networks by modeling the robust representation for each node.> However, the current GRL-based methods learn the representation of nodes by aggregating the features of their neighbors with equal weights. Furthermore, they also fail to identify which features of higher-order neighbors are integrated into the representation of the central node. In this work, we propose a novel graph representation learning framework: a multi-order graph neural network based on reconstructed specific subgraphs (MGRS) for biomedical interaction prediction. In the MGRS, we apply the multi-order graph aggregation module (MOGA) to learn the wide-view representation by integrating the multi-hop neighbor features. Besides, we propose a subgraph selection module (SGSM) to reconstruct the specific subgraph with adaptive edge weights for each node. SGSM can clearly explore the dependency of the node representation on the neighbor features and learn the subgraph-based representation based on the reconstructed weighted subgraphs. Extensive experimental results on four public biomedical networks demonstrate that the MGRS performs better and is more robust than the latest baselines. Huaxin Pang, Shikui Wei, Zhuoran Du, Shengxing Cai, Yao Zhao 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | ESNet: An Efficient Framework for Superpixel SegmentationabstractSuperpixel segmentation divides an original image into mid-level regions to reduce the number of computational primitives for subsequent tasks. The two-stage approaches work better but have high computational complexity among the existing deep superpixel algorithms. In contrast, the FCN style approaches cannot extract specific image features for the superpixel task. To combine the advantages of both types of methods, we propose a carefully designed framework termed Efficient Superpixel Network (ESNet) to explicitly enhance the capability of the network to describe clustering-friendly features and simultaneously preserve the simple network structure. Concretely, two points are concerned with ESNet. First, meaningful features need to be constructed for effective superpixel clustering; hence we propose the Pyramid-gradient Superpixel Generator(PSG) to decouple the ESNet into two joint parts, i.e., the feature extractor and the superpixel generator. Second, the superpixel generator is designed in an efficient manner, which performs multi-scale sampling of input images, and can work independently by replacing the introduced feature extractor with two initial convolutional layers. Extensive experiments show that our framework achieves state-of-the-art performances on multi-datasets and is 5.3× smaller on inference than the best existing one-stage FCN-based methods. Shikui Wei, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Model-Free Rectification via Cascaded Distortion Model and Enhanced Backward Flow NetworkabstractModel-free rectification methods are limited by poor rectification quality and low generalization. This paper introduces a novel framework for enhancing model-free distortion rectification by addressing the limitations of existing methods. Our proposed method incorporates a Cascaded Distortion Model (CDM) inspired by fisheye lenses, which combines multiple reversible distortion models to create a versatile and comprehensive framework. By utilizing backward warping instead of forward warping, our approach overcomes the limitations of non-integer pixel positions and grid artifacts. Furthermore, our data synthesis method facilitates the fusion of different distortion models, bridging the distribution gap and improving generalization. To improve flow prediction accuracy, we introduce a two-stream network that incorporates both forward and backward flow branches. This approach enhances the prediction of backward flow and improves overall distortion rectification performance. We evaluate our method on large-scale synthetic datasets and real distorted images, and the results demonstrate its superior performance in both qualitative and quantitative experiments. Jie Zhao 0035, Shikui Wei, Yakun Chang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Toward Accurate Human Parsing Through Edge Guided DiffusionabstractExisting human parsing frameworks commonly employ joint learning of semantic edge detection and human parsing to facilitate the localization around boundary regions. Nevertheless, the parsing prediction within the interior of the part contour may still exhibit inconsistencies due to the inherent ambiguity of fine-grained semantics. In contrast, binary edge detection does not suffer from such fine-grained semantic ambiguity, leading to a typical failure case where misclassification occurs inner the part contour while the semantic edge is accurately detected. To address these challenges, we develop a novel diffusion scheme that incorporates guidance from the detected semantic edge to mitigate this problem by propagating corrected classified semantics into the misclassified regions. Building upon this diffusion scheme, we present an Edge Guided Diffusion Network (EGDNet) for human parsing, which can progressively refine the parsing predictions to enhance the accuracy and coherence of human parsing results. Moreover, we design a horizontal-vertical aggregation to exploit inherent correlations among body parts along both the horizontal and vertical axes, which aims at enhancing the initial parsing results. Extensive experimental evaluations on various challenging datasets demonstrate the effectiveness of the proposed EGDNet. Remarkably, our EGDNet shows impressive performances on six benchmark datasets, including four human body parsing datasets (LIP, CIHP, ATR, and PASCAL-Person-Part), and two human face parsing datasets (CelebAMask-HQ and LaPa). Ting Liu 0012, Hongkun Zhu, Yunchao Wei, Shikui Wei, Yao Zhao 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Multimodal Composition Example Mining for Composed Query Image RetrievalabstractComposed query image retrieval task aims to retrieve the target image in the database by a query that composes two different modalities: a reference image and a sentence declaring that some details of the reference image need to be modified and replaced by new elements. Tackling this task needs to learn a multimodal embedding space, which can make semantically similar targets and queries close but dissimilar targets and queries as far away as possible. Most of the existing methods start from the perspective of model structure and design some clever interactive modules to promote the better fusion and embedding of different modalities. However, their learning objectives use conventional query-level examples as negatives while neglecting the composed query's multimodal characteristics, leading to the inadequate utilization of the training data and suboptimal construction of metric space. To this end, in this paper, we propose to improve the learning objective by constructing and mining hard negative examples from the perspective of multimodal fusion. Specifically, we compose the reference image and its logically unpaired sentences rather than paired ones to create component-level negative examples to better use data and enhance the optimization of metric space. In addition, we further propose a new sentence augmentation method to generate more indistinguishable multimodal negative examples from the element level and help the model learn a better metric space. Massive comparison experiments on four real-world datasets confirm the effectiveness of the proposed method. Gangjian Zhang, Shikun Li, Shikui Wei, Shiming Ge, Na Cai, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Exploring the Applicability of Spectral Recovery in Semantic Segmentation of RGB ImagesabstractCompared with RGB images, hyperspectral images (HSIs) offer a distinct advantage in that they can record continuous spectral bands of light reflectance in each pixel, reflecting the physical and chemical characteristics of materials. This capability enables differentiation between objects that may have similar textures but different spectral characteristics. It is desirable to recover spectral information from RGB images to improve semantic segmentation accuracy. Additionally, semantic information can serve as a guide for spectral information recovery, thereby ensuring the quality of the recovered spectral information. The two tasks are mutually beneficial in this regard. In light of these considerations, we propose a multi-task framework that exploits the complementary relationship between spectral recovery and semantic segmentation tasks, comprising a complementary spectral-semantic attentive fusion model (CSSF) that enables the two tasks to mutually facilitate each other by fusing information from both branches. Specifically, the proposed CSSF incorporates a window-based spectral-semantic attentive fusion (WSSAF) module to incorporate recovered spectral information into the segmentation process effectively, and a pixel-shuffle-based fusion (PSF) module to provide semantic guidance for spectral recovery. To evaluate the effectiveness of our approach, we built the first flower hyperspectral image dataset (FHRS) with corresponding segmentation annotations and RGB images. By doing so, we have made the first attempt to explore the complementary relationship between semantic segmentation and spectral recovery. Experimental results on both the FHRS dataset and the publicly available LIB-HSI dataset demonstrate that our proposed method has the ability to enhance both tasks by utilizing their complementary relationship, indicating the generalization ability of our method. Zhuoran Du, Shikui Wei, Ting Liu 0012, Shunli Zhang 0005, Shiyin Zhang, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Each Performs Its Functions: Task Decomposition and Feature Assignment for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) aims to segment the object instances that produce sound at the time of the video frames. Existing related solutions focus on designing cross-modal interaction mechanisms, which try to learn audio-visual correlations and simultaneously segment objects. Despite effectiveness, the close-coupling network structures become increasingly complex and hard to analyze. To address these problems, we propose a simple but effective method, ‘EachPerformsItsFunctions (PIF),’ which focuses on task decomposition and feature assignment. Inspired by human sensory experiences, PIF decouples AVS into two subtasks, correlation learning, and segmentation refinement, via two branches. Correlation learning aims to learn the correspondence between sound and visible individuals and provide the positional prior. Segmentation refinement focuses on fine segmentation. Then we assign different level features to perform the appropriate duties, i.e., using deep features for cross-modal interaction due to their semantic advantages; using rich textures of shallow features to improve segmentation results. Moreover, we propose the recurrent collaboration block to enhance interbranch communication. Experimental results on AVSBench show that our method outperforms related state-of-the-art methods by a large margin (e.g., +6.0% mIoU and +7.6% F-score on the Multi-Source subset). In addition, by purposely boosting subtasks' performance, our approach can serve as a strong baseline for audio-visual segmentation. Shikui Wei, Lixin Liao, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Enhance Composed Image Retrieval via Multi-Level Collaborative Localization and Semantic Activeness PerceptionabstractComposed image retrieval (CIR) is an emerging and challenging research task that combines two modalities, a reference image, and a modification text, into one query to retrieve the target image. In online shopping scenarios, the user would use the modification text as feedback to describe the difference between the reference and the desired image. In order to handle the task, there must be two main problems needed to be addressed. One is the localization problem: how to precisely find those spatial areas of the image mentioned by the text. The other is the modification problem: how to effectively modify the image semantics based on the text. However, existing methods merely fuse information coarsely from the two-modality, while the accurate spatial and semantic correspondence between these two heterogeneous features tends to be neglected. Therefore, image details cannot be precisely located and modified. To this end, we consider integrating information from the two modalities more accurately from spatial and semantic aspects. Thus, we propose an end-to-end framework for the CIR task, which contains three key components, i.e., Multi-level Collaborative Localization module (MCL), Differential Semantics Discrimination module (DSD), and Image Difference Enhancement constraints (IDE). Specifically, to solve the localization problem, MCL precisely locates the text to the image areas by collaboratively using text positioning information on multiple image layers. For the modification problem, DSD builds a distribution to evaluate the modification possibility of each image semantic dimension, and IDE effectively learns the modification patterns of text against image embedding based on the distribution. Extensive experiments on three datasets show that the proposed method achieves outstanding performance against the SOTA methods. Gangjian Zhang, Shikui Wei, Huaxin Pang, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images DetectionabstractRecently, there has been a significant advancement in image generation technology, known as GAN. It can easily generate realistic fake images, leading to an increased risk of abuse. However, most image detectors suffer from sharp performance drops in unseen domains. The key of fake image detection is to develop a generalized representation to describe the artifacts produced by generation models. In this work, we introduce a novel detection framework, named Learning on Gradients (LGrad), designed for identifying GAN-generated images, with the aim of constructing a generalized detector with cross-model and cross-data. Specifically, a pretrained CNN model is employed as a transformation model to convert images into gradients. Subsequently, we leverage these gradients to present the generalized artifacts, which are fed into the classifier to ascertain the authenticity of the images. In our framework, we turn the data-dependent problem into a transformation-model-dependent problem. To the best of our knowledge, this is the first study to utilize gradients as the representation of artifacts in GAN-generated images. Extensive experiments demonstrate the effectiveness and robustness of gradients as generalized artifact representations. Our detector achieves a new state-of-the-art performance with a remarkable gain of 11.4%. The code is released at https://github.com/chuangchuangtan/LGrad. Chuangchuang Tan, Yao Zhao 0001, Shikui Wei, Guanghua Gu, Yunchao Wei |
CVPR | 3 |
| 2023 | Rethinking Parking Slot Detection with Rotated Bounding BoxabstractParking slot detection is an essential yet challenging task in the field of self-driving perception. During parking, vehicles often block part of the parking slots which makes the corners occluded. In addition, due to the impact of the external environment, the corners of the parking slot may be blurred. Existing parking slot detection algorithms based on parking slot markings are sensitive to the corners of the parking slots, which makes it difficult to cope with the above scenario. To address this problem, we propose a parking slot entrance line detection algorithm called RPSED, which is the first to apply rotating object detection to the parking slot entrance line. RPSED takes a different route from traditional corner detection methods by focusing on the entrance lines of parking slots to grasp the intricate geometric details inherent to parking slots, which solves the problem that existing parking slot detection algorithms cannot detect parking slots with blurred corners. To further improve the precision and recall of the model and make the model more generalizable, we propose a model ensemble strategy to match and select the results of multiple models. Moreover, we propose two manually optimized parking slot dataset named RPS2.0 and RPSV, which adds more annotations with obstructed corners or obscured configurations to the datasets ps2.0 and psv, making the model evaluation more reasonable and realistic. Experimental results on the RPS2.0 and RPSV benchmarks demonstrate the superiority of our approach compared to existing state-of-the-art methods. Shikui Wei, Shiyin Zhang, Weiyan Xu, Yao Zhao 0001 |
MMAsia | 2 |
| 2023 | ℒ풪2net: Global-Local Semantics Coupled Network for scene-specific video foreground extraction with less supervision
Shikui Wei, Yao Zhao 0001, Baoqing Guo, Zujun Yu |
Pattern Anal. Appl. | 2 |
| 2023 | Interactive Object Segmentation With Inside-Outside GuidanceabstractThis article explores how to harvest precise object segmentation masks while minimizing the human interaction cost. To achieve this, we propose a simple yet effective interaction scheme, named Inside-Outside Guidance (IOG). Concretely, we leverage an inside point that is clicked near the object center and two outside points at the symmetrical corner locations (top-left and bottom-right or top-right and bottom-left) of an almost-tight bounding box that encloses the target object. The interaction results in a total of one foreground click and four background clicks for segmentation. The advantages of our IOG are four-fold: 1) the two outside points can help remove distractions from other objects or background; 2) the inside point can help eliminate the unrelated regions inside the bounding box; 3) the inside and outside points are easily identified, reducing the confusion raised by the state-of-the-art DEXTR Maninis et al. 2018, in labeling some extreme samples; 4) it naturally supports additional click annotations for further correction. Despite its simplicity, our IOG not only achieves state-of-the-art performance on several popular benchmarks such as GrabCut Rother et al. 2004, PASCAL Everingham et al. 2010 and MS COCO Russakovsky et al. 2015, but also demonstrates strong generalization capability across different domains such as street scenes (Cityscapes Cordts et al. 2016), aerial imagery (Rooftop Sun et al. 2014 and Agriculture-Vision Chiu et al. 2020) and medical images (ssTEM Gerhard et al. 2013). Code is available at https://github.com/shiyinzhang/Inside-Outside-Guidancehttps://github.com/shiyinzhang/Inside-Outside-Guidance. Shiyin Zhang, Shikui Wei, Jun Hao Liew, Kunyang Han, Yao Zhao 0001, Yunchao Wei |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Heterogeneous Feature Alignment and Fusion in Cross-Modal Augmented Space for Composed Image RetrievalabstractComposed image retrieval (CIR) aims at fusing a reference image and text feedback to search for the desired images. Compared to general image retrieval, it can model the users' search intent more comprehensively and search the target images more accurately, which has significant impacts in various real-world applications, such as E-commerce and Internet search. However, because of the existing heterogeneous semantic gap, the synthetic understanding and fusion of both image and text are difficult to implement. In this work, to tackle this difficult problem, we propose an end-to-end framework MCR, which uses text and images as retrieval queries. The framework mainly includes four pivotal modules. Specifically, we introduce the Relative Caption-aware Consistency (RCC) constraint to align text pieces and images in the database, which can effectually bridge the heterogeneous gap. The Multi-modal Complementary Fusion (MCF) and Cross-modal Guided Pooling (CGP) are constructed to mine multiple interactions between image local features and text word features and learn the complementary representation of the composed query. Furthermore, we develop a plug-and-play Weak-text Semantic Augment (WSA) module for datasets with short or incomplete query texts, which can supplement the weak-text features and is conducive to modeling an augmented semantic space. Extensive experiments demonstrate the practical superior performance over the existing state-of-the-art empirical algorithms on several benchmarks. Huaxin Pang, Shikui Wei, Gangjian Zhang, Shiyin Zhang, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Dual-Gradients Localization Framework With Skip-Layer Connections for Weakly Supervised Object LocalizationabstractThis article focuses on generating object locations in a given image while only using image-level annotations. Towards this end, we present a simple and effective training-free framework, named Dual-Gradients Localization (DGL) framework. The key idea of the proposed DGL framework is to leverage two kinds of gradients to achieve precise localization on any convolutional layer of a classification model during the testing stage. Concretely, the DGL framework is developed based on two branches: 1) Pixel-level Class Selection, leveraging gradients of the target class to identify the correlation ratio of pixels to the target class within any convolutional feature maps, and 2) Class-aware Enhanced Maps, utilizing linear relationship in gradients of the classification loss function to mine entire target object regions. To further polish the details of objects, we apply the skip-layer connections to the classification model, which concatenates the high- and low-level layers to achieve classification. In such a case, DGL with Skip-layer Connections (DGL-SC) can capture more edge information on the high-level layer. In addition, we propose a Localization Maps Selection method to evaluate the quality of the localization map and provide a way for automatically selecting localization maps produced on different layers. Extensive experiments on public ILSVRC and CUB-200-2011 datasets show the effectiveness of the proposed DGL framework. Especially, our DGL-SC obtains a new state-of-the-art gt-known localization error of 27.35% on the ILSVRC benchmark. Chuangchuang Tan, Guanghua Gu, Shikui Wei, Yao Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Weakly Supervised Object Localization with Noisy-Label Learning
Yuming Fan, Shikui Wei, Chuangchuang Tan, Yao Zhao 0001 |
PRCV (4) | 2 |
| 2022 | Multi-Source Aggregation Transformer for Concealed Object Detection in Millimeter-Wave ImagesabstractThe active millimeter wave scanner has been widely used for detecting objects concealed underneath a person’s clothing in the field of security inspection and anti-terrorism. However, the active millimeter wave (AMMW) images always suffer from low signal-noise ratio, motion blur, and small size objects, making it challenging to detect concealed objects efficiently and accurately. The scanner usually captures a sequence of images in different views around a human body at once, while the existing algorithms only utilize the single image without considering the relationships among images. In this paper, we design a multi-source aggregation transformer (MATR) with two different attention mechanisms to model spatial correlations within an image and contextual interactions across images. Specifically, a self-attention module is introduced to encode local relationships between the region proposals in each image, while a cross-attention mechanism is built to focus on modeling the cross-correlations between different images. Besides, to handle the problem of small objects in size and suppress the noise in AMMW images, we present a selective context module (SCM). It designs a dynamic selection mechanism to enhance the high-resolution feature with spatial details and make it more distinguishable from the noisy background. Experiments on two AMMW image datasets demonstrate that the proposed methods lead to a remarkable improvement compared to previous state-of-the-art and will benefit the concealed object detection in practice. Ting Liu 0012, Shiyin Zhang, Yao Zhao 0001, Shikui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Composed Image Retrieval via Explicit Erasure and Replenishment With Semantic AlignmentabstractComposed image retrieval aims at retrieving the desired images, given a reference image and a text piece. To handle this task, two important subprocesses should be modeled reasonably. One is to erase irrelated details of the reference image against the text piece, and the other is to replenish the desired details in the image against the text piece. Nowadays, the existing methods neglect to distinguish between the two subprocesses and implicitly put them together to solve the composed image retrieval task. To explicitly and orderly model the two subprocesses of the task, we propose a novel composed image retrieval method which contains three key components, i.e., Multi-semantic Dynamic Suppression module (MDS), Text-semantic Complementary Selection module (TCS), and Semantic Space Alignment constraints (SSA). Concretely, MDS is to erase irrelated details of the reference image by suppressing its semantic features. TCS aims to select and enhance the semantic features of the text piece and then replenish them to the reference image. In the end, to facilitate the erasure and replenishment subprocesses, SSA aligns the semantics of the two modality features in the final space. Extensive experiments on three benchmark datasets (Shoes, FashionIQ, and Fashion200K) show the superior performance of our approach against state-of-the-art methods. Gangjian Zhang, Shikui Wei, Huaxin Pang, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | GradingNet: Towards Providing Reliable Supervisions for Weakly Supervised Object Detection by Grading the Box CandidatesabstractWeakly-Supervised Object Detection (WSOD) aims at training a model with limited and coarse annotations for precisely locating the regions of objects. Existing works solve the WSOD problem by using a two-stage framework, i.e., generating candidate bounding boxes with weak supervision information and then refining them by directly employing supervised object detection models. However, most of such works mainly focus on the performance boosting of the first stage, while ignoring the better usage of generated candidate bounding boxes. To address this issue, we propose a new two-stage framework for WSOD, named GradingNet, which can make good use of the generated candidate bounding boxes. Specifically, the proposed GradingNet consists of two modules: Boxes Grading Module (BGM) and Informative Boosting Module (IBM). BGM generates proposals of the bounding boxes by using standard one-stage weakly-supervised methods, then utilizes Inclusion Principle to pick out highly-reliable boxes and evaluate the grade of each box. With the above boxes and their grade information, an effective anchor generator and a grade-aware loss are carefully designed to train the IBM. Taking the advantages of the grade information, our GradingNet achieves state-of-the-art performance on COCO, VOC 2007 and VOC 2012 benchmarks. Qifei Jia, Shikui Wei, Yao Zhao 0001 |
AAAI | 2 |
| 2021 | Heterogeneous Feature Fusion and Cross-modal Alignment for Composed Image RetrievalabstractComposed image retrieval aims at performing image retrieval task by giving a reference image and a complementary text piece. Since composing both image and text information can accurately model the users' search intent, composed image retrieval can perform target-specific image retrieval task and be potentially applied to many scenarios such as interactive product search. However, two key challenging issues must be addressed in composed image retrieval occasion. One of them is how to fuse heterogeneous image and text piece in the query into a complementary feature space. The other is how to bridge the heterogeneous gap between text pieces in the query and images in the database. To address the issues, we propose an end-to-end framework for composed image retrieval, which consists of three key components including Multi-modal Complementary Fusion (MCF), Cross-modal Guided Pooling (CGP), and Relative Caption-aware Consistency (RCC). By incorporating MCF and CGP modules, we can fully integrate the complementary information of image and text piece in the query through multiple deep interactions and aggregate obtained local features into an embedding vector. To bridge the heterogeneous gap, we introduce the RCC constraint to align text pieces in the query and images in the database. Extensive experiments on four public benchmark datasets show that the proposed composed image retrieval framework achieves outstanding performance against the state-of-the-art methods. Gangjian Zhang, Shikui Wei, Huaxin Pang, Yao Zhao 0001 |
ACM Multimedia | 2 |
| 2021 | Towards Transferable 3D Adversarial AttackabstractCurrently, most of the adversarial attacks focused on perturbation adding on 2D images. In this way, however, the adversarial attacks cannot easily be involved in a real-world AI system, since it is impossible for the AI system to open an interface to attackers. Therefore, it is more practical to add perturbation on real-world 3D objects’ surface, i.e., 3D adversarial attacks. The key challenges for 3D adversarial attacks are how to effectively deal with viewpoint changing and keep strong transferability across different state-of-the-art networks. In this paper, we mainly focus on improving the robustness and transferability of 3D adversarial examples generated by perturbing the surface textures of 3D objects. Towards this end, we propose an effective method, named Momentum Gradient-Filter Sign Method (M-GFSM), to generate 3D adversarial examples. Specially, the momentum is introduced into the procedure of 3D adversarial examples generation, which results in multiview robustness of 3D adversarial examples and high efficiency of attacking by updating the perturbation and stabilizing the update directions. In addition, filter operation is involved to improve the transferability of 3D adversarial examples by filtering gradient images selectively and completing the gradients of neglected pixels caused by downsampling in the rendering stage. Experimental results show the effectiveness and good transferability of the proposed method. Besides, we show that the 3D adversarial examples generated by our method still be robust under different illuminations. Qiming Lu, Shikui Wei, Haoyu Chu, Yao Zhao 0001 |
MMAsia | 2 |
| 2021 | DQN-based gradual fisheye image rectification
Jie Zhao 0035, Shikui Wei, Lixin Liao, Yao Zhao 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | Spatial-Aware Texture Transformer for High-Fidelity Garment TransferabstractGarment transfer aims to transfer the desired garment from a model image with the desired clothing to a target person, which has attracted a great deal of attention due to its wider potential applications. However, considering the model and target persons are often given at different views, body shapes and poses, realistic garment transfer is facing the following challenges that have not been well addressed: 1) deforming the garment; 2) inferring unobserved appearance; 3) preserving fine texture details. To tackle these challenges, we propose a novel SPatial-Aware Texture Transformer (SPATT) model. Different from existing models, SPATT establishes correspondence and infers unobserved clothing appearance by leveraging the spatial prior information of a UV-space. Specifically, the source image is transformed into a partial UV texture map guided by the extracted dense pose. To better infer the unseen appearance utilizing seen region, we first propose a novel coordinate-prior map that defines the spatial relationship between the coordinates in the UV texture map, and design an algorithm to compute it. Based on the proposed coordinate-prior map, we present a novel spatial-aware texture generation network to complete the partial UV texture. In the second stage, we first transform the completed UV texture to fit the target person. To polish the details and improve realism, we introduce a refinement generative network conditioned on the warped image and source input. Compared with existing frameworks as shown experimentally, the proposed framework can generate more realistic images with better-preserved texture details. Furthermore, difficult cases where two persons have large pose and view differences can also be well handled by SPATT. Ting Liu 0012, Xuecheng Nie, Yunchao Wei, Shikui Wei, Yao Zhao 0001, Jiashi Feng |
IEEE Trans. Image Process. | 5 |
| 2021 | Blind Image Clustering for Camera Source Identification via Row-Sparsity OptimizationabstractGiven a set of images with the number of cameras providing those images unknown, how to blindly identify the sources of the images has been a critical problem in digital forensics. Although state-of-the-art methods have achieved impressive results, they have failed at suppressing outliers. When they deal with a noisy dataset, the performance is significantly degraded. To address this issue, we propose an optimization approach with sparsity constraints to simultaneously handle the how-many subproblem (i.e., the number of cameras) and the which-from-which subproblem (i.e., the image–camera relationship). In our approach, we first formulate the blind camera source clustering as a row-sparsity optimization problem, in which the representation errors are minimized and the outliers caused by noisy features are suppressed. Then, a new two-stage refinement method based on inter- and the intra-class differences is proposed to achieve a more accurate estimation of the number of cameras. Because strong sparsity constraints have been adopted and the interactive relationship among data points can be fully explored to distinguish the images originated from different cameras, the proposed method can effectively handle outliers. Extensive experiments on the popular Dresden dataset show that the proposed method outperforms existing methods in both identification accuracy and efficiency. Xiang Jiang 0005, Shikui Wei, Ting Liu 0012, Ruizhen Zhao, Yao Zhao 0001, Heng Huang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Interactive Object Segmentation With Inside-Outside GuidanceabstractThis paper explores how to harvest precise object segmentation masks while minimizing the human interaction cost. To achieve this, we propose an Inside-Outside Guidance (IOG) approach in this work. Concretely, we leverage an inside point that is clicked near the object center and two outside points at the symmetrical corner locations (top-left and bottom-right or top-right and bottom-left) of a tight bounding box that encloses the target object. This results in a total of one foreground click and four background clicks for segmentation. The advantages of our IOG is four-fold: 1) the two outside points can help to remove distractions from other objects or background; 2) the inside point can help to eliminate the unrelated regions inside the bounding box; 3) the inside and outside points are easily identified, reducing the confusion raised by the state-of-the-art DEXTR in labeling some extreme samples; 4) our approach naturally supports additional clicks annotations for further correction. Despite its simplicity, our IOG not only achieves state-of-the-art performance on several popular benchmarks, but also demonstrates strong generalization capability across different domains such as street scenes, aerial imagery and medical images, without fine-tuning. In addition, we also propose a simple two-stage solution that enables our IOG to produce high quality instance segmentation masks from existing datasets with off-the-shelf bounding boxes such as ImageNet and Open Images, demonstrating the superiority of our IOG as an annotation tool. Shiyin Zhang, Jun Hao Liew, Yunchao Wei, Shikui Wei, Yao Zhao 0001 |
CVPR | 4 |
| 2020 | Dual-Gradients Localization Framework for Weakly Supervised Object LocalizationabstractWeakly Supervised Object Localization (WSOL) aims to learn object locations in a given image while only using image-level annotations. For highlighting the whole object regions instead of the discriminative parts, previous works often attempt to train classification model for both classification and localization tasks. However, it is hard to achieve a good tradeoff between the two tasks, if only classification labels are employed for training on a single classification model. In addition, all of recent works just perform localization based on the last convolutional layer of classification model, ignoring the localization ability of other layers. In this work, we propose an offline framework to achieve precise localization on any convolutional layer of a classification model by exploiting two kinds of gradients, called Dual-Gradients Localization (DGL) framework. DGL framework is developed based on two branches: 1) Pixel-level Class Selection, leveraging gradients of the target class to identify the correlation ratio of pixels to the target class within any convolutional feature maps, and 2) Class-aware Enhanced Maps, utilizing gradients of classification loss function to mine entire target object regions, which would not damage classification performance. Extensive experiments on public ILSVRC and CUB-200-2011 datasets show the effectiveness of the proposed DGL framework. Especially, our DGL obtains a new state-of-the-art Top-1 localization error of 43.55% on the ILSVRC benchmark. Chuangchuang Tan, Guanghua Gu, Shikui Wei, Yao Zhao 0001 |
ACM Multimedia | 4 |
| 2020 | R-PFN: Towards Precise Object Detection by Recurrent Pyramidal Feature Fusion
Qifei Jia, Shikui Wei |
PRCV (1) | 2 |
| 2020 | Rich Features Embedding for Cross-Modal Retrieval: A Simple BaselineabstractDuring the past few years, significant progress has been made on cross-modal retrieval, benefiting from the development of deep neural networks. Meanwhile, the overall frameworks are becoming more and more complex, making the training as well as the analysis more difficult. In this paper, we provide a Rich Features Embedding (RFE) approach to tackle the cross-modal retrieval tasks in a simple yet effective way. RFE proposes to construct rich representations for both images and texts, which is further leveraged to learn the rich features embedding in the common space according to a simple hard triplet loss. Without any bells and whistles in constructing complex components, the proposed RFE is concise and easy to implement. More importantly, our RFE obtains the state-of-the-art results on several popular benchmarks such as MS COCO and Flickr 30 K. In particular, the image-to-text and text-to-image retrieval achieve 76.1% and 61.1% (R@1) on MS COCO, which outperform others more than 3.4% and 2.3%, respectively. We hope our RFE will serve as a solid baseline and help ease future research in cross-modal retrieval. Xin Fu 0009, Yao Zhao 0001, Yunchao Wei, Shikui Wei |
IEEE Trans. Multim. | 5 |
| 2020 | Referring Image Segmentation by Generative Adversarial LearningabstractReferring expression is a kind of language expression being used for referring to particular objects. In this paper, we focus on the problem of image segmentation from natural language referring expressions. Existing works tackle this problem by augmenting the convolutional semantic segmentation networks with an LSTM sentence encoder, which is optimized by a pixel-wise classification loss. We argue that the distribution similarity between the inference and ground truth plays an important role in referring image segmentation. Therefore we introduce a complementary loss considering the consistency between the two distributions. To this end, we propose to train the referring image segmentation model in a generative adversarial fashion, which well addresses the distribution similarity problem. In particular, the proposed adversarial semantic guidance network (ASGN) includes the following advantages: a) more detailed visual information is incorporated by the detail enhancement; b) semantic information counteracts the word embedding impact; c) the proposed adversarial learning approach relieves the distribution inconsistencies. Experimental results on four standard datasets show significant improvements over all the compared baseline models, demonstrating the effectiveness of our method. Yao Zhao 0001, Jianbo Jiao, Yunchao Wei, Shikui Wei |
IEEE Trans. Multim. | 5 |
| 2020 | Parameter Distribution Balanced CNNsabstractConvolutional neural network (CNN) is the primary technique that has greatly promoted the development of computer vision technologies. However, there is little research on how to allocate parameters in different convolution layers when designing CNNs. We research mainly on revealing the relationship between CNN parameter distribution, i.e., the allocation of parameters in convolution layers, and the discriminative performance of CNN. Unlike previous works, we do not append more elements into the network, such as more convolution layers or denser short connections. We focus on enhancing the discriminative performance of CNN through varying its parameter distribution under strict size constraint. We propose an energy function to represent the CNN parameter distribution, which establishes the connection between the allocation of parameters and the discriminative performance of CNN. Extensive experiments with shallow CNNs on three public image classification data sets demonstrate that the CNN parameter distribution with a higher energy value will promote the model to obtain better performance. According to the motivated observation, the problem of finding the optimal parameter distribution can be transformed into an optimization problem of finding the biggest energy value. We present a simple yet effective guideline that uses balanced parameter distribution to design CNNs. Extensive experiments on ImageNet with three popular backbones, i.e., AlexNet, ResNet34, and ResNet101, demonstrate that the proposed guideline can make consistent improvements upon different baselines under strict size constraint. Lixin Liao, Yao Zhao 0001, Shikui Wei, Yunchao Wei, Jingdong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Devil in the Details: Towards Accurate Single and Multiple Human ParsingabstractHuman parsing has received considerable interest due to its wide application potentials. Nevertheless, it is still unclear how to develop an accurate human parsing system in an efficient and elegant way. In this paper, we identify several useful properties, including feature resolution, global context information and edge details, and perform rigorous analyses to reveal how to leverage them to benefit the human parsing task. The advantages of these useful properties finally result in a simple yet effective Context Embedding with Edge Perceiving (CE2P) framework for single human parsing. Our CE2P is end-to-end trainable and can be easily adopted for conducting multiple human parsing. Benefiting the superiority of CE2P, we won the 1st places on all three human parsing tracks in the 2nd Look into Person (LIP) Challenge. Without any bells and whistles, we achieved 56.50% (mIoU), 45.31% (mean APr) and 33.34% (APp0.5) in Track 1, Track 2 and Track 5, which outperform the state-of-the-arts more than 2.06%, 3.81% and 1.87%, respectively. We hope our CE2P will serve as a solid baseline and help ease future research in single/multiple human parsing. Code has been made available at https://github.com/liutinglt/CE2P. Ting Liu 0012, Yunchao Wei, Shikui Wei, Yao Zhao 0001 |
AAAI | 5 |
| 2019 | A Visual Perspective for User Identification Based on Camera Fingerprint
Xiang Jiang 0005, Shikui Wei, Ruizhen Zhao, Ruoyu Liu, Yao Zhao 0001 |
ICIG (2) | 2 |
| 2019 | Face Verification Between ID Document Photos and Partial Occluded Spot Photos
Shikui Wei, Xiang Jiang 0005, Yao Zhao 0001 |
ICIG (2) | 2 |
| 2019 | Adversarial task-specific learning
Xin Fu 0009, Yao Zhao 0001, Ting Liu 0012, Yunchao Wei, Jianan Li 0001, Shikui Wei |
Neurocomputing | 6 |
| 2019 | Improving image similarity estimation via global distance distribution information
Lixin Liao, Yao Zhao 0001, Shikui Wei |
Neurocomputing | 3 |
| 2019 | Magic-Wall: Visualizing Room Decoration by Enhanced Wall SegmentationabstractThis paper presents an intelligent system named Magic-wall, which enables visualization of the effect of room decoration automatically. Concretely, given an image of the indoor scene and a preferred color, the Magic-wall can automatically locate the wall regions in the image and smoothly replace the existing wall with the required one. The key idea of the proposed Magic-wall is to leverage visual semantics to guide the entire process of color substitution, including wall segmentation and replacement. To strengthen the reality of visualization, we make the following contributions. First, we propose an edge-aware fully convolutional neural network (Edge-aware-FCN) for indoor semantic scene parsing, in which a novel edge-prior branch is introduced to identify the boundary of different semantic regions better. To further polish the details between the wall and other semantic regions, we leverage the output of Edge-aware-FCN as the prior knowledge, concatenating with the image to form a new input for the Enhanced-Net. In such a case, the Enhanced-Net is able to capture more semantic-aware information from the input and polish some ambiguous regions. Finally, to naturally replace the color of the original walls, a simple yet effective color space conversion method is proposed for replacement with brightness reserved. We build a new indoor scene dataset upon ADE20K for training and testing, which includes six semantic labels. Extensive experimental evaluations and visualizations well demonstrate that the proposed Magic-wall is effective and can automatically generate a set of visually pleasing results. Ting Liu 0012, Yunchao Wei, Yao Zhao 0001, Si Liu 0001, Shikui Wei |
IEEE Trans. Image Process. | 5 |
| 2019 | Rearranging Online Tubes for Streaming Video Synopsis: A Dynamic Graph Coloring ApproachabstractTo efficiently browse long surveillance videos, the video synopsis technique is often used to rearrange tubes (i.e., tracks of moving objects) along the temporal axis to form a much shorter video. In this process, two key issues need to be addressed, i.e., the minimization of spatial tube collision and the maximization of temporal video condensation. In addition, when a surveillance video comes as a stream, an online algorithm with the capability of dynamically rearranging tubes is also required. Toward this end, this paper proposes a novel graph-based tube rearrangement approach for online video synopsis. The relationships among tubes are modeled with a dynamic graph, whose nodes (i.e., object masks of tubes) and edges (i.e., relationships) can be progressively inserted and updated. Based on this graph, we propose a dynamic graph coloring algorithm to efficiently rearrange all tubes by determining when they should appear. Extensive experimental results show that our approach can condense online surveillance video streams in real time with less tube collision and high compact ratio. Shikui Wei, Jia Li 0003, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Saliency Inside: Learning Attentive CNNs for Content-Based Image RetrievalabstractIn content-based image retrieval (CBIR), one of the most challenging and ambiguous tasks are to correctly understand the human query intention and measure its semantic relevance with images in the database. Due to the impressive capability of visual saliency in predicting human visual attention that is closely related to the query intention, this paper attempts to explicitly discover the essential effect of visual saliency in CBIR via qualitative and quantitative experiments. Toward this end, we first generate the fixation density maps of images from a widely used CBIR dataset by using an eye-tracking apparatus. These ground-truth saliency maps are then used to measure the influence of visual saliency to the task of CBIR by exploring several probable ways of incorporating such saliency cues into the retrieval process. We find that visual saliency is indeed beneficial to the CBIR task, and the best saliency involving scheme is possibly different for different image retrieval models. Inspired by the findings, this paper presents two-stream attentive CNNs with saliency embedded inside for CBIR. The proposed network has two streams that simultaneously handle two tasks. The main stream focuses on extracting discriminative visual features that are tightly related to semantic attributes. Meanwhile, the auxiliary stream aims to facilitate the main stream by redirecting the feature extraction to the salient image content that human may pay attention to. By fusing these two streams into the Main and Auxiliary CNNs (MAC), image similarity can be computed as the human being does by reserving conspicuous content and suppressing irrelevant regions. Extensive experiments show that the proposed model achieves impressive performance in image retrieval on four public datasets. Shikui Wei, Lixin Liao, Jia Li 0003, Qinjie Zheng, Yao Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Modality-Invariant Image-Text Embedding for Image-Sentence MatchingabstractPerforming direct matching among different modalities (like image and text) can benefit many tasks in computer vision, multimedia, information retrieval, and information fusion. Most of existing works focus on class-level image-text matching, called cross-modal retrieval , which attempts to propose a uniform model for matching images with all types of texts, for example, tags, sentences, and articles (long texts). Although cross-model retrieval alleviates the heterogeneous gap among visual and textual information, it can provide only a rough correspondence between two modalities. In this article, we propose a more precise image-text embedding method, image-sentence matching, which can provide heterogeneous matching in the instance level. The key issue for image-text embedding is how to make the distributions of the two modalities consistent in the embedding space. To address this problem, some previous works on the cross-model retrieval task have attempted to pull close their distributions by employing adversarial learning. However, the effectiveness of adversarial learning on image-sentence matching has not been proved and there is still not an effective method. Inspired by previous works, we propose to learn a modality-invariant image-text embedding for image-sentence matching by involving adversarial learning. On top of the triplet loss--based baseline, we design a modality classification network with an adversarial loss, which classifies an embedding into either the image or text modality. In addition, the multi-stage training procedure is carefully designed so that the proposed network not only imposes the image-text similarity constraints by ground-truth labels, but also enforces the image and text embedding distributions to be similar by adversarial learning. Experiments on two public datasets (Flickr30k and MSCOCO) demonstrate that our method yields stable accuracy improvement over the baseline model and that our results compare favorably to the state-of-the-art methods. Ruoyu Liu, Yao Zhao 0001, Shikui Wei, Liang Zheng 0001, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Indexing of the CNN features for the large scale image search
Ruoyu Liu, Shikui Wei, Yao Zhao 0001, Yi Yang 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Enhancing heterogeneous similarity estimation via neighborhood reversibility
Shikui Wei, Yao Zhao 0001, Shiming Ge |
Multim. Tools Appl. | 1 |
| 2018 | 3-D Surround View for Advanced Driver Assistance SystemsabstractAs the primary means of transportations in modern society, the automobile is developing toward the trend of intelligence, automation, and comfort. In this paper, we propose a more immersive 3-D surround view covering the automobiles around for advanced driver assistance systems. The 3-D surround view helps drivers to become aware of the driving environment and eliminates visual blind spots. The system first uses four fish-eye lenses mounted around a vehicle to capture images. Then, according to the pattern of image acquisition, camera calibration, image stitching, and scene generation, the 3-D surround driving environment is created. To achieve the real-time and easy-to-handle performance, we only use one image to finish the camera calibration through a special designed checkerboard. Furthermore, in the process of image stitching, a 3-D ship model is built to be the supporter, where texture mapping and image fusion algorithms are utilized to preserve the real texture information. The algorithms used in this system can reduce the computational complexity and improve the stitching efficiency. The fidelity of the surround view is also improved, thereby optimizing the immersion experience of the system under the premise of preserving the information of the surroundings. Chunyu Lin, Yao Zhao 0001, Xin Wang 0046, Shikui Wei |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | A Vehicle-Mounted Multi-camera 3D Panoramic Imaging Algorithm Based on Ship-Shaped Model
Xin Wang 0046, Chunyu Lin, Shikui Wei, Yao Zhao 0001 |
ICIG (3) | 5 |
| 2017 | Enhanced isomorphic semantic representation for cross-media retrievalabstractNowadays cross-media retrieval is an useful technology that helps people find expected information from the huge amount of multimodal data more efficiently. A common cross-media retrieval framework is first to map features of different modalities into an isomorphic semantic space so that the similarity between heterogeneous data can be measured. For most of semantic space based methods, the mapping mechanism from original to semantic space of each modality is optimized independently, yet the more discriminative characteristic of a certain modality is not taken into account. In this paper, we propose a deep framework which introduces a latent embedding layer to learn joint parameters to obtain semantically meaningful representations of images and texts. Specifically, the discriminative characteristic embedded in the textual modality can be transferred to images through the latent embedding layer and joint parameters to enhance the consistency between semantic representations. Extensive experiments on the three popular publicly available datasets well demonstrate the superiority of the proposed method, which achieves the new state-of-the-arts. Ting Liu 0012, Yao Zhao 0001, Shikui Wei, Yunchao Wei, Lixin Liao |
ICME | 3 |
| 2017 | Finding the Secret of CNN Parameter Layout under Strict Size ConstraintabstractAlthough deep convolutional neural networks (CNNs) have significantly boosted the performance of many computer vision tasks, their complexities~(the size or the number of parameters) are also dramatically increased even with slight performance improvement. However, the larger network leads to more computation requirements, which are unfavorable to resource-constrained scenarios, such as the widely used embedded systems. In this paper, we tentatively explore the essential effect of CNN parameter layout, ıe, the allocation of parameters in the convolution layers, on the discriminative capability of CNN. Instead of enlarging the breadth or depth of networks, we attempt to improve the discriminative ability of CNN by changing its parameter layout under strict size constraint. Toward this end, a novel energy function is proposed to represent the CNN parameter layout, which makes it possible to model the relationship between the allocation of parameters in the convolution layers and the discriminative ability of CNN. According to extensive experimental results with plain CNN models and Residual Nets, we find that the higher the energy of a specific CNN parameter layout is, the better its discriminative ability is. Following this finding, we propose a novel approach to learn the better parameter layout. Experimental results on two public image classification datasets show that the CNN models with the learned parameter layouts achieve the better image classification results under strict size constraint. Lixin Liao, Yao Zhao 0001, Shikui Wei, Jingdong Wang 0001, Ruoyu Liu |
ACM Multimedia | 3 |
| 2017 | Magic-wall: Visualizing Room DecorationabstractThis work focuses on Magic-wall, an automatic system for visualizing the effect of room decoration. Given an image of the indoor scene and a preferred color, the Magic-wall can automatically locate the wall regions in the image and smoothly replace the existing color with the required one. The key idea of the proposed Magic-wall is to leverage visual semantics to guide the entire process of color substitution including wall segmentation and color replacement. We propose an edge-aware fully convolutional neural network (FCN) for indoor semantic scene parsing, in which a novel edge-prior branch is introduced to better identify the boundary of different semantic regions. To accurately localize the wall regions, we adapt a semantic-dependent optimized strategy, which pays more attention to those pixels belonging to the wall by adapting larger optimization weights compared with those from other semantic regions. Finally, to naturally replace the color of original walls, a simple yet effective color space conversion method is proposed for replacement with brightness reservation. We build a new indoor scene dataset upon ADE20K for training and testing, which includes 6 semantic labels. Extensive experimental evaluations and visualizations well demonstrate that the proposed Magic-wall is effective and can automatically generate a set of visually pleasing results. Ting Liu 0012, Yunchao Wei, Yao Zhao 0001, Si Liu 0001, Shikui Wei |
ACM Multimedia | 5 |
| 2017 | Two-stream Attentive CNNs for Image RetrievalabstractIn content-based image retrieval, the most challenging (and ambiguous) part is to define the similarity between images. For the human-being, such similarity can be defined with respect to where they pay attention to and what semantic attributes they understand. Inspired by this fact, this paper presents two-stream attentive CNNs for image retrieval. As the human-being does, the proposed network has two streams that simultaneously handle two tasks. The Main stream focuses on extracting discriminative visual features that are tightly correlated with semantic attributes. Meanwhile, the Auxiliary stream aims to facilitate the main stream by redirecting the feature extraction operation mainly to the image content that human may pay attention to. By fusing these two streams into the Main and Auxiliary CNNs (MAC), image similarity can be computed as the human-being does by reserving the conspicuous content and suppressing the irrelevant regions. Extensive experiments show that the proposed model achieves impressive performance in image retrieval on four public datasets. Jia Li 0003, Shikui Wei, Qinjie Zheng, Ting Liu 0012, Yao Zhao 0001 |
ACM Multimedia | 3 |
| 2017 | Cross-Modal Retrieval With CNN Visual Features: A New BaselineabstractRecently, convolutional neural network (CNN) visual features have demonstrated their powerful ability as a universal representation for various recognition tasks. In this paper, cross-modal retrieval with CNN visual features is implemented with several classic methods. Specifically, off-the-shelf CNN visual features are extracted from the CNN model, which is pretrained on ImageNet with more than one million images from 1000 object categories, as a generic image representation to tackle cross-modal retrieval. To further enhance the representational ability of CNN visual features, based on the pretrained CNN model on ImageNet, a fine-tuning step is performed by using the open source Caffe CNN library for each target data set. Besides, we propose a deep semantic matching method to address the cross-modal retrieval problem with respect to samples which are annotated with one or multiple labels. Extensive experiments on five popular publicly available data sets well demonstrate the superiority of CNN visual features for cross-modal retrieval. Yunchao Wei, Yao Zhao 0001, Canyi Lu, Shikui Wei, Luoqi Liu, Zhenfeng Zhu, Shuicheng Yan |
IEEE Trans. Cybern. | 4 |
| 2016 | Improving the similarity estimation via score distributionabstractGenerally distance-based similarity estimation between two images is not always reliable due to the limitations in both image understanding techniques and distance measure methods. This paper presents a novel approach for improving the similarity estimation through introducing the distribution information of similarity scores. The key idea is based on an underlying assumption that the distributions of similarity scores are similar for true-relevant images when they query an independent database. By representing each distribution with the area under the corresponding similarity score curve, the difference between different distributions can be easily calculated and employed to update the original distance measure. Experiments on three public datasets with various feature representations show that the enhanced similarity estimation remarkably outperforms the original distance measure and the proposed approach also keeps a good generalization ability on various datasets and feature representations. Lixin Liao, Shikui Wei, Yao Zhao 0001, Guanghua Gu |
ICME | 2 |
| 2016 | A comparative evaluation: Different methods for simplifying the deep compositional featuresabstractAlthough deep compositional features have achieved an amazing performance in many application scenarios, it is not easy for them to directly tackle big data case due to their high computing time and memory usage. This paper presents a comparative evaluation for the simplification of deep compositional features by exploring the existing vector quantization and binarization techniques. Different techniques display different capabilities in similarity-preserving or memory-saving when projecting the original deep compositional features into more compact visual words or binary codes. We propose a dedicated image searching framework to evaluate all the techniques in terms of computational cost, memory usage and discrimination preserving. Extensive experiments demonstrate that it is feasible to greatly reduce computational cost and memory usage of the deep compositional features while preserving enough discriminative power. In addition, some useful conclusions are derived to guide the design of better simplifying schemes. Shikui Wei, Yao Zhao 0001 |
ICME | 2 |
| 2016 | Light-weight binary code embedding of local feature distribution in image search
Shikui Wei, Yao Zhao 0001, Jia Li 0003 |
Neurocomputing | 1 |
| 2016 | Modality-Dependent Cross-Media RetrievalabstractIn this article, we investigate the cross-media retrieval between images and text, that is, using image to search text (I2T) and using text to search images (T2I). Existing cross-media retrieval methods usually learn one couple of projections, by which the original features of images and text can be projected into a common latent space to measure the content similarity. However, using the same projections for the two different retrieval tasks (I2T and T2I) may lead to a tradeoff between their respective performances, rather than their best performances. Different from previous works, we propose a modality-dependent cross-media retrieval (MDCR) model, where two couples of projections are learned for different cross-media retrieval tasks instead of one couple of projections. Specifically, by jointly optimizing the correlation between images and text and the linear regression from one modal space (image or text) to the semantic space, two couples of mappings are learned to project images and text from their original feature spaces into two common latent subspaces (one for I2T and the other for T2I). Extensive experiments show the superiority of the proposed MDCR compared with other methods. In particular, based on the 4,096-dimensional convolutional neural network (CNN) visual feature and 100-dimensional Latent Dirichlet Allocation (LDA) textual feature, the mAP of the proposed method achieves the mAP score of 41.5%, which is a new state-of-the-art performance on the Wikipedia dataset. Yunchao Wei, Yao Zhao 0001, Zhenfeng Zhu, Shikui Wei, Yanhui Xiao, Jiashi Feng, Shuicheng Yan |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Cross-media hashing with Centroid ApproachingabstractCross-media retrieval has received increasing interest in recent years, which aims to addressing the semantic correlation issues within rich media. As two key aspects, cross-media representation and indexing have been studied for dealing with cross-media similarity measure and the scalability issue, respectively. In this paper, we propose a new cross-media hashing scheme, called Centroid Approaching Cross-Media Hashing (CAMH), to handle both cross-media representation and indexing simultaneously. Different from existing indexing methods, the proposed method introduces semantic category information into the learning procedure, leading to more exact hash codes of multiple media type instances. In addition, we present a comparative study of cross-media indexing methods under a unique evaluation framework. Extensive experiments on two commonly used datasets demonstrate the good performance in terms of search accuracy and time complexity. Ruoyu Liu, Yao Zhao 0001, Shikui Wei, Zhenfeng Zhu |
ICME | 3 |
| 2015 | Redundancy filtering and fusion verification for video copy detection
Shikui Wei, Su Jiang, Wenxian Jin, Yao Zhao 0001, Zhenfeng Zhu |
Multim. Syst. | 1 |
| 2015 | Accumulated reconstruction error vector (AREV): a semantic representation for cross-media retrieval
Shikui Wei, Yao Zhao 0001, Zhenfeng Zhu, Yunchao Wei, Changsheng Xu |
Multim. Tools Appl. | 2 |
| 2015 | Kernel Reconstruction ICA for Sparse RepresentationabstractIndependent component analysis with soft reconstruction cost (RICA) has been recently proposed to linearly learn sparse representation with an overcomplete basis, and this technique exhibits promising performance even on unwhitened data. However, linear RICA may not be effective for the majority of real-world data because nonlinearly separable data structure pervasively exists in original data space. Meanwhile, RICA is essentially an unsupervised method and does not employ class information. Motivated by the success of the kernel trick that maps a nonlinearly separable data structure into a linearly separable case in a high-dimensional feature space, we propose a kernel RICA (kRICA) model to nonlinearly capture sparse representation in feature space. Furthermore, we extend the unsupervised kRICA to a supervised one by introducing a class-driven discrimination constraint, such that the data samples from the same class are well represented on the basis of the corresponding subset of basis vectors. This discrimination constraint minimizes inhomogeneous representation energy and maximizes homogeneous representation energy simultaneously, which is essentially equivalent to maximizing between-class scatter and minimizing within-class scatter at the same time in an implicit manner. Experimental results demonstrate that the proposed algorithm is more effective than other state-of-the-art methods on several datasets. Yanhui Xiao, Zhenfeng Zhu, Yao Zhao 0001, Yunchao Wei, Shikui Wei |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2014 | Learning a mid-level feature space for cross-media regularizationabstractIn this paper, we propose a cross-media regularization framework to enhance image understanding which can benefit image retrieval, classification and so on. The goal of cross-media regularization is to find regularization projections by exploiting the correlations between visual features and textual features. Thus, the original noisy distribution of visual features can be refined by leveraging the discriminative distribution of the corresponding textual features. Within the proposed cross-media regularization framework, a mid-level representation is built by jointly projecting both visual and textual features into a shared feature subspace, which leads to transferring of the discriminative semantic characteristic embedded in the textual modality into the corresponding visual modality. Meanwhile, the discriminative characteristic of textual features can also be boosted simultaneously. The experimental results demonstrate that the proposed mid-level space learning process can remarkably improve the search quality and outperform the existing semantic regularization methods. Yunchao Wei, Yao Zhao 0001, Zhenfeng Zhu, Yanhui Xiao, Shikui Wei |
ICME | 5 |
| 2014 | Topographic NMF for Data RepresentationabstractNonnegative matrix factorization (NMF) is a useful technique to explore a parts-based representation by decomposing the original data matrix into a few parts-based basis vectors and encodings with nonnegative constraints. It has been widely used in image processing and pattern recognition tasks due to its psychological and physiological interpretation of natural data whose representation may be parts-based in human brain. However, the nonnegative constraint for matrix factorization is generally not sufficient to produce representations that are robust to local transformations. To overcome this problem, in this paper, we proposed a topographic NMF (TNMF), which imposes a topographic constraint on the encoding factor as a regularizer during matrix factorization. In essence, the topographic constraint is a two-layered network, which contains the square nonlinearity in the first layer and the square-root nonlinearity in the second layer. By pooling together the structure-correlated features belonging to the same hidden topic, the TNMF will force the encodings to be organized in a topographical map. Thus, the feature invariance can be promoted. Some experiments carried out on three standard datasets validate the effectiveness of our method in comparison to the state-of-the-art approaches. Yanhui Xiao, Zhenfeng Zhu, Yao Zhao 0001, Yunchao Wei, Shikui Wei, Xuelong Li 0001 |
IEEE Trans. Cybern. | 5 |
| 2014 | Mining Semantically Consistent Patterns for Cross-View DataabstractIn some real world applications, like information retrieval and data classification, we often are confronted with the situation that the same semantic concept can be expressed using different views with similar information. Thus, how to obtain a certain Semantically Consistent Patterns (SCP) for cross-view data, which embeds the complementary information from different views, is of great importance for those applications. However, the heterogeneity among cross-view representations brings a significant challenge on mining the SCP. In this paper, we propose a general framework to discover the SCP for cross-view data. Specifically, aiming at building a feature-isomorphic space among different views, a novel Isomorphic Relevant Redundant Transformation (IRRT) is first proposed. The IRRT linearly maps multiple heterogeneous low-level feature spaces to a high-dimensional redundant feature-isomorphic one, which we name as mid-level space. Thus, much more complementary information from different views can be captured. Furthermore, to mine the semantic consistency among the isomorphic representations in the mid-level space, we propose a new Correlation-based Joint Feature Learning (CJFL) model to extract a unique high-level semantic subspace shared across the feature-isomorphic data. Consequently, the SCP for cross-view data can be obtained. Comprehensive experiments on three data sets demonstrate the advantages of our framework in classification and retrieval. Lei Zhang 0116, Yao Zhao 0001, Zhenfeng Zhu, Shikui Wei, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Neighborhood reversibility verifying for image searchabstractThe neighborhood structure can significantly impact the effectiveness of image search, and fulfilling the reversibility of neighborhood may improve the image search quality. This paper proposes an effective and efficient scheme for reconstructing the symmetry relationship of k-nearest neighborhood (KNN). In particular, we design a verifying function to learn the prior knowledge of neighborhood reversibility among images. By exploiting the prior knowledge, the image search system will give higher rank to those images that satisfy the reversibility of KNN relationship with the query. In addition, we systematically investigate the sensitivity of neighborhood size on image search quality and propose an adaptive selection scheme for improving robustness of neighborhood reversibility learning methods. The extensive experimental results show that the proposed scheme remarkably improves the image search quality and give a comparable but more stable performance to the state-of-the-art method for various image datasets. Yao Zhao 0001, Shikui Wei, Zhenfeng Zhu |
ICME | 3 |
| 2013 | Orthogonal graph-regularized matrix factorization and its application for recommendationabstractAs one of the most successful approaches for recommendation, matrix factorization based Collaborative Filtering (CF) technique has received considerable attentions over the past years. In this paper, we propose an orthogonal matrix factorization model with graph regularization to preserve the consistency of the local structure both in user and item spaces, respectively. Instead of traditional alternating optimization method, a greedy sequential one is introduced to optimize a pair of coupled factor vector and its corresponding loading vector simultaneously each time, thus the original optimization problem is converted into the well-studied Multivariate Eigen Problem (MEP). Furthermore, multiple pairs of coupled eigen-vectors can be obtained in sequence. To guarantee nonrecurring of repetition of solutions, a novel dual-deflation technique is developed to incorporate into the sequential optimization. Experimental results on MovieLens and Each Movie data sets demonstrate that the proposed method is much more competitive compared with the state of the art matrix factorization based collaborative filtering methods. Zhenfeng Zhu, Peilu Xin, Shikui Wei, Yao Zhao 0001 |
ICME | 3 |
| 2013 | Joint Optimization Toward Effective and Efficient Image SearchabstractThe bag-of-words (BoW) model has been known as an effective method for large-scale image search and indexing. Recent work shows that the performance of the model can be further improved by using the embedding method. While different variants of the BoW model and embedding method have been developed, less effort has been made to discover their underlying working mechanism. In this paper, we systematically investigate the image search performance variation with respect to a few factors of the BoW model, and study how to employ the embedding method to further improve the image search performance. Subsequently, we summarize several observations based on the experiments on descriptor matching. To validate these observations in a real image search, we propose an effective and efficient image search scheme, in which the BoW model and embedding method are jointly optimized in terms of effectiveness and efficiency by following these observations. Our comprehensive experiments demonstrate that it is beneficial to employ these observations to develop an image search algorithm, and the proposed image search scheme outperforms state-of-the art methods in both effectiveness and efficiency. Shikui Wei, Dong Xu 0001, Xuelong Li 0001, Yao Zhao 0001 |
IEEE Trans. Cybern. | 1 |
| 2012 | Discriminative ICA model with reconstruction constraint for image classificationabstractIndependent Component Analysis (ICA) is an effective unsupervised tool to learn statistically independent representations. However, ICA is not only sensitive to whitening but also difficult to learn an over-complete basis set. Consequently, ICA with soft Reconstruction cost(RICA) was presented to learn sparse representations with over-complete basis even on unwhitened data. Nevertheless, this model may not be an optimal discriminative model for classification tasks, because it failed to consider the association between the training sample and its class. In this paper, we propose a supervised Discriminative ICA model with Reconstruction constraint for image classification, named DRICA. DRICA brings in class information to learn the over-complete basis by incorporating inhomogeneous representation cost constraint into the RICA framework. This constraint leads to partition the set of basis vectors into several subsets corresponding to the sample classes, where each subset could sparsely model data samples from the same class but not others. Therefore, the proposed ICA model can learn an over-complete basis and an optimal multi-class classifier jointly. Some experiments carried out on several standard image databases validate the effectiveness of DRICA for image classification. Yanhui Xiao, Zhenfeng Zhu, Shikui Wei, Yao Zhao 0001 |
ACM Multimedia | 3 |
| 2011 | Copy detection towards semantic mining for video retrievalabstractIn large-scale video database, lots of different videos frequently share the similar content copied from the same source. Generally, those videos have certain semantic correlations, such as being of similar events and sharing the same topic. Mining these semantic correlations can greatly facilitate video search. However, as a preprocessing step, detecting and localizing the copy pair among videos, i.e. copy detection problem, plays a key role for precise semantic mining. To meet the requirements in semantic mining scenario, we propose a frame fusion based copy detection scheme. In this scheme, the copy detection problem is converted to HMM decoding problem with three relaxed constraints, where Viterbi algorithm is employed to automatically detect the copy pair. The experimental results show that the proposed approach achieves high localization accuracy even when the copied clip undergoes some complex transformations, while achieving comparable performance compared with state-of-the-art copy detection methods. Shikui Wei, Yao Zhao 0001, Changsheng Xu, Dong Xu 0001 |
ICIP | 1 |
| 2011 | Frame Fusion for Video Copy DetectionabstractContent-based video copy detection is very important for copyright protection in view of the growing popularity of video sharing websites, which deals with not only whether a copy occurs in a query video stream but also where the copy is located and where the copy is originated from. While a lot of work has addressed the problem with good performance, less effort has been made to consider the copy detection problem in the case of a continuous query stream, for which precise temporal localization and some complex video transformations like frame insertion and video editing need to be handled. We attempt to attack the problem by presenting a frame fusion based copy detection approach, which converts video copy detection to frame similarity search and frame fusion under a temporal consistency assumption. Our work focuses mainly on the frame fusion stage due to its critical role in copy detection performance. The proposed frame fusion scheme is based on a Viterbi-like algorithm, comprising an online back-tracking strategy with three relaxed constraints. The experimental results show that the proposed approach achieves high localization accuracy in both the query stream and the reference database even when a query video stream undergoes some complex transformations, while achieving comparable performance compared with state-of-the-art copy detection methods. Shikui Wei, Yao Zhao 0001, Ce Zhu, Changsheng Xu, Zhenfeng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Multimodal Fusion for Video Search RerankingabstractAnalysis on click-through data from a very large search engine log shows that users are usually interested in the top-ranked portion of returned search results. Therefore, it is crucial for search engines to achieve high accuracy on the top-ranked documents. While many methods exist for boosting video search performance, they either pay less attention to the above factor or encounter difficulties in practical applications. In this paper, we present a flexible and effective reranking method, called CR-Reranking, to improve the retrieval effectiveness. To offer high accuracy on the top-ranked results, CR-Reranking employs a cross-reference (CR) strategy to fuse multimodal cues. Specifically, multimodal features are first utilized separately to rerank the initial returned results at the cluster level, and then all the ranked clusters from different modalities are cooperatively used to infer the shots with high relevance. Experimental results show that the search quality, especially on the top-ranked results, is improved significantly. Shikui Wei, Yao Zhao 0001, Zhenfeng Zhu, Nan Liu 0007 |
IEEE Trans. Knowl. Data Eng. | 1 |