VLDB 2026 Research / reviewers in the wild / expert
Weimin Tan
dblp:64/11238
· DBLP profile ↗
64ranked-venue papers
9as first author
44since 2021 · last 2025
0000-0001-7677-4772ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 7 first-author · 37 since 2021Artificial intelligence and machine learning · 22 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Online Training for Zero-Shot Time-Lapse Microscopy Denoising and Super-ResolutionabstractIn time-lapse microscopy, inherent noise significantly limits imaging sensitivity and increases measurement uncertainty. Due to the scarcity of clean data, zero-shot approaches have emerged as highly data-efficient solutions for microscopy denoising. However, existing methods typically process video frames independently, resulting in long training times and issues such as temporal noise and over-smoothing. In this paper, we introduce MDSR-Zero, a zero-shot online learning method designed for plug-and-play noise suppression and super-resolution of microscopy videos. Our approach leverages an efficient online training strategy that reuses denoising models from previous frames. By treating the video as a continuous stream, our model significantly reduces training time and ensures temporally consistent denoising. Additionally, we propose a novel loss function tailored for denoising in the context of super-resolution, which enhances the detail in the denoised results. Extensive experiments on both synthetic and real-world noise demonstrate that our method achieves state-of-the-art performance among zero-shot denoising approaches and is competitive with self-supervised methods. Notably, our method can reduce training time by up to 10x compared to the previous SOTA method. Ruian He, Ri Cheng, Xinkai Lyu, Weimin Tan, Bo Yan 0001 |
AAAI | 4 |
| 2025 | Multimodal Inference with Incremental Tabular AttributesabstractMultimodal Learning with visual and tabular modalities has become more and more popular nowadays, especially in the healthcare area. Due to the adaptation of new equipment or new factors being introduced, the tabular modality keeps changing. However, the standard process of training multimodal AI models requires tables to have fixed columns in training and inference; thus, it is not suitable for handling dynamically changed tables. Therefore, new methods are needed for efficiently handling such tables in multimodal learning. In this paper, we introduce a new task, multimodal inference with incremental tabular attributes, which aims to enable trained multimodal models to leverage incremental attributes in tabular modality during the inference stage efficiently. We implement a specialized encoder to disentangle the latent representation of incremental tabular attributes inside itself and with the old attributes to reduce information redundancy and further align the incremental attributes with the visual modality with consistency loss to improve information richness. Experimental results across five public datasets show that our method effectively utilizes incremental tabular attributes, achieving state-of-the-art performance in general scenarios. Beyond the inference, we also find that our method achieved better performance in fully supervised settings, evoking a new training style for multimodal learning with tables. Xinda Chen, Zixian Zhang, Weimin Tan |
IJCAI | 4 |
| 2025 | Efficient Trajectory Space-Time Super-Resolution for Fast Live-cell ImagingabstractLive-cell imaging is a powerful tool for studying dynamic subcellular processes by capturing the spatiotemporal organization of the biological microenvironment. However, limitations due to phototoxicity and photobleaching prevent microscopes from achieving high frame rates and high-quality images. Although current deep learning methods can enhance both frame rates and image resolution without compromising cell health, they often overlook the continuity of subcellular trajectories, which leads to discontinuous temporal modeling. It also incurs prohibitive computational costs due to exhaustive correlation computation that hinder real-time applications. To address these issues with high efficiency, we propose Trajectory Space-Time Super-Resolution (T-STSR), a method designed to boost frame rates and resolution in fast subcellular imaging while significantly reducing computational overhead. Our approach incorporates Spatial-Temporal Trajectory Modeling (STTM), which learns a state-space model over spatiotemporal slices to reconstruct particle trajectories at low cost. In addition, our novel Trajectory-Aware Loss randomly subsamples trajectory data during training, promoting continuous trajectory representation and mitigating noise with minimal additional computation. We validated T-STSR on both synthesized and real-world datasets with various particle types and noise conditions, demonstrating that our method achieves superior restoration results while saving 75% inference time compared to the previous SOTA model. Ruian He, Zixian Zhang, Ri Cheng, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 4 |
| 2025 | Scaling Laws for Data-Efficient Visual Transfer Learning
Wenxuan Yang, Qingqv Wei, Chenxi Ma, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 4 |
| 2025 | MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from TextbooksabstractMedical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis remains underdeveloped, primarily due to less specialized text descriptions in current dermatology multimodal datasets. To address this issue, we propose MM-Skin, the first large-scale multimodal dermatology dataset that encompasses 3 imaging modalities, including clinical, dermoscopic, and pathological and nearly 10k high-quality image-text pairs collected from professional textbooks. In addition, we generate over 27k diverse, instruction-following vision question answering (VQA) samples (9× the size of current largest dermatology VQA dataset). Leveraging public datasets and MM-Skin, we developed SkinVL, a dermatology-specific VLM designed for precise and nuanced skin disease interpretation. Comprehensive benchmark evaluations of SkinVL on VQA, supervised fine-tuning (SFT) and zero-shot classification tasks across 8 datasets, reveal its exceptional performance for skin diseases in comparison to both general and medical VLM models. The introduction of MM-Skin and SkinVL offers a meaningful contribution to advancing the development of clinical dermatology VLM assistants. Code and dataset are available at https://github.com/ZwQ803/MM-Skin. Chenxi Ma, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 4 |
| 2025 | TabiMed: Tabularizing Medical Images for Few-Shot In-Context DiagnosisabstractAchieving accurate predictions with limited samples is a key challenge in biomedical image artificial intelligence. Previous methods rely on pre-trained image foundation models with supervised fine-tuning (SFT) or zero-shot inference to enhance small-data performance. However, SFT is time-consuming and prone to overfitting, whereas zero-shot inference fails to fully exploit available data. Inspired by recent tabular foundation models, which show superior performance on small-sample tasks with in-context learning (ICL), we propose TabiMed, a novel framework that transforms visual representations into structured tabular data, leveraging pre-trained tabular models for fast and accurate analysis on small data. TabiMed consists of three key components: dynamic modality-aware representation engine, tabularization adapter and in-context inference module. Experiments on 10 datasets from different fields demonstrate three major advantages of TabiMed: 1) excellent performance on small datasets, with an average AUC of 14.1% higher than zero-shot; 2) high efficiency, with a training time 250x faster than SFT; 3) scalability to larger datasets through our tabularization adapter. TabiMed proposes a novel pathway to address the challenges of analyzing biomedical images with few samples. Wanying Zhou, Yu Ling, Chenxi Ma, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 6 |
| 2025 | A hybrid genetic algorithm for the vehicle relocation problem with ride-sharing options in one-way car-sharing systemsabstractThe imbalance of idle cars at different stations remains a critical challenge in one-way car-sharing systems. This paper proposes a novel mixed user-operator-based relocation strategy for this problem. In this one-way car-sharing system, ride-sharing service is allowed, and customers can share trips with others by a rental vehicle. Ride-sharing, as a supplement to operator-based relocation, can relieve the pressure of vehicle relocation, lowering the relocation fee and reducing the required fleet size. In this study, the operators must determine a mixed vehicle relocation scheme, including operator-based vehicle relocation routes and user-based ride-sharing matches. This problem can be defined as a bi-objective mixed-integer linear programming model to minimize total user fees and maximize system benefits. The linear weighting method can combine those two objectives into one objective. To solve this problem, we propose a meta-heuristic algorithm based on the state-of-the-art hybrid genetic search with adaptive diversity control (HGSADC). The computational results show that the proposed algorithm can produce high-quality solutions within acceptable computing time. We also show that the proposed mixed vehicle relocation strategy can benefit operators and users. Weimin Tan, Min Kong, Muhammet Deveci, Witold Pedrycz |
Knowl. Based Syst. | 1 |
| 2025 | Spatiotemporal-Aware Self-Supervised Fluorescence Microscopy Image DenoisingabstractFluorescence microscopy has been an indispensable tool in many scientific disciplines. However, the expensive imaging cost and the photo-toxicity problem make it difficult to obtain high-quality images. The independent shot noise in fluorescence microscopy images always overwhelms signals and limits the imaging resolution, hindering progress in related research. Recently, self-supervised image denoising has received wide attention for its ability to train a denoiser without paired low Signal-to-Noise Ratio (SNR) and high SNR images. Existing self-supervised fluorescence microscopy image denoising works either suffer from high imaging/computational cost or large training difficulty. Here, we propose a Single-image based Self-Supervised Denoising approach (TriS-D) by utilizing the spatiotemporal redundancy of the fluorescence microscopy imaging data, which facilitates the low-cost and convenient training. The TriS-D can generate the training data from a raw image, not only releasing the demand for multiple low SNR time-lapse imaging data but also enabling the building of a 2D convolution-based model. Comprehensive experiments across different imaging modalities and biological samples verify the effectiveness of the TriS-D. Chenxi Ma, Weimin Tan, Zhaohui Zhou, Bo Yan 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Context-Aware Iteration Policy Network for Efficient Optical Flow EstimationabstractExisting recurrent optical flow estimation networks are computationally expensive since they use a fixed large number of iterations to update the flow field for each sample. An efficient network should skip iterations when the flow improvement is limited. In this paper, we develop a Context-Aware Iteration Policy Network for efficient optical flow estimation, which determines the optimal number of iterations per sample. The policy network achieves this by learning contextual information to realize whether flow improvement is bottlenecked or minimal. On the one hand, we use iteration embedding and historical hidden cell, which include previous iterations information, to convey how flow has changed from previous iterations. On the other hand, we use the incremental loss to make the policy network implicitly perceive the magnitude of optical flow improvement in the subsequent iteration. Furthermore, the computational complexity in our dynamic network is controllable, allowing us to satisfy various resource preferences with a single trained model. Our policy network can be easily integrated into state-of-the-art optical flow networks. Extensive experiments show that our method maintains performance while reducing FLOPs by about 40%/20% for the Sintel/KITTI datasets. Ri Cheng, Ruian He, Xuhao Jiang, Shili Zhou, Weimin Tan, Bo Yan 0001 |
AAAI | 5 |
| 2024 | Low-Latency Space-Time Supersampling for Real-Time RenderingabstractWith the rise of real-time rendering and the evolution of display devices, there is a growing demand for post-processing methods that offer high-resolution content in a high frame rate. Existing techniques often suffer from quality and latency issues due to the disjointed treatment of frame supersampling and extrapolation. In this paper, we recognize the shared context and mechanisms between frame supersampling and extrapolation, and present a novel framework, Space-time Supersampling (STSS). By integrating them into a unified framework, STSS can improve the overall quality with lower latency. To implement an efficient architecture, we treat the aliasing and warping holes unified as reshading regions and put forth two key components to compensate the regions, namely Random Reshading Masking (RRM) and Efficient Reshading Module (ERM). Extensive experiments demonstrate that our approach achieves superior visual fidelity compared to state-of-the-art (SOTA) methods. Notably, the performance is achieved within only 4ms, saving up to 75\% of time against the conventional two-stage pipeline that necessitates 17ms. Ruian He, Shili Zhou, Ri Cheng, Weimin Tan, Bo Yan 0001 |
AAAI | 5 |
| 2024 | MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation LocalizationabstractDeep learning-based models have made great progress in image tampering localization, which aims to distinguish between manipulated and authentic regions. However, these models suffer from inefficient training. This is because they use ground-truth mask labels mainly through the cross-entropy loss, which prioritizes per-pixel precision but disregards the spatial location and shape details of manipulated regions. To address this problem, we propose a Mask-Guided Query-based Transformer Framework (MGQFormer), which uses ground-truth masks to guide the learnable query token (LQT) in identifying the forged regions. Specifically, we extract feature embeddings of ground-truth masks as the guiding query token (GQT) and feed GQT and LQT into MGQFormer to estimate fake regions, respectively. Then we make MGQFormer learn the position and shape information in ground-truth mask labels by proposing a mask-guided loss to reduce the feature distance between GQT and LQT. We also observe that such mask-guided training strategy has a significant impact on the convergence speed of MGQFormer training. Extensive experiments on multiple benchmarks show that our method significantly improves over state-of-the-art methods. Kunlun Zeng, Ri Cheng, Weimin Tan, Bo Yan 0001 |
AAAI | 3 |
| 2024 | SAMFlow: Eliminating Any Fragmentation in Optical Flow with Segment Anything ModelabstractOptical Flow Estimation aims to find the 2D dense motion field between two frames. Due to the limitation of model structures and training datasets, existing methods often rely too much on local clues and ignore the integrity of objects, resulting in fragmented motion estimation. Through theoretical analysis, we find the pre-trained large vision models are helpful in optical flow estimation, and we notice that the recently famous Segment Anything Model (SAM) demonstrates a strong ability to segment complete objects, which is suitable for solving the fragmentation problem. We thus propose a solution to embed the frozen SAM image encoder into FlowFormer to enhance object perception. To address the challenge of in-depth utilizing SAM in non-segmentation tasks like optical flow estimation, we propose an Optical Flow Task-Specific Adaption scheme, including a Context Fusion Module to fuse the SAM encoder with the optical flow context encoder, and a Context Adaption Module to adapt the SAM features for optical flow task with Learned Task-Specific Embedding. Our proposed SAMFlow model reaches 0.86/2.10 clean/final EPE and 3.55/12.32 EPE/F1-all on Sintel and KITTI-15 training set, surpassing Flowformer by 8.5%/9.9% and 13.2%/16.3%. Furthermore, our model achieves state-of-the-art performance on the Sintel and KITTI-15 benchmarks, ranking #1 among all two-frame methods on Sintel clean pass. Shili Zhou, Ruian He, Weimin Tan, Bo Yan 0001 |
AAAI | 3 |
| 2024 | Bridging The Domain Gap Arising from Text Description Differences for Stable Text-To-Image GenerationabstractGenerating high-quality images that conform to the semantics of captions has numerous potential applications. However, text-to-image generation is a challenging task due to its cross-modality nature. Current generative models are typically unstable, meaning that complex sentences can result in poor image quality. In this paper, we propose a novel model to bridge the domain gap arising from sentence complexity to achieve stable text-to-image generation. Our model includes two key modules, the attribute extraction module and the attribute fusion module. These modules can extract attributes from the captions and fuse them with image features to encourage the model to accurately understand the semantics. Our modules are plug-and-play and extensive experiments demonstrate that our approach outperforms the state-of-the-art GAN model. Our code and trained model are available at https://github.com/tantian21/stable-t2i-generation. Tian Tan 0016, Weimin Tan, Xuhao Jiang, Yueming Jiang, Bo Yan 0001 |
ICASSP | 2 |
| 2024 | Facial Micro-Motion-Aware Mixup for Micro-Expression RecognitionabstractData-driven learning models have demonstrated strong benefits in capturing subtle facial movements for micro-expression recognition (MER), but are limited by the available data. Generative models can generate a variety of new data, but are typically computationally prohibitive compared to efficient Mixup-like methods. In this paper, we propose a novel Facial Micro-Motion-Aware Mixup approach for MER, namely MEMix. Our MEMix constructs a micro-motion-aware mask to select the most salient facial motions and generate a new sample with a mixed motion feature. This mixed motion feature can effectively expand the data distribution, leading to smoother decision boundaries for MER models. To demonstrate the good generality of MEMix, we integrate it with three advanced vision transformer-based models. The results show that the three integrated models consistently achieve performance improvements ranging from 4.07% to 7.32% in accuracy and from 6.54% to 9.18% in F1-score. Besides, to further explore the ability of MEMix, we propose a two-stream network called MixMeFormer, which unlocks the potential of the transformer by simply integrating mixed motion features with facial semantics for MER. Extensive experiments demonstrate that our MixMeFormer outperforms other state-of-the-art methods on three well-known micro-expression datasets. Zhuoyao Gu, Miao Pang, Weimin Tan, Xuhao Jiang, Bo Yan 0001 |
ICASSP | 4 |
| 2024 | Addressing Imbalance for Class Incremental Learning in Medical Image ClassificationabstractDeep convolutional neural networks have made significant breakthroughs in medical image classification, under the assumption that training samples from all classes are simultaneously available. However, in real-world medical scenarios, there's a common need to continuously learn about new diseases, leading to the emerging field of class incremental learning (CIL) in the medical domain. Typically, CIL suffers from catastrophic forgetting when trained on new classes. This phenomenon is mainly caused by the imbalance between old and new classes, and it becomes even more challenging with imbalanced medical datasets. In this work, we introduce two simple yet effective plug-in methods to mitigate the adverse effects of the imbalance. First, we propose a CIL-balanced classification loss to mitigate the classifier bias toward majority classes via logit adjustment. Second, we propose a distribution margin loss that not only alleviates the inter-class overlap in embedding space but also enforces the intra-class compactness. We evaluate the effectiveness of our method with extensive experiments on three benchmark datasets (CCH5000, HAM10000, and EyePACS). The results demonstrate that our approach outperforms state-of-the-art methods. Xuze Hao, Wenqian Ni, Xuhao Jiang, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 4 |
| 2024 | FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed ModelabstractFacial movements play a crucial role in conveying altitude and intentions, and facial optical flow provides a dynamic and detailed representation of it. However, the scarcity of datasets and a modern baseline hinders the progress in facial optical flow research. This paper proposes FacialFlowNet (FFN), a novel large-scale facial optical flow dataset, and the Decomposed Facial Flow Model (DecFlow), the first method capable of decomposing facial flow. FFN comprises 9,635 identities and 105,970 image pairs, offering unprecedented diversity for detailed facial and head motion analysis. DecFlow features a facial semantic-aware encoder and a decomposed flow decoder, excelling in accurately estimating and decomposing facial flow into head and expression components. Comprehensive experiments demonstrate that FFN significantly enhances the accuracy of facial flow estimation across various optical flow methods, achieving up to an 11% reduction in Endpoint Error (EPE) (from 3.91 to 3.48). Moreover, DecFlow, when coupled with FFN, outperforms existing methods in both synthetic and real-world scenarios, enhancing facial expression analysis. The decomposed expression flow achieves a substantial accuracy improvement of 18% (from 69.1% to 82.1%) in micro-expressions recognition. These contributions represent a significant advancement in facial motion analysis and optical flow estimation. Codes and datasets can be found. Jianzhi Lu, Ruian He, Shili Zhou, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 4 |
| 2024 | Learning Cross-Spectral Prior for Image Super-ResolutionabstractWith the rising interest in multi-camera cross-spectral systems, cross-spectral images have been widely used in computer vision and image processing. Therefore, an effective super-resolution (SR) method provides high-resolution (HR) cross-spectral images for different research and applications. However, existing SR methods rarely consider utilizing cross-spectral information to assist the SR of visible images. They cannot handle complex degradation (noise, high brightness, low light) and misalignment problems in low-resolution (LR) cross-spectral images. Here, we first explore the potential of using near-infrared (NIR) image guidance for better SR, based on the observation that NIR images can preserve valuable information for recovering adequate image details. To take full advantage of the cross-spectral prior, we propose a novel Cross-Spectral Prior guided image SR approach (CSPSR). The cross-view matching (CVM) module and the dynamic multi-modal fusion (DMF) module can enhance the spatial correlation between cross-spectral images and bridge the multi-modal feature gap, respectively. Extensive experiments demonstrate the effectiveness of our CSPSR. Chenxi Ma, Weimin Tan, Shili Zhou, Bo Yan 0001 |
ACM Multimedia | 2 |
| 2024 | Audio-Driven Identity Manipulation for Face InpaintingabstractRecent advances in multimodal artificial intelligence have greatly improved the integration of vision-language-audio cues to enrich the content creation process. Inspired by these developments, in this paper, we first integrate audio into the face inpainting task to facilitate identity manipulation. Our main insight is that a person's voice carries distinct identity markers, such as age and gender, which provide an essential supplement for identity-aware face inpainting. By extracting identity information from audio as guidance, our method can naturally support tasks of identity preservation and identity swapping in face inpainting. Specifically, we introduce a dual-stream network architecture comprising a face branch and an audio branch. The face branch is tasked with extracting deterministic information from the visible parts of the input masked face, while the audio branch is designed to capture heuristic identity priors from the speaker's voice. The identity codes from two streams are integrated using a multi-layer perceptron (MLP) to create a virtual unified identity embedding that represennts comprehensive identity features. In addition, to explicitly exploit the information from audio, we introduce an audio-face generator to generate an 'fake' audio face directly from audio and fuse the multi-scale intermediate features from the audio-face generator into face inpainting network through an audio-visual feature fusion (AVFF) module. Extensive experiments demonstrate the positive impact of extracting identity information from audio on face inpainting task, especially in identity preservation. Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 3 |
| 2024 | A Medical Data-Effective Learning Benchmark for Highly Efficient Pre-training of Foundation Models
Wenxuan Yang, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 2 |
| 2024 | A Double Deep Q-Network framework for a flexible job shop scheduling problem with dynamic job arrivals and urgent job insertions
Shaojun Lu, Min Kong, Weimin Tan, Yingxin Song |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Prompt-Guided Semantic-Aware Distillation for Weakly Supervised Incremental Semantic SegmentationabstractWeakly Supervised Incremental Semantic Segmentation (WISS) aims to enable deep neural networks to incrementally learn new classes using only image-level labels without catastrophic forgetting. Despite WISS eliminating the usage of costly and time-consuming pixel-by-pixel annotations, the image-level labels can not provide details about the location of new classes, resulting in inferior performance. To address these issues, we take inspiration from zero-shot learning to model the inter-class semantic relation utilizing class names as text prompts, thereby facilitating knowledge transfer between classes. However, some class names of the segmentation datasets are polysemous. Thus, we design a new prompt template to better capture the semantic relation by appending synonyms and definitions of the corresponding classes. Guided by this semantic relation, we propose semantic relation weighted distillation to transfer the knowledge from old to new classes, significantly improving plasticity while reducing forgetting. Additionally, we introduce a novel superclass-level distillation aimed at preserving shared global knowledge within the superclass, further alleviating catastrophic forgetting. We extensively evaluate our method by integrating it into state-of-the-art WISS approaches on Pascal VOC and COCO datasets. We observe consistent gains in performance across diverse experimental scenarios. Code is available athttps://github.com/Magic-Nova77/PGSD. Xuze Hao, Xuhao Jiang, Wenqian Ni, Weimin Tan, Bo Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Motion Distillation Framework for Video Frame InterpolationabstractIn recent years, we have seen the success of deep video enhancement models. However, the performance improvement of new methods has gradually entered a bottleneck period. Optimizing model structures or increasing training data brings less and less improvement. We argue that existing models with advanced structures have not fully demonstrated their performance and demand further exploration. In this study, we statistically analyze the relationship between motion estimation accuracy and video interpolation quality of existing video frame interpolation methods, and find that only supervising the final output leads to inaccurate motion and further affects the interpolation performance. Based on this important observation, we propose a general motion distillation framework that can be widely applied to flow-based and kernel-based video frame interpolation methods. Specifically, we begin by training a teacher model, which uses the ground-truth target frame and adjacent frames to estimate motion. These motion estimates then guide the training of a student model for video frame interpolation. Our experimental results demonstrate the effectiveness of this approach in enhancing performance across diverse advanced video interpolation model structures. For example, after applying our motion distillation framework, the CtxSyn model achieves a PSNR gain of 3.047 dB. Shili Zhou, Weimin Tan, Bo Yan 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Lesion-Decoupling-Based Segmentation With Large-Scale Colon and Esophageal Datasets for Early Cancer DiagnosisabstractLesions of early cancers often show flat, small, and isochromatic characteristics in medical endoscopy images, which are difficult to be captured. By analyzing the differences between the internal and external features of the lesion area, we propose a lesion-decoupling-based segmentation (LDS) network for assisting early cancer diagnosis. We introduce a plug-and-play module called self-sampling similar feature disentangling module (FDM) to obtain accurate lesion boundaries. Then, we propose a feature separation loss (FSL) function to separate pathological features from normal ones. Moreover, since physicians make diagnoses with multimodal data, we propose a multimodal cooperative segmentation network with two different modal images as input: white-light images (WLIs) and narrowband images (NBIs). Our FDM and FSL show a good performance for both single-modal and multimodal segmentations. Extensive experiments on five backbones prove that our FDM and FSL can be easily applied to different backbones for a significant lesion segmentation accuracy improvement, and the maximum increase of mean Intersection over Union (mIoU) is 4.58. For colonoscopy, we can achieve up to mIoU of 91.49 on our Dataset A and 84.41 on the three public datasets. For esophagoscopy, mIoU of 64.32 is best achieved on the WLI dataset and 66.31 on the NBI dataset. Weimin Tan, Shilun Cai, Bo Yan 0001, Yunshi Zhong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Multi-Modality Deep Network for Extreme Learned Image CompressionabstractImage-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2x to 4x bitrates of ours. Xuhao Jiang, Weimin Tan, Tian Tan 0016, Bo Yan 0001, Liquan Shen |
AAAI | 2 |
| 2023 | Fine-Grained Blind Face Inpainting with 3D Face Component DisentanglementabstractInpainting is a task to restore occlusion or other corruption on images. However, previous works require mask of the occluded area to restore the occluded image, which is inconvenient for application. Blind face inpainting aims to automatically restore the occluded face without position information of the corrupt region. In this paper, we propose a novel fine-grained blind face inpainting framework, combining 3D face components disentanglement with generative network. Canonical face texture and shape disentangled by unsupervised 3D face model is restored separately to get occlusion-free rendered result. Finally, the pixel-to-pixel generative module utilize the occlusion image and the coarse de-occlusion face to get refined inpainted result. We also build up a new dataset called CelebO-3D which consists of occluded face images synthesized with 3D occlusion and rendered by 3D face model. Extensive experiments show that the proposed method is effective and robust in face blind inpainting both in synthesized and real images. Extensive evaluations and comparison with previous methods also show our superior effectiveness and lightweight architecture. Ruian He, Weimin Tan, Bo Yan 0001, Yangle Lin |
ICASSP | 3 |
| 2023 | Uncer2Natural: Uncertainty-Aware Unsupervised Image DenoisingabstractRecently, unsupervised image denoising methods learning from paired noisy samples have received increasing attention. These methods build on the idea that the mean of multiple noisy images of the same scene is the ideal clean image. However, these methods ignore the effect of Aleatoric uncertainty in the noisy image (e.g., pixels deviating from the expected distribution). The presence of Aleatoric uncertainty causes degradation of the reconstructed target pixels, resulting in high uncertainty for these pixels (i.e., low confidence), which in turn leads to sub-optimal denoising results. To address this problem, we propose a novel uncertainty-aware unsupervised image denoising method named Uncer2Natural (U2N). It dynamically predicts the Aleatoric uncertainty for each noisy sample and produces satisfactory denoising results by reducing the effect of Aleatoric uncertainty. Extensive experimental results show that U2N outperforms state-of-the- art unsupervised image denoising methods in terms of both quantitative metrics and qualitative visual quality. Weimin Tan, Jiaxing Shi, Bo Yan 0001 |
ICASSP | 2 |
| 2023 | Multi-Modality Deep Network for JPEG Artifacts ReductionabstractIn recent years, many convolutional neural network-based models are designed for JPEG artifacts reduction, and have achieved notable progress. However, few methods are suitable for extreme low-bitrate image compression artifacts reduction. The main challenge is that the highly compressed image loses too much information, resulting in reconstructing high-quality image difficultly. To address this issue, we propose a multimodal fusion learning method for text-guided JPEG artifacts reduction, in which the corresponding text description not only provides the potential prior information of the highly compressed image, but also serves as supplementary information to assist in image deblocking. We fuse image features and text semantic features from the global and local perspectives respectively, and design a contrastive loss built upon contrastive learning to produce visually pleasing results. Extensive experiments, including a user study, prove that our method can obtain better deblocking results compared to the state-of-the-art methods. Xuhao Jiang, Weimin Tan, Chenxi Ma, Bo Yan 0001, Liquan Shen |
IJCAI | 2 |
| 2023 | Learning Survival Distribution with Implicit Survival FunctionabstractSurvival analysis aims at modeling the relationship between covariates and event occurrence with some untracked (censored) samples. In implementation, existing methods model the survival distribution with strong assumptions or in a discrete time space for likelihood estimation with censorship, which leads to weak generalization. In this paper, we propose Implicit Survival Function (ISF) based on Implicit Neural Representation for survival distribution estimation without strong assumptions, and employ numerical integration to approximate the cumulative distribution function for prediction and optimization. Experimental results show that ISF outperforms the state-of-the-art methods in three public datasets and has robustness to the hyperparameter controlling estimation precision. Yu Ling, Weimin Tan, Bo Yan 0001 |
IJCAI | 2 |
| 2023 | Uncertainty-Guided Spatial Pruning Architecture for Efficient Frame InterpolationabstractThe video frame interpolation (VFI) model applies the convolution operation to all locations, leading to redundant computations in regions with easy motion. We can use dynamic spatial pruning method to skip redundant computation, but this method cannot properly identify easy regions in VFI tasks without supervision. In this paper, we develop an Uncertainty-Guided Spatial Pruning (UGSP) architecture to skip redundant computation for efficient frame interpolation dynamically. Specifically, pixels with low uncertainty indicate easy regions, where the calculation can be reduced without bringing undesirable visual results. Therefore, we utilize uncertainty-generated mask labels to guide our UGSP in properly locating the easy region. Furthermore, we propose a self-contrast training strategy that leverages an auxiliary non-pruning branch to improve the performance of our UGSP. Extensive experiments show that UGSP maintains performance but reduces FLOPs by 34%/52%/30% compared to baseline without pruning on Vimeo90K/UCF101/MiddleBury datasets. In addition, our method achieves state-of-the-art performance with lower FLOPs on multiple benchmarks. Ri Cheng, Xuhao Jiang, Ruian He, Shili Zhou, Weimin Tan, Bo Yan 0001 |
ACM Multimedia | 5 |
| 2023 | MVFlow: Deep Optical Flow Estimation of Compressed Videos with Motion Vector PriorabstractIn recent years, many deep learning-based methods have been proposed to tackle the problem of optical flow estimation and achieved promising results. However, they hardly consider that most videos are compressed and thus ignore the pre-computed information in compressed video streams. Motion vectors, one of the compression information, record the motion of the video frames. They can be directly extracted from the compression code stream without computational cost and serve as a solid prior for optical flow estimation. Therefore, we propose an optical flow model, MVFlow, which uses motion vectors to improve the speed and accuracy of optical flow estimation for compressed videos. In detail, MVFlow includes a key Motion-Vector Converting Module, which ensures that the motion vectors can be transformed into the same domain of optical flow and then be utilized fully by the flow estimation module. Meanwhile, we construct four optical flow datasets for compressed videos containing frames and motion vectors in pairs. The experimental results demonstrate the superiority of our proposed MVFlow, which can reduce the AEPE by 1.09 compared to existing models or save 52% time to achieve similar accuracy to existing models. Shili Zhou, Xuhao Jiang, Weimin Tan, Ruian He, Bo Yan 0001 |
ACM Multimedia | 3 |
| 2023 | Self-Supervised Digital Histopathology Image Disentanglement for Arbitrary Domain Stain TransferabstractDiagnosis of cancerous diseases relies on digital histopathology images from stained slides. However, the staining varies among medical centers, which leads to a domain gap of staining. Existing generative adversarial network (GAN) based stain transfer methods highly rely on distinct domains of source and target, and cannot handle unseen domains. To overcome these obstacles, we propose a self-supervised disentanglement network (SDN) for domain-independent optimization and arbitrary domain stain transfer. SDN decomposes an image into features of content and stain. By exchanging the stain features, the staining style of an image is transferred to the target domain. For optimization, we propose a novel self-supervised learning policy based on the consistency of stain and content among augmentations from one instance. Therefore, the process of training SDN is independent on the domain of training data, and thus SDN is able to tackle unseen domains. Exhaustive experiments demonstrate that SDN achieves the top performance in intra-dataset and cross-dataset stain transfer compared with the state-of-the-art stain transfer models, while the number of parameters in SDN is three orders of magnitude smaller parameters than that of compared models. Through stain transfer, SDN improves AUC of downstream classification model on unseen data without fine-tuning. Therefore, the proposed disentanglement framework and self-supervised learning policy have significant advantages in eliminating the stain gap among multi-center histopathology images. Yu Ling, Weimin Tan, Bo Yan 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Rethinking and Improving Few-Shot Segmentation From a Contour-Aware PerspectiveabstractExisting few-shot segmentation approaches basically adopt the idea of comparing the semantic prototype vector of the query image and support images, and then obtaining the segmentation result. However, recent studies have shown that a single feature vector in feature map cannot accurately represent pixel-level categories, thus leading to poor segmentation of object boundary and semantic ambiguity. To address this common problem, we propose a novel contour-aware network (CTANet) for few-shot segmentation in this paper. Unlike the usual practice of classifying each pixel separately, CTANet regards all pixels within the same contour as a whole, which can take advantage of the internal consistency of objects to obtain a more accurate representation of category information. To obtain the accurate object contour, our network consists of a contour generation module and a contour refinement module, where the former exploits multiple levels of features to generate a primary contour map and the latter learns to refine the primary contour map. Furthermore, a novel contour-aware mixed loss is proposed to fuse the common BCE loss and our contour-aware loss to supervise the training process on two levels, pixel-level and contour-level. Extensive experiments demonstrate that our CTANet achieves a new state-of-the-art performance on$ \text{PASCAL-5}^{i}$and$ \text{COCO-20}^{i}$. Hopefully, our new perspective could provide more clues for future research on few-shot segmentation. Our code is freely available at:https://github.com/hardtogetA/CTANet. Weimin Tan, Ganghui Ru, Yueming Jiang, Bo Yan 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Promoting Single-Modal Optical Flow Network for Diverse Cross-Modal Flow EstimationabstractIn recent years, optical flow methods develop rapidly, achieving unprecedented high performance. Most of the methods only consider single-modal optical flow under the well-known brightness-constancy assumption. However, in many application systems, images of different modalities need to be aligned, which demands to estimate cross-modal flow between the cross-modal image pairs. A lot of cross-modal matching methods are designed for some specific cross-modal scenarios. We argue that the prior knowledge of the advanced optical flow models can be transferred to the cross-modal flow estimation, which may be a simple but unified solution for diverse cross-modal matching tasks. To verify our hypothesis, we design a self-supervised framework to promote the single-modal optical flow networks for diverse corss-modal flow estimation. Moreover, we add a Cross-Modal-Adapter block as a plugin to the state-of-the-art optical flow model RAFT for better performance in cross-modal scenarios. Our proposed Modality Promotion Framework and Cross-Modal Adapter have multiple advantages compared to the existing methods. The experiments demonstrate that our method is effective on multiple datasets of different cross-modal scenarios. Shili Zhou, Weimin Tan, Bo Yan 0001 |
AAAI | 2 |
| 2022 | Learning Robust Image-Based Rendering on Sparse Scene Geometry via Depth CompletionabstractRecent image-based rendering (IBR) methods usually adopt plenty of views to reconstruct dense scene geometry. However, the number of available views is limited in prac-tice. When only few views are provided, the performance of these methods drops off significantly, as the scene geometry becomes sparse as well. Therefore, in this paper, we propose Sparse-IBRNet (SIBRNet) to perform robust IBR on sparse scene geometry by depth completion. The SIBR-Net has two stages, geometry recovery (GR) stage and light blending (LB) stage. Specifically, GR stage takes sparse depth map and RGB as input to predict dense depth map by exploiting the correlation between two modals. As in-accuracy of the complete depth map may cause projection biases in the warping process, LB stage first uses a bias-corrected module (BCM) to rectify deviations, and then ag-gregates modified features from different views to render a novel view. Extensive experimental results demonstrate that our method performs best on sparse scene geometry than re-cent IBR methods, and it can generate better or comparable results as well when the geometric information is dense.1 Shili Zhou, Ri Cheng, Weimin Tan, Bo Yan 0001, Lang Fu |
CVPR | 4 |
| 2022 | Geometry-Aware Reference Synthesis for Multi-View Image Super-ResolutionabstractRecent multi-view multimedia applications struggle between high-resolution (HR) visual experience and storage or bandwidth constraints. Therefore, this paper proposes a Multi-View Image Super-Resolution (MVISR) task. It aims to increase the resolution of multi-view images captured from the same scene. One solution is to apply image or video super-resolution (SR) methods to reconstruct HR results from the low-resolution (LR) input view. However, these methods cannot handle large-angle transformations between views and leverage information in all multi-view images. To address these problems, we propose the MVSRnet, which uses geometry information to extract sharp details from all LR multi-view to support the SR of the LR input view. Specifically, the proposed Geometry-Aware Reference Synthesis module in MVSRnet uses geometry information and all multi-view LR images to synthesize pixel-aligned HR reference images. Then, the proposed Dynamic High-Frequency Search network fully exploits the high-frequency textural details in reference images for SR. Extensive experiments on several benchmarks show that our method significantly improves over the state-of-the-art approaches. Ri Cheng, Bo Yan 0001, Weimin Tan, Chenxi Ma |
ACM Multimedia | 4 |
| 2022 | Learning Parallax Transformer Network for Stereo Image JPEG Artifacts RemovalabstractUnder stereo settings, the performance of image JPEG artifacts removal can be further improved by exploiting the additional information provided by a second view. However, incorporating this information for stereo image JPEG artifacts removal is a huge challenge, since the existing compression artifacts make pixel-level view alignment difficult. In this paper, we propose a novel parallax transformer network (PTNet) to integrate the information from stereo image pairs for stereo image JPEG artifacts removal. Specifically, a well-designed symmetric bi-directional parallax transformer module is proposed to match features with similar textures between different views instead of pixel-level view alignment. Due to the issues of occlusions and boundaries, a confidence-based cross-view fusion module is proposed to achieve better feature fusion for both views, where the cross-view features are weighted with confidence maps. Especially, we adopt a coarse-to-fine design for the cross-view interaction, leading to better performance. Comprehensive experimental results demonstrate that our PTNet can effectively remove compression artifacts and achieves superior performance than other testing state-of-the-art methods. Xuhao Jiang, Weimin Tan, Ri Cheng, Shili Zhou, Bo Yan 0001 |
ACM Multimedia | 2 |
| 2022 | Rethinking Super-Resolution as Text-Guided Details GenerationabstractDeep neural networks have greatly promoted the performance of single image super-resolution (SISR). Conventional methods still resort to restoring the single high-resolution (HR) solution only based on the input of image modality. However, the image-level information is insufficient to predict adequate details and photo-realistic visual quality facing large upscaling factors (×8, ×16). In this paper, we propose a new perspective that regards the SISR as a semantic image detail enhancement problem to generate semantically reasonable HR image that are faithful to the ground truth. To enhance the semantic accuracy and the visual quality of the reconstructed image, we explore the multi-modal fusion learning in SISR by proposing a Text-Guided Super-Resolution (TGSR) framework, which can effectively utilize the information from the text and image modalities. Different from existing methods, the proposed TGSR could generate HR image details that match the text descriptions through a coarse-to-fine process. Extensive experiments and ablation studies demonstrate the effect of the TGSR, which exploits the text reference to recover realistic images. Chenxi Ma, Bo Yan 0001, Weimin Tan, Siming Chen 0001 |
ACM Multimedia | 4 |
| 2022 | Co-Completion for Occluded Facial Expression RecognitionabstractThe existence of occlusions brings in semantically irrelevant visual patterns and leads to the content loss of occluded regions. Although previous works have made improvement on occluded facial expression recognition, they do not explicitly handle the interference factors aforementioned. In this paper, we propose an intuitive and simplified workflow, Co-Completion, which combines occlusion discarding and feature completion together to reduce the impact of occlusions on facial expression recognition. To protect key features from being contaminated and reduce the dependency of feature completion on occlusion discarding, guidance from discriminative regions is also introduced for joint feature completion. Moreover, we release the COO-RW database for occlusion simulation and refine the occlusion generation protocol for fair comparison in this filed. Experiments on synthetic and realistic databases demonstrate the superiority of our method. The COO-RW database can be downloaded from https://github.com/loveSmallOrange/COO-RW. Weimin Tan, Ruian He, Yangle Lin, Bo Yan 0001 |
ACM Multimedia | 2 |
| 2022 | Prior embedding multi-degradations super resolution network
Chenxi Ma, Weimin Tan, Bo Yan 0001, Shili Zhou |
Neurocomputing | 2 |
| 2021 | Perceptual Variousness Motion Deblurring with Light Global Context RefinementabstractDeep learning algorithms have made significant progress in dynamic scene deblurring. However, several challenges are still unsettled: 1) The degree and scale of blur in different regions of a blurred image can have a considerable variation in a large range. However, the traditional input pyramid or downscaling-upscaling, is designed to have limited and inflexible perceptual variousness to cope with large blur scale variation. 2) The nonlocal block is proved to be effective in the image enhancement tasks, but it requires high computation and memory cost. In this paper, we are the first to propose a light-weight globally-analyzing module into the image deblurring field, named Light Global Context Refinement (LGCR) module. With exponentially lower cost, it achieves even better performance than the nonlocal unit. Moreover, we propose the Perceptual Variousness Block (PVB) and PVB-piling strategy. By placing PVB repeatedly, the whole method possesses abundant reception field spectrum to be aware of the blur with various degrees and scales. Comprehensive experimental results from the different benchmarks and assessment metrics show that our method achieves excellent performance to set a new state-of-the-art in motion deblurring.1 Weimin Tan, Bo Yan 0001 |
ICCV | 2 |
| 2021 | Multimodal Asymmetric Dual Learning for Unsupervised Eyeglasses RemovalabstractGlasses removal is a challenging task due to the diversity of glasses species and the difficulty of obtaining paired datasets. Most existing methods need to build different models for different glasses or expensive paired datasets for supervised training, which lacks universality. In this paper, we propose a multimodal asymmetric dual learning method for unsupervised glasses removal. This method uses large-scale face images with and without glasses for dual feature learning, which does not require intensive manual marking of the glasses. Given a face image with glasses, we aim to generate a glasses-free image preserving the person identity. Thus, in order to make up for the lack of semantic features in the glasses region, we introduce the text description of the target image into the task, and propose a text-guided multimodal feature fusion method. We adaptively select the glasses-free image closest to the target one for better dual feature learning. We also propose a exchange residual loss to generate more precise mask of glasses. Extensive experiments prove that our method can generate real glasses-free images, and better retain the person identity, which can be useful for face recognition. Bo Yan 0001, Weimin Tan |
ACM Multimedia | 3 |
| 2021 | Perception-Oriented Stereo Image Super-ResolutionabstractRecent studies of deep learning based stereo image super-resolution (StereoSR) have promoted the development of StereoSR. However, existing StereoSR models mainly concentrate on improving quantitative evaluation metrics and neglect the visual quality of super-resolved stereo images. To improve the perceptual performance, this paper proposes the first perception-oriented stereo image super-resolution approach by exploiting the feedback, provided by the evaluation on the perceptual quality of StereoSR results. To provide accurate guidance for the StereoSR model, we develop the first special stereo image super-resolution quality assessment (StereoSRQA) model, and further construct a StereoSRQA database. Extensive experiments demonstrate that our StereoSR approach significantly improves the perceptual quality and enhances the reliability of stereo images for disparity estimation. Chenxi Ma, Bo Yan 0001, Weimin Tan, Xuhao Jiang |
ACM Multimedia | 3 |
| 2021 | Research on hybrid feature selection method of power transformer based on fuzzy information entropy
Weimin Tan |
Adv. Eng. Informatics | 2 |
| 2021 | JROTM: Jointly reinforced object tracking with temporal content reference and motion guidance
Bo Yan 0001, Chuming Lin, Weimin Tan |
Neurocomputing | 4 |
| 2020 | Assessing Eye Aesthetics for Automatic Multi-Reference Eye In-PaintingabstractWith the wide use of artistic images, aesthetic quality assessment has been widely concerned. How to integrate aesthetics into image editing is still a problem worthy of discussion. In this paper, aesthetic assessment is introduced into eye in-painting task for the first time. We construct an eye aesthetic dataset, and train the eye aesthetic assessment network on this basis. Then we propose a novel eye aesthetic and face semantic guided multi-reference eye inpainting GAN approach (AesGAN), which automatically selects the best reference under the guidance of eye aesthetics. A new aesthetic loss has also been introduced into the network to learn the eye aesthetic features and generate highquality eyes. We prove the effectiveness of eye aesthetic assessment in our experiments, which may inspire more applications of aesthetics assessment. Both qualitative and quantitative experimental results show that the proposed AesGAN can produce more natural and visually attractive eyes compared with state-of-the-art methods. Bo Yan 0001, Weimin Tan, Shili Zhou |
CVPR | 3 |
| 2020 | Disparity-Aware Domain Adaptation in Stereo Image RestorationabstractUnder stereo settings, the problems of disparity estimation, stereo magnification and stereo-view synthesis have gathered wide attention. However, the limited image quality brings non-negligible difficulties in developing related applications and becomes the main bottleneck of stereo images. To the best of our knowledge, stereo image restoration is rarely studied. Towards this end, this paper analyses how to effectively explore disparity information, and proposes a unified stereo image restoration framework. The proposed framework explicitly learn the inherent pixel correspondence between stereo views and restores stereo image with the cross-view information at image and feature level. A Feature Modulation Dense Block (FMDB) is introduced to insert disparity prior throughout the whole network. The experiments in terms of efficiency, objective and perceptual quality, and the accuracy of depth estimation demonstrates the superiority of the proposed framework on various stereo image restoration tasks. Bo Yan 0001, Chenxi Ma, Bahetiyaer Bare, Weimin Tan, Steven C. H. Hoi |
CVPR | 4 |
| 2020 | MMFL: Multimodal Fusion Learning for Text-Guided Image InpaintingabstractPainters can successfully recover severely damaged objects, yet current inpainting algorithms still can not achieve this ability. Generally, painters will have a conjecture about the seriously missing image before restoring it, which can be expressed in a text description. This paper imitates the process of painters' conjecture, and proposes to introduce the text description into the image inpainting task for the first time, which provides abundant guidance information for image restoration through the fusion of multimodal features. We propose a multimodal fusion learning method for image inpainting (MMFL). To make better use of text features, we construct an image-adaptive word demand module to reasonably filter the effective text features. We introduce a text guided attention loss and a text-image matching loss to make the network pay more attention to the entities in the text description. Extensive experiments prove that our method can better predict the semantics of objects in the missing regions and generate fine grained textures. Bo Yan 0001, Weimin Tan |
ACM Multimedia | 4 |
| 2020 | Effective image restoration for semantic segmentation
Xuejing Niu, Bo Yan 0001, Weimin Tan |
Neurocomputing | 3 |
| 2020 | Cycle-IR: Deep Cyclic Image RetargetingabstractSupervised deep learning techniques have achieved great success in various fields due to getting rid of the limitation of handcrafted representations. However, most previous image retargeting algorithms still employ fixed design principles such as using gradient map or handcrafted features to compute saliency map, which inevitably restricts its generality. Deep learning techniques may help to address this issue, but the challenging problem is that we need to build a large-scale image retargeting dataset for the training of deep retargeting models. However, building such a dataset requires huge human efforts. In this paper, we propose a novel deep cyclic image retargeting approach, called Cycle-IR, to firstly implement image retargeting with a single deep model, without relying on any explicit user annotations. Our idea is built on the reverse mapping from the retargeted images to the given images. If the retargeted image has serious distortion or excessive loss of important visual information, the reverse mapping is unlikely to restore the input image well. We constrain this forward-reverse consistency by introducing a cyclic perception coherence loss. In addition, we propose a simple yet effective image retargeting network (IRNet) to implement the image retargeting process. Our IRNet contains a spatial and channel attention layer, which is able to discriminate visually important regions of input images effectively, especially in cluttered images. Given arbitrary sizes of input images and desired aspect ratios, our Cycle-IR can produce visually pleasing target images directly. Extensive experiments on the standard RetargetMe dataset show the superiority of our Cycle-IR. Weimin Tan, Bo Yan 0001, Chuming Lin, Xuejing Niu |
IEEE Trans. Multim. | 1 |
| 2020 | Semantic Segmentation Guided Pixel Fusion for Image RetargetingabstractImage retargeting aims to obtain high visual quality of target images for human vision. Through semantic segmentation and understanding of input images, we can better preserve the important semantic regions, so as to effectively improve the performance of image retargeting. Benefit from the successful application of deep neural network in the field of semantic segmentation, in this paper, we propose a novel image retargeting approach using semantic segmentation and pixel fusion. Compared with existing image retargeting methods, our approach can effectively reduce geometric distortion during image retargeting by finely reallocating scaling factors for each region based on the semantic segmentation results. Experimental results demonstrate that the proposed approach can well preserve important semantic regions while leaving less unnatural geometric distortion. Our approach also shows the important role of semantic segmentation and understanding of scenes in image retargeting in detail. Bo Yan 0001, Xuejing Niu, Bahetiyaer Bare, Weimin Tan |
IEEE Trans. Multim. | 4 |
| 2019 | Frame and Feature-Context Video Super-ResolutionabstractFor video super-resolution, current state-of-the-art approaches either process multiple low-resolution (LR) frames to produce each output high-resolution (HR) frame separately in a sliding window fashion or recurrently exploit the previously estimated HR frames to super-resolve the following frame. The main weaknesses of these approaches are: 1) separately generating each output frame may obtain high-quality HR estimates while resulting in unsatisfactory flickering artifacts, and 2) combining previously generated HR frames can produce temporally consistent results in the case of short information flow, but it will cause significant jitter and jagged artifacts because the previous super-resolving errors are constantly accumulated to the subsequent frames.In this paper, we propose a fully end-to-end trainable frame and feature-context video super-resolution (FFCVSR) network that consists of two key sub-networks: local network and context network, where the first one explicitly utilizes a sequence of consecutive LR frames to generate local feature and local SR frame, and the other combines the outputs of local network and the previously estimated HR frames and features to super-resolve the subsequent frame. Our approach takes full advantage of the inter-frame information from multiple LR frames and the context information from previously predicted HR frames, producing temporally consistent highquality results while maintaining real-time speed by directly reusing previous features and frames. Extensive evaluations and comparisons demonstrate that our approach produces state-of-the-art results on a standard benchmark dataset, with advantages in terms of accuracy, efficiency, and visual quality over the existing approaches. Bo Yan 0001, Chuming Lin, Weimin Tan |
AAAI | 3 |
| 2019 | A Multi-level Aggregated Network for Image RestorationabstractRecently, significant progress has been witnessed in image restoration benefited from the development of deep convolutional neural networks (CNN). However, we note that many state-of-the-art image restoration networks can be unfolded as a one-level architecture, which is constructed by stacking multiple convolution layers or blocks. As the depth of network grows, the information flow is weakened. And, existing skip connection has restricted ability to pass the previous state to the latter layers in the network. Based on above observations, we explore a multi-level aggregated network (MLAN) to fully exploit features of deeper layers. The proposed MLAN can extract more features by merging layers at different levels progressively, and can better aggregate the features by the augmented connection manner. Experimental results demonstrate a satisfactory performance of the proposed model on different image restoration tasks. Chenxi Ma, Weimin Tan, Bahetiyaer Bare, Bo Yan 0001 |
ICME | 2 |
| 2019 | RDGAN: Retinex Decomposition Based Adversarial Learning for Low-Light EnhancementabstractPictures taken under the low-light condition often suffer from low contrast and loss of image details, thus an approach that can effectively improve low-light images is demanded. Traditional Retinex-based methods assume that the reflectance components of low-light images keep unchanged, which neglect the color distortion and lost details. In this paper, we propose an end-to-end learning-based framework that first decomposes the low-light image and then learns to fuse the decomposed results to obtain the high quality enhanced result. Our framework can be divided into a RDNet (Retinex Decomposition Network) for decomposition and a FENet (Fusion Enhancement Network) for fusion. Specific multi-term losses are respectively designed for the two networks. We also present a new RDGAN (Retinex Decomposition based Generative Adversarial Network) loss, which is computed on the decomposed reflectance components of the enhanced and the reference images. Experiments demonstrate that our approach is good at color and detail restoration, which outperforms other state-of-the-art methods. Weimin Tan, Xuejing Niu, Bo Yan 0001 |
ICME | 2 |
| 2019 | Deep Objective Quality Assessment Driven Single Image Super-ResolutionabstractSingle-image super-resolution (SISR) is a classic problem in the image processing community, which aims at generating a high-resolution image from a low-resolution one. In recent years, deep learning based SISR methods emerged and achieved a performance leap than previous methods. However, because the evaluation metrics of SISR methods is peak signal-to-noise ratio (PSNR), previous methods usually choose L2-norm as the loss function. This leads to a significant improvement in the final PSNR value but little improvement in perceptual quality. In this paper, in order to achieve better results in both perceptual quality and PSNR values, we propose an objective quality assessment driven SISR method. First, we propose a novel full-reference image quality assessment approach for SISR and employ it as a loss function, namely super-resolution image quality assessment (SR-IQA) loss. Then, we combine SR-IQA loss with L2-norm to guide our proposed SISR method to achieve better results. Besides that, our proposed SISR method consists of several proposed highway units. Furthermore, in order to verify the generalization ability of our new kind of loss function, we integrate SR-IQA loss to generative adversarial networks based SR method and achieve better perceptual quality. Experimental results prove that our proposed SISR method achieves better performance than other methods both qualitatively and quantitatively in most of the cases. Bo Yan 0001, Bahetiyaer Bare, Chenxi Ma, Ke Li 0010, Weimin Tan |
IEEE Trans. Multim. | 5 |
| 2019 | Naturalness-Aware Deep No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) is a non-trivial task, because it is hard to find a pristine counterpart for an image in real applications, such as image selection, high quality image recommendation, etc. In recent years, deep learning-based NR-IQA methods emerged and achieved better performance than previous methods. In this paper, we present a novel deep neural networks-based multi-task learning approach for NR-IQA. Our proposed network is designed by a multi-task learning manner that consists of two tasks, namely, natural scene statistics (NSS) features prediction task and the quality score prediction task. NSS features prediction is an auxiliary task, which helps the quality score prediction task to learn better mapping between the input image and its quality score. The main contribution of this work is to integrate the NSS features prediction task to the deep learning-based image quality prediction task to improve the representation ability and generalization ability. To the best of our knowledge, it is the first attempt. We conduct the same database validation and cross database validation experiments on LIVE1, TID20132, CSIQ3, LIVE multiply distorted image quality database (LIVE MD)4, CID20135, and LIVE in the wild image quality challenge (LIVE challenge)6databases to verify the superiority and generalization ability of the proposed method. Experimental results confirm the superior performance of our method on the same database validation; our method especially achieves 0.984 and 0.986 on the LIVE image quality assessment database in terms of the Pearson linear correlation coefficient (PLCC) and Spearman rank-order correlation coefficient (SROCC), respectively. Also, experimental results from cross database validation verify the strong generalization ability of our method. Specifically, our method gains significant improvement up to 21.8% on unseen distortion types. Bo Yan 0001, Bahetiyaer Bare, Weimin Tan |
IEEE Trans. Multim. | 3 |
| 2018 | Feature Super-Resolution: Make Machine See More ClearlyabstractIdentifying small size images or small objects is a notoriously challenging problem, as discriminative representations are difficult to learn from the limited information contained in them with poor-quality appearance and unclear object structure. Existing research works usually increase the resolution of low-resolution image in the pixel space in order to provide better visual quality for human viewing. However, the improved performance of such methods is usually limited or even trivial in the case of very small image size (we will show it in this paper explicitly). In this paper, different from image super-resolution (ISR), we propose a novel super-resolution technique called feature super-resolution (FSR), which aims at enhancing the discriminatory power of small size image in order to provide high recognition precision for machine. To achieve this goal, we propose a new Feature Super-Resolution Generative Adversarial Network (FSR-GAN) model that transforms the raw poor features of small size images to highly discriminative ones by performing super-resolution in the feature space. Our FSR-GAN consists of two subnetworks: a feature generator network G and a feature discriminator network D. By training the G and the D networks in an alternative manner, we encourage the G network to discover the latent distribution correlations between small size and large size images and then use G to improve the representations of small images. Extensive experiment results on Oxford5K, Paris, Holidays, and Flick100k datasets demonstrate that the proposed FSR approach can effectively enhance the discriminatory ability of features. Even when the resolution of query images is reduced greatly, e.g., 1/64 original size, the query feature enhanced by our FSR approach achieves surprisingly high retrieval performance at different image resolutions and increases the retrieval precision by 25% compared to the raw query feature. Weimin Tan, Bo Yan 0001, Bahetiyaer Bare |
CVPR | 1 |
| 2018 | Deep Residual Network for Enhancing Quality of the Decoded Intra Frames of HevcabstractHigh Efficiency Video Coding (HEVC) has been widely used to encode video sequences and output streams for its outstanding performance on compression. However, the lossy compression process still results in blocking, blurring and ring effects, which are quite significant at low bit-rates. In this paper, we propose a post-processing method based on a residual convolutional neural network (CNN) to effectively improve the visual quality of the decoded intra frames of HEVC. By learning the spatial mapping between high and low quality image patches, our approach can effectively predict lost image details. Since the proposed approach only works on decoded frames, it does not require to modify the original HEVC, i.e., the enhancement can be done in parallel with HEVC baseline, which makes it flexible to be included into a decoding system. Experimental results show that the proposed approach outperforms the state-of-the-arts by a large margin and makes a 5.5% bit-rate reduction on average compared to HEVC baseline. Furthermore, the PSNR measures of intra frames increase about 0.395 dB when encoded at high quality parameters (QP). Weimin Tan, Bo Yan 0001 |
ICIP | 2 |
| 2018 | Foreground Detection in Surveillance Video with Fully Convolutional Semantic NetworkabstractForeground detection is an important part of surveillance video analysis, and also has challenges. For example, the classical methods are difficult to distinguish the foreground, which is similar to the background. In recent years, Convolutional Neural Networks (CNNs) have been widely used in image processing and achieved better performance. In this paper, we proposed an efficient deep Fully Convolutional Semantic Networks (FCSN) model for foreground detection in surveillance video. Our model aimed at learning the global differences between the video frame and the background image, and the semantic information by utilizing the pre-trained weights on semantic segmentation. In the experiment, unlike other related work, we proposed a reasonable method, which is able to avoid overfitting results to construct training data with 20 videos and test data with 6 videos on the dataset of 2014 ChangeDetection.net (CDnet 2014). Experimental results verified that our model outperforms the state-of-the-art methods in the foreground detection of surveillance video. Chuming Lin, Bo Yan 0001, Weimin Tan |
ICIP | 3 |
| 2018 | Beyond Visual Retargeting: A Feature Retargeting Approach for Visual Recognition and Its ApplicationsabstractThe popularity of mobile applications has greatly enriched and facilitated our lives. However, the rapid increase of digital images and the problem of narrow bandwidth of the wireless network call for an appropriate approach to reduce the amount of data transmitted over the wireless network (i.e., low bit-rate transmission) while ensuring high recognition accuracy at the cloud. We propose a simple and effective feature retargeting (FR) approach for retargeting an image while preserving the representative local features (e.g., SIFT, SURF, and BRIEF) in the image. Our feature retargeting approach aims at low bit-rate visual recognition instead of high-quality visual perception that visual retargeting methods dedicate to. Our algorithm consists of two key novelties: estimating feature saliency and retargeting image: Estimating feature saliency focuses on predicting the relative importance of different features in an image by analyzing uniqueness in a specific context; Retargeting image aims at finding the optimal resolution for the retargeted image to maximize feature-saliency energy. We evaluate the proposed approach for two different applications in three large data sets and observe that our FR approach consistently outperforms state-of-the-art retargeting algorithms, resulting in both higher precision and lower bit-rates. We also demonstrate that even when the resolution of source image is reduced greatly, e.g., 1/7 original size, our algorithm produces superior results as compared with other approaches. Weimin Tan, Bo Yan 0001, Chuming Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Salient Object Detection via Google Image Retrieval
Weimin Tan, Bo Yan 0001 |
ICIG (1) | 1 |
| 2017 | Salient object detection via multiple saliency weights
Weimin Tan, Bo Yan 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Codebook Guided Feature-Preserving for Recognition-Oriented Image RetargetingabstractTraditional image resizing methods, such as uniform scaling and content-aware image retargeting, are designed to preserve the visually salient contents of an image while resizing it. In this paper, we propose a novel image resizing approach called recognition-oriented image retargeting. Its goal is to preserve the distinctive local features for recognition instead of the traditional visual saliency during resizing. Moreover, we also apply our approach to image matching and image retrieval applications to verify its performance. Meanwhile, using our approach to these applications is able to solve some of the challenging problems in their fields. In image matching application, we find that our approach shows promising preservation of local feature descriptors. In image retrieval task, extensive experiments on Oxford5K, Holidays, Paris, and Flickr100k data sets demonstrate that our approach consistently outperforms other image retargeting methods by large margins in the aspects of retrieval precision and query bits. Bo Yan 0001, Weimin Tan, Ke Li 0010, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | A survey on high coherence visual media retargeting: recent advances and applications
Weimin Tan, Bo Yan 0001 |
Frontiers Comput. Sci. | 1 |
| 2016 | Image Retargeting for Preserving Robust Local Feature: Application to Mobile Visual SearchabstractWith the sharp increasing of mobile devices, conducting search on mobile devices becomes pervasive, and one of the most popular applications is mobile visual search. To achieve low bit-rate visual search, most of the existing works focus on addressing local descriptor coding and BoW histogram compression . In this paper, we extend the concept of image retargeting and propose a new image resizing approach that is devoted to preserving the robust local features in the query image while resizing it. Based on the extended concept, we introduce a novel mobile-visual-search scheme that conducts the proposed approach to reduce the size of the query image for achieving low bit-rate visual search. Extensive experiments on Oxford 5 K and Flickr 100k datasets show that our approach obtains superior retrieval performance than state-of-the-art image resizing approaches at the similar query size; meanwhile, it is cost effective in terms of processing time. Weimin Tan, Bo Yan 0001, Ke Li 0010, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |