VLDB 2026 Research / reviewers in the wild / expert
Shifeng Chen
dblp:84/4529
· DBLP profile ↗
71ranked-venue papers
6as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 33 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsabstractPersonalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation (LoRA) enable efficient single-concept customization by injecting lightweight, concept-specific adapters into pre-trained diffusion models. However, combining multiple LoRA modules for multi-concept generation often leads to identity missing and visual feature leakage. In this work, we identify two key issues behind these failures: (1) token-wise interference among different LoRA modules, and (2) spatial misalignment between the attention map of a rare token and its corresponding concept-specific region. To address these issues, we propose Token-Aware LoRA (TARA), which introduces a token mask to explicitly constrain each module to focus on its associated rare token to avoid interference, and a training objective that encourages the spatial attention of a rare token to align with its concept region. Our method enables training-free multi-concept composition by directly injecting multiple independently trained TARA modules at inference time. Experimental results demonstrate that TARA enables efficient multi-concept inference and effectively preserving the visual identity of each concept by avoiding mutual interference between LoRA modules. Yuqi Peng, Lingtao Zheng, Yi Huang 0035, Mingfu Yan, Jianzhuang Liu, Shifeng Chen |
AAAI | 7 |
| 2026 | Revisiting color-event based tracking: A unified network, dataset, and metric
Chuanming Tang, Xiao Wang 0014, Ju Huang, Bo Jiang 0002, Lin Zhu 0012, Shifeng Chen, Jianlin Zhang 0001, Yaowei Wang 0001, Yonghong Tian 0001 |
Pattern Recognit. | 6 |
| 2025 | Efficient Document Shadow Removal with Contrast-Aware Guidance
Yifan Liu 0001, Jiyu Wu, Jiancheng Huang, Mingfu Yan, Yi Huang 0035, Shifeng Chen |
CGI (3) | 7 |
| 2025 | IDOL: Instant Photorealistic 3D Human Creation from a Single ImageabstractCreating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the perspectives of dataset, model, and representation. First, we introduce a large-scale HUman-centric GEnerated dataset, HuGe100K, consisting of 100K diverse, photorealistic sets of human images. Each set contains 24-view frames in specific human poses, generated using a pose-controllable image-to-multi-view model. Next, leveraging the diversity in views, poses, and appearances within HuGe100K, we develop a scalable feed-forward transformer model to predict a 3D human Gaussian representation in a uniform space from a given human image. This model is trained to disentangle human pose, body shape, clothing geometry, and texture. The estimated Gaussians can be animated without post-processing. We conduct comprehensive experiments to validate the effectiveness of the proposed dataset and method. Our model demonstrates the ability to efficiently reconstruct photorealistic humans at 1K resolution from a single input image using a single GPU instantly. Additionally, it seamlessly supports various applications, as well as shape and texture editing tasks. Yiyu Zhuang, Jiaxi Lv, Hao Wen 0005, Qing Shuai, Ailing Zeng, Hao Zhu 0004, Shifeng Chen, Yujiu Yang 0001, Xun Cao, Wei Liu 0005 |
CVPR | 7 |
| 2025 | DIVE: Taming DINO for Subject-Driven Video EditingabstractBuilding on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these issues, this paper proposes DINO-guided Video Editing (DIVE), a framework designed to facilitate subject-driven editing in source videos conditioned on either target text prompts or reference images with specific identities. The core of DIVE lies in leveraging the powerful semantic features extracted from a pretrained DINOv2 model as implicit correspondences to guide the editing process. Specifically, to ensure temporal motion consistency, DIVE employs DINO features to align with the motion trajectory of the source video. For precise subject editing, DIVE incorporates the DINO features of reference images into a pretrained text-to-image model to learn Low-Rank Adaptations (LoRAs), effectively registering the target subject's identity. Extensive experiments on diverse real-world videos demonstrate that our framework can achieve high-quality editing results with robust motion consistency, highlighting the potential of DINO to contribute to video editing. Project page: https://dino-video-editing.github.io Yi Huang 0035, Wei Xiong 0008, He Zhang 0004, Chaoqi Chen, Jianzhuang Liu, Mingfu Yan, Shifeng Chen |
ICCV | 7 |
| 2025 | WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructabstractLarge language models (LLMs), such as GPT-4, have shown remarkable performance in natural language processing (NLP) tasks, including challenging mathematical reasoning. However, most existing open-source models are only pre-trained on large-scale internet data and without math-related optimization. In this paper, we present WizardMath, which enhances the mathematical reasoning abilities of LLMs, by applying our proposed Reinforcement Learning from Evol-Instruct Feedback (RLEIF) method to the domain of math. Through extensive experiments on two mathematical reasoning benchmarks, namely GSM8k and MATH, we reveal the extraordinary capabilities of our model. Remarkably, WizardMath-Mistral 7B surpasses all other open-source LLMs by a substantial margin. Furthermore, WizardMath 70B even outperforms ChatGPT-3.5, Claude Instant, Gemini Pro and Mistral Medium. Additionally, our preliminary exploration highlights the pivotal role of instruction evolution and process supervision in achieving exceptional math performance. Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Jian-Guang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, Yansong Tang, Dongmei Zhang 0001 |
ICLR | 9 |
| 2025 | Component Adaptive Clustering for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) tackles the challenging problem of categorizing unlabeled images into both known and novel classes within a partially labeled dataset, without prior knowledge of the number of unknown categories. Traditional methods often rely on rigid assumptions, such as predefining the number of classes, which limits their ability to handle the inherent variability and complexity of real-world data. To address these shortcomings, we propose AdaGCD, a cluster-centric contrastive learning framework that incorporates Adaptive Slot Attention (AdaSlot) into the GCD framework. AdaSlot dynamically determines the optimal number of slots based on data complexity, removing the need for predefined slot counts. This adaptive mechanism facilitates the flexible clustering of unlabeled data into known and novel categories by dynamically allocating representational capacity. By integrating adaptive representation with dynamic slot allocation, our method captures both instance-specific and spatially clustered features, improving class discovery in open-world scenarios. Extensive experiments on public and fine-grained datasets validate the effectiveness of our framework, emphasizing the advantages of leveraging spatial local information for category discovery in unlabeled image datasets. Mingfu Yan, Jiancheng Huang, Yifan Liu 0001, Shifeng Chen |
ICME | 4 |
| 2025 | Enhancing Image-Text Retrieval with Phrase-aware and Modality Difference-aware EmbeddingsabstractImage-to-text retrieval aims to search for relevant images given a textual description or vice versa, which is a crucial task in multi-modal intelligence. The core idea is to learn visual and textual embeddings to ensure the similarity of matched image-text pairs. Its challenges originate from grasping the intricate relationships between various parts of images and texts, and the difficulty of understanding modality differences. To alleviate these problems, we propose a novel phrase-aware and modality difference-aware embedding learning network (PMD-Net) by jointly using an optimal transport-based phrase-aware module and a modality difference-aware module in a unified model. To the best of our knowledge, this is the first work that uses an optimal transport algorithm to learn phrase-aware features with varied structures to effectively represent relationships between multi-modality data. Additionally, the proposed modality difference-aware module leverages the variations between textual and visual features to bridge the semantic gap effectively. Extensive experimental results with two backbones demonstrate that our PMD-Net performs favorably against state-of-the-art methods on two standard datasets. Shifeng Chen |
IJCNN | 4 |
| 2025 | Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image EditingabstractText-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity. Jiancheng Huang, Yi Huang 0035, Jianzhuang Liu, Yifan Liu 0001, Shifeng Chen |
WACV | 6 |
| 2025 | Diffusion Model-Based Image Editing: A SurveyabstractDenoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research. Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2025 | DCD-Net: Weakly supervised decomposition learning for real-world image dehazing
Yi Huang 0035, Jiancheng Huang, Mingfu Yan, Shifeng Chen |
Signal Process. | 5 |
| 2025 | Edge Cascading Failure Model for Command and Control Networks: Integrating Edge Importance and Edge HierarchyabstractMost existing cascading failure models for command and control networks (C2Ns) primarily focus on node-level failures while overlooking edge-level failures. To address this issue, this paper proposes a novel edge-based cascading failure model that integrates edge importance and hierarchical structure. First, both local and global edge importance are quantified through edge degree and ego network edge betweenness, and an improved definition of edge importance is introduced by incorporating an enhanced bridging coefficient. Second, based on the hierarchical structure of network nodes, edge hierarchy is defined, and a new initial edge load formulation is proposed by combining edge hierarchy and importance. Third, to reflect load imbalance in realistic networks, a nonlinear edge capacity model is constructed, and a non-uniform, adjustable load redistribution strategy based on the residual capacity of neighboring edges is developed to accommodate network hierarchy. Simulation experiments are conducted to validate the proposed model. The results show that, in two representative C2N networks, the degradation rates of average efficiency and connectivity coefficient are 22.42% and 73.70% for Network 1, and 13.84% and 14.52% for Network 2, respectively. These results demonstrate that the proposed model significantly improves the robustness and resilience of C2Ns under cascading failure scenarios. Bo Chen 0007, Lingdong Sun, Yufeng Chen 0008, Xiu-e Gao, Zhengtao Xiang, Shifeng Chen |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | A Dual-Level Bio-Inspired Optimization Algorithm for Cloud Manufacturing Service Evaluation on Industrial Internet of Things (IIoT) PlatformsabstractInspired by the immune-endocrine system, an improved biological comprehensive optimization algorithm (IBCOA) is proposed for industrial big data analysis and cloud manufacturing service matching. IBCOA employs a dual-level strategy: a bottom-level global optimization immune algorithm (GOIA) narrows down the search space to optimize short-term parameters, while a top-level fuzzy weighted comprehensive evaluation (FWCE) refines the solutions by incorporating long-term performance metrics. Experimental results demonstrate IBCOA’s superior performance, showing higher accuracy, recall, and F1 scores compared to least squares, decision trees, andK-means clustering, along with longer execution time and lower error rates. When tested on standard benchmarks including Iris (classification), MNIST (handwritten digits), and CIFAR-10 (image recognition), IBCOA achieves remarkable accuracies, highlighting its strong generalization and adaptability. The algorithm not only addresses immediate production requirements in polyester fiber industrial data analysis but also enhances long-term operational efficiency and product quality. By balancing stakeholder interests (suppliers, consumers, operators), it promotes sustainable development on industrial internet platforms. This work provides a robust solution for the industrial Internet of Things (IIoT) service evaluation and classification tasks, demonstrating transformative potential for cloud manufacturing resource allocation across diverse applications. Chunli Jiang, Kuangrong Hao, Witold Pedrycz, Haoliang Zhu, Shifeng Chen |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | LeSkill: Structured Skill Learning for Long-Horizon Robotic Manipulation TasksabstractIn long-horizon tasks, leveraging prior knowledge to streamline task execution is essential. However, navigating complex environments to achieve long-term objectives poses a significant challenge due to the vast exploration space involved. To address this issue, we propose a skill-based hierarchical reinforcement learning (RL) framework, termed LeSkill. This framework utilizes a conditional generative model to pretrain a comprehensive and generalizable skill repository from heterogeneous datasets, facilitating skill inference across diverse contexts. This strategy enhances transferability to novel tasks, thereby minimizing the need for extensive, task-specific training. Subsequently, a concise set of task-specific demonstrations is employed to guide the selection process, allowing the model to efficiently sample relevant skills from the pre-existing skill repository, which effectively reduces the exploration space. This approach accelerates the acquisition of highly effective policies tailored for task completion. Our framework undergoes rigorous evaluation on two challenging long horizon, multistep tasks: a standard task and a distribution mismatch task. The results highlight the framework’s superior performance in mastering intricate tasks and its remarkable generalization capabilities. Xiucai Huang, Shifeng Chen, Yongduan Song 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | MambaDW: Semantic-Aware Mamba for Document Watermark Removal
Yifan Liu 0001, Mingfu Yan, He Hua, Jiancheng Huang, Shifeng Chen |
CGI (1) | 5 |
| 2024 | BK-Editer: Body-Keeping Text-Conditioned Real Image Editing
Jiancheng Huang, Linxiao Shi, Shifeng Chen |
CVM (1) | 5 |
| 2024 | Entwined Inversion: Tune-Free Inversion For Real Image Faithful Reconstruction and EditingabstractText-conditional image editing is a very practical AIGC task that has recently emerged with great commercial and academic research value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing, but DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for all downstream edits. In order to solve this problem, we first mathematically analyze the reason for the reconstruction failure of DDIM Inversion, and then propose a new inversion and sampling method named Entwined Inversion that can achieve satisfactory reconstruction and editing performance, which can solve two major problems: 1) the object can retain the main content of the original image; 2) the edited object can conform to the semantics of the text prompt. In addition, our method does not require training the diffusion model itself on a large dataset, nor does it require any fine-tuning for some particular images. Jiancheng Huang, Yifan Liu 0001, Jiaxi Lv, Shifeng Chen |
ICASSP | 4 |
| 2024 | Color-SD: Stable Diffusion Model Already has a Color Style Noisy Latent SpaceabstractWe present Color-SD, a comprehensive color style transfer framework that utilizes either image or text references. Built on the pretrained Stable Diffusion Model, Color-SD exploits an existing color style space, enabling a training-free and tuning-free zero-shot color style transfer method without introducing new parameters. For image references, we first invert the source and reference images to the noisy latent space, followed by parallel sampling. During this process, we execute distribution transformation in the noisy latent space, effectively completing the color style transfer and generating the stylized result. For text references, we capitalize on the Stable Diffusion model’s inherent text-to-image capability. We only invert the source image to the noisy latent, and the given text reference prompt is utilized during the parallel sampling. This approach eliminates the need for training or tuning, yet produces impressive open-set transfer results. Comprehensive experiments validate the effectiveness of our method, demonstrating significant superiority over existing methods in both qualitative and quantitative evaluations. Jiancheng Huang, Mingfu Yan, Shifeng Chen |
ICME | 4 |
| 2024 | Feature Augmentation for Self-supervised Contrastive Learning: A Closer LookabstractSelf-supervised contrastive learning heavily relies on the view variance brought by data augmentation, so that it can learn a view-invariant pre-trained representation. Beyond increasing the view variance for contrast, this work focuses on improving the diversity of training data, to improve the generalization and robustness of the pre-trained models. To this end, we propose a unified framework to conduct data augmentation in the feature space, known as feature augmentation. This strategy is domain-agnostic, which augments similar features to the original ones and thus improves the data diversity. We perform a systematic investigation of various feature augmentation architectures, the gradient-flow skill, and the relationship between feature augmentation and traditional data augmentation. Our study reveals some practical principles for feature augmentation in self-contrastive learning. By integrating feature augmentation on the instance discrimination or the instance similarity paradigm, we consistently improve the performance of pre-trained feature learning and gain better generalization over the downstream image classification and object detection task. Shifeng Chen |
IJCNN | 5 |
| 2024 | SBCR: Stochasticity Beats Content Restriction Problem in Training and Tuning Free Image EditingabstractText-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as a first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. Many inversion-based works modify the formula to address this problem but this leads to another content restriction problem. To solve the content restriction problem, we first analyze why the reconstruction via DDIM Inversion fails and then propose Reconstruction-and-Generation Balancing Noises (R&G-B noises) that can achieve superior reconstruction and editing performance with the following advantages: 1) It can perfectly reconstruct real images without fine-tuning. 2) It can overcome the content restriction problem and generate diverse content. Jiancheng Huang, Mingfu Yan, Shifeng Chen |
ICMR | 4 |
| 2024 | MagicFight: Personalized Martial Arts Combat Video GenerationabstractAmid the surge in generic text-to-video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single-person scenarios. However, to our knowledge, the domain of two-person interactions, particularly in the context of martial arts combat, remains uncharted. We identify a significant gap: existing models for single-person dancing generation prove insufficient for capturing the subtleties and complexities of two engaged fighters, resulting in challenges such as identity confusion, anomalous limbs, and action mismatches. To address this, we introduce a pioneering new task, Personalized Martial Arts Combat Video Generation. Our approach, MagicFight, is specifically crafted to overcome these hurdles. Given this pioneering task, we face a lack of appropriate datasets. Thus, we generate a bespoke dataset using the game physics engine Unity, meticulously crafting a multitude of 3D characters, martial arts moves, and scenes designed to represent the diversity of combat. MagicFight refines and adapts existing models and strategies to generate high-fidelity two-person combat videos that maintain individual identities and ensure seamless, coherent action sequences, thereby laying the groundwork for future innovations in the realm of interactive video content creation. Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang 0035, Shifeng Chen |
ACM Multimedia | 5 |
| 2024 | WizardArena: Post-training Large Language Models via Simulated Offline Chatbot ArenaabstractRecent work demonstrates that, post-training large language models with open-domain instruction following data have achieved colossal success. Simultaneously, human Chatbot Arena has emerged as one of the most reasonable benchmarks for model evaluation and developmental guidance. However, the processes of manually curating high-quality training data and utilizing online human evaluation platforms are both expensive and limited. To mitigate the manual and temporal costs associated with post-training, this paper introduces a Simulated Chatbot Arena named WizardArena, which is fully based on and powered by open-source LLMs. For evaluation scenario, WizardArena can efficiently predict accurate performance rankings among different models based on offline test set. For training scenario, we simulate arena battles among various state-of-the-art models on a large scale of instruction data, subsequently leveraging the battle results to constantly enhance target model in both the supervised fine-tuning and reinforcement learning . Experimental results demonstrate that our WizardArena aligns closely with the online human arena rankings, and our models trained on offline extensive battle data exhibit significant performance improvements during SFT, DPO, and PPO stages. Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Qingwei Lin, Jian-Guang Lou, Shifeng Chen, Yansong Tang, Weizhu Chen |
NeurIPS | 7 |
| 2024 | IFAST: Weakly Supervised Interpretable Face Anti-Spoofing From Single-Shot Binocular NIR ImagesabstractSingle-shot face anti-spoofing (FAS) is a key technique for securing face recognition systems, relying solely on static images as input. However, single-shot FAS remains a challenging and under-explored problem due to two reasons: 1) On the data side, learning FAS from RGB images is largely context-dependent, and single-shot images without additional annotations contain limited semantic information. 2) On the model side, existing single-shot FAS models struggle to provide proper evidence for their decisions, and FAS methods based on depth estimation require expensive per-pixel annotations. To address these issues, we construct and release a large binocular NIR image dataset named BNI-FAS, which contains more than 300,000 real face and plane attack images, and propose an Interpretable FAS Transformer (IFAST) that requires only weak supervision to produce interpretable predictions. Our IFAST generates pixel-wise disparity maps using the proposed disparity estimation Transformer with Dynamic Matching Attention (DMA) blocks. Besides, we design a confidence map generator to work in tandem with a dual-teacher distillation module to obtain the final discriminant results. Comprehensive experiments show that our IFAST achieves state-of-the-art performance on BNI-FAS, verifying its effectiveness of single-shot FAS on binocular NIR images. The project page is available athttps://ifast-bni.github.io/. Jiancheng Huang, Jianzhuang Liu, Linxiao Shi, Shifeng Chen |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | WaveDM: Wavelet-Based Diffusion Models for Image RestorationabstractLatest diffusion-based methods for many image restoration tasks outperform traditional models, but they encounter the long-time inference problem. To tackle it, this paper proposes a Wavelet-Based Diffusion Model (WaveDM). WaveDM learns the distribution of clean images in the wavelet domain conditioned on the wavelet spectrum of degraded images after wavelet transform, which is more time-saving in each step of sampling than modeling in the spatial domain. To ensure restoration performance, a unique training strategy is proposed where the low-frequency and high-frequency spectrums are learned using distinct modules. In addition, an Efficient Conditional Sampling (ECS) strategy is developed from experiments, which reduces the number of total sampling steps to around 5. Evaluations on twelve benchmark datasets including image raindrop removal, rain steaks removal, dehazing, defocus deblurring, demoiréing, and denoising demonstrate that WaveDM achieves state-of-the-art performance with the efficiency that is comparable to traditional one-pass methods and over 100× faster than existing image restoration methods using vanilla diffusion models. The code is available athttps://github.com/stayalive16/WaveDM Yi Huang 0035, Jiancheng Huang, Jianzhuang Liu, Mingfu Yan, Jiaxi Lv, Chaoqi Chen, Shifeng Chen |
IEEE Trans. Multim. | 8 |
| 2024 | V4D: Voxel for 4D Novel View SynthesisabstractNeural radiance fields have made a remarkable breakthrough in the novel view synthesis task at the 3D static scene. However, for the 4D circumstance (e.g., dynamic scene), the performance of the existing method is still limited by the capacity of the neural network, typically in a multilayer perceptron network (MLP). In this article, we utilize 3D Voxel to model the 4D neural radiance field, short as V4D, where the 3D voxel has two formats. The first one is to regularly model the 3D space and then use the sampled local 3D feature with the time index to model the density field and the texture field by a tiny MLP. The second one is in look-up tables (LUTs) format that is for the pixel-level refinement, where the pseudo-surface produced by the volume rendering is utilized as the guidance information to learn a 2D pixel-level refinement mapping. The proposed LUTs-based refinement module achieves the performance gain with little computational cost and could serve as the plug-and-play module in the novel view synthesis task. Moreover, we propose a more effective conditional positional encoding toward the 4D data that achieves performance gain with negligible computational burdens. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance at a low computational cost. Wanshui Gan, Yi Huang 0035, Shifeng Chen, Naoto Yokoya |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | KV Inversion: KV Embeddings Learning for Text-Conditioned Real Image Action Editing
Jiancheng Huang, Shifeng Chen |
PRCV (1) | 4 |
| 2023 | Bootstrap Diffusion Model Curve Estimation for High Resolution Low-Light Image Enhancement
Jiancheng Huang, Shifeng Chen |
PRICAI (3) | 3 |
| 2023 | Rethinking 3D cost aggregation in stereo matchingabstractIn the stereo matching task, the 3D convolution network can effectively aggregate the cost volume with the strong representation ability to model the spatial and depth dimensions but with the disadvantage of a high computational cost. In this letter, we revisit the 3D convolution network and its common variant, and then propose the Depth Shift Module (DSM) to model the cost volume in the depth dimension which could imitate the 3D convolution function with the computational complexity of the 2D convolution. The proposed DSM is easy to extend to present 3D cost aggregation methods in stereo matching with less inference time, lower computational complexity, and minor precision loss. Moreover, a novel compact but efficient stereo matching framework named HybridNet is proposed. This framework can hybridize the 2D convolution layer with the proposed DSM to effectively aggregate the cost volume. The proposed HybridNet achieves a better trade-off between the performance, computational complexity, and model size ( e.g. , 30% less than the size of AANet and 25% less than the size of PSMNet) in public open-source datasets ( e.g. , Scene Flow and KITTI Stereo 2015). The relevant code is available at https://github.com/GANWANSHUI/HybridNet . Wanshui Gan, Shifeng Chen, Pak-Kin Wong 0001 |
Pattern Recognit. Lett. | 3 |
| 2023 | DeSeal: Semantic-Aware Seal2Clear Attention for Document Seal RemovalabstractSeal removal aims to eliminate the seal portion from documents to facilitate better OCR and document reconstruction. However, existing seal removal methods often lack publicly available code and pre-trained models, and suffer from a lack of publicly available seal datasets. To address these issues, we propose DeSeal for seal removal and introduce a SealBank dataset containing 100K paired seal images. In DeSeal, we introduce the Semantic-Aware Seal2Clear Attention and Color-Adapter Module, where the former identifies seal regions in the entire image and focuses on removing seals from these areas, while the latter significantly improves the model's generalization performance, enabling it to perform well on both real and synthetic data. Experimental results on the SealBank dataset demonstrate the effectiveness of our proposed DeSeal. Jiancheng Huang, Shifeng Chen |
IEEE Signal Process. Lett. | 3 |
| 2023 | Semi-Supervised Domain Alignment Learning for Single Image DehazingabstractConvolutional neural networks (CNNs) have attracted much research attention and achieved great improvements in single-image dehazing. However, previous learning-based dehazing methods are mainly trained on synthetic data, which greatly degrades their generalization capability on natural hazy images. To address this issue, this article proposes a semi-supervised learning approach for single-image dehazing, where both synthetic and realistic images are leveraged during training. Considering the situation that it is hard to obtain the realistic pairs of hazy and haze-free images, how to utilize the realistic data is not a trivial work. In this article, a domain alignment module is introduced to narrow the distribution distance between synthetic data and realistic hazy images in a latent feature space. Meanwhile, a haze-aware attention module is designed to describe haze densities of different regions in the image, thus adaptively responds for different hazy areas. Furthermore, the dark channel prior is introduced to the framework to improve the quality of the unsupervised learning results by considering the statistical characters of haze-free images. Such a semi-supervised design can significantly address the domain shift issue between the synthetic and realistic data, and improve generalization performance in the real world. Experiments indicate that the proposed method obtains state-of-the-art performance on both public synthetic and realistic hazy images with better visual results. Yunan Li 0001, He Zhang 0004, Shifeng Chen |
IEEE Trans. Cybern. | 5 |
| 2023 | Density Distillation for Fast Nonparametric Density EstimationabstractNonparametric density estimation has been extensively used in various application scenarios and theoretical models. However, the modeling of these powerful methods is inseparable from the sample data and comes at the cost of repeated and intensive kernel calculations, which makes their efficiency greatly affected by the sample scale, data dimension, and evaluation scale. Inspired by the knowledge distillation method, a student-teacher paradigm model named density convolutional neural network (DCNN) is proposed in this article. The method extracts the density knowledge of the samples based on the density convolution rule and transfers it to a compact and small deep neural network, in order to separate the sample data from the modeling and avoid the cumbersome kernel calculations. Experimental results show the superiority of the proposed method to various nonparametric estimation methods in terms of accuracy, stability, processing efficiency, and low-storage advantage. Especially, for the estimation speed, a univariate density estimation on 1.0E + 08 evaluation points using GPU only takes 1.57 s, and a 10-D multivariate density estimation on 1.0E + 08 evaluation points only takes 10.50 s, which makes our method very suitable for real-time and large-scale repetitive density estimation tasks. Bopeng Fang, Shifeng Chen, Zhurong Dong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | ES6D: A Computation Efficient and Symmetry-Aware 6D Pose Regression FrameworkabstractIn this paper, a computation efficient regression framework is presented for estimating the 6D pose of rigid objects from a single RGB-D image, which is applicable to handling symmetric objects. This framework is designed in a simple architecture that efficiently extracts point-wise features from RGB-D data using a fully convolutional network, called XYZNet, and directly regresses the 6D pose without any post refinement. In the case of symmetric object, one object has multiple ground-truth poses, and this one-to-many relationship may lead to estimation ambiguity. In order to solve this ambiguity problem, we design a symmetry-invariant pose distance metric, called average (maximum) grouped primitives distance or A(M)GPD. The proposed A(M)GPD loss can make the regression network converge to the correct state, i.e., all minima in the A(M)GPD loss surface are mapped to the correct poses. Extensive experiments on YCB-Video and TLESS datasets demonstrate the proposed framework's substantially superior performance in top accuracy and low computational cost. The relevant code is available in https://github.com/GANWANSHUI/ES6D.git. Ningkai Mo, Wanshui Gan, Naoto Yokoya, Shifeng Chen |
CVPR | 4 |
| 2022 | Semi-Supervised Segmentation of Radiation-Induced Pulmonary Fibrosis From Lung CT Scans With Multi-Scale Guided Dense AttentionabstractComputed Tomography (CT) plays an important role in monitoring radiation-induced Pulmonary Fibrosis (PF), where accurate segmentation of the PF lesions is highly desired for diagnosis and treatment follow-up. However, the task is challenged by ambiguous boundary, irregular shape, various position and size of the lesions, as well as the difficulty in acquiring a large set of annotated volumetric images for training. To overcome these problems, we propose a novel convolutional neural network called PF-Net and incorporate it into a semi-supervised learning framework based on Iterative Confidence-based Refinement And Weighting of pseudo Labels (I-CRAWL). Our PF-Net combines 2D and 3D convolutions to deal with CT volumes with large inter-slice spacing, and uses multi-scale guided dense attention to segment complex PF lesions. For semi-supervised learning, our I-CRAWL employs pixel-level uncertainty-based confidence-aware refinement to improve the accuracy of pseudo labels of unannotated images, and uses image-level uncertainty for confidence-based image weighting to suppress low-quality pseudo labels in an iterative training process. Extensive experiments with CT scans of Rhesus Macaques with radiation-induced PF showed that: 1) PF-Net achieved higher segmentation accuracy than existing 2D, 3D and 2.5D neural networks, and 2) I-CRAWL outperformed state-of-the-art semi-supervised learning methods for the PF lesion segmentation task. Our method has a potential to improve the diagnosis of PF and clinical assessment of side effects of radiotherapy for lung cancers. Guotai Wang, Shuwei Zhai, Giovanni Lasio, Baoshe Zhang, Byong Yi, Shifeng Chen, Thomas J. Macvittie, Dimitris N. Metaxas, Jinghao Zhou, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Deep Network Quantization via Error CompensationabstractFor portable devices with limited resources, it is often difficult to deploy deep networks due to the prohibitive computational overhead. Numerous approaches have been proposed to quantize weights and/or activations to speed up the inference. Loss-aware quantization has been proposed to directly formulate the impact of weight quantization on the model's final loss. However, we discover that, under certain circumstances, such a method may not converge and end up oscillating. To tackle this issue, we introduce a novel loss-aware quantization algorithm to efficiently compress deep networks with low bit-width model weights. We provide a more accurate estimation of gradients by leveraging the Taylor expansion to compensate for the quantization error, which leads to better convergence behavior. Our theoretical analysis indicates that the gradient mismatch issue can be fixed by the newly introduced quantization error compensation term. Experimental results for both linear models and convolutional networks verify the effectiveness of our proposed method. Hanyu Peng, Jiaxiang Wu 0001, Zhiwei Zhang 0012, Shifeng Chen, Hai-Tao Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Deep center-based dual-constrained hashing for discriminative face image retrieval
Ming Zhang 0023, Xuefei Zhe, Shifeng Chen, Hong Yan 0001 |
Pattern Recognit. | 3 |
| 2020 | FD-GAN: Generative Adversarial Networks with Fusion-Discriminator for Single Image DehazingabstractRecently, convolutional neural networks (CNNs) have achieved great improvements in single image dehazing and attained much attention in research. Most existing learning-based dehazing methods are not fully end-to-end, which still follow the traditional dehazing procedure: first estimate the medium transmission and the atmospheric light, then recover the haze-free image based on the atmospheric scattering model. However, in practice, due to lack of priors and constraints, it is hard to precisely estimate these intermediate parameters. Inaccurate estimation further degrades the performance of dehazing, resulting in artifacts, color distortion and insufficient haze removal. To address this, we propose a fully end-to-end Generative Adversarial Networks with Fusion-discriminator (FD-GAN) for image dehazing. With the proposed Fusion-discriminator which takes frequency information as additional priors, our model can generator more natural and realistic dehazed images with less color distortion and fewer artifacts. Moreover, we synthesize a large-scale training dataset including various indoor and outdoor hazy images to boost the performance and we reveal that for learning-based dehazing methods, the performance is strictly influenced by the training data. Experiments have shown that our method reaches state-of-the-art performance on both public synthetic datasets and real-world images with more visually pleasing dehazed results. Yihao Liu 0001, He Zhang 0004, Shifeng Chen, Yu Qiao 0001 |
AAAI | 4 |
| 2020 | P-KDGAN: Progressive Knowledge Distillation with GANs for One-class Novelty DetectionabstractOne-class novelty detection is to identify anomalous instances that do not conform to the expected normal instances. In this paper, the Generative Adversarial Networks (GANs) based on encoder-decoder-encoder pipeline are used for detection and achieve state-of-the-art performance. However, deep neural networks are too over-parameterized to deploy on resource-limited devices. Therefore, Progressive Knowledge Distillation with GANs (P-KDGAN) is proposed to learn compact and fast novelty detection networks. The P-KDGAN is a novel attempt to connect two standard GANs by the designed distillation loss for transferring knowledge from the teacher to the student. The progressive learning of knowledge distillation is a two-step approach that continuously improves the performance of the student GAN and achieves better performance than single step methods. In the first step, the student GAN learns the basic knowledge totally from the teacher via guiding of the pre-trained teacher GAN with fixed weights. In the second step, joint fine-training is adopted for the knowledgeable teacher and student GANs to further improve the performance and stability. The experimental results on CIFAR-10, MNIST, and FMNIST show that our method improves the performance of the student GAN by 2.44%, 1.77%, and 1.73% when compressing the computation at ratios of 24.45:1, 311.11:1, and 700:1, respectively. Shifeng Chen |
IJCAI | 2 |
| 2020 | Multi-scale fully convolutional network for gland segmentation using three-class classification
Huijun Ding, Zhanpeng Pan, Qian Cen, Shifeng Chen |
Neurocomputing | 5 |
| 2020 | Deep Class-Wise Hashing: Semantics-Preserving Hashing via Class-Wise LossabstractDeep supervised hashing has emerged as an effective solution to large-scale semantic image retrieval problems in computer vision. Convolutional neural network-based hashing methods typically seek pairwise or triplet labels to conduct similarity-preserving learning. However, complex semantic concepts of visual contents are hard to capture by similar/dissimilar labels, which limits the retrieval performance. Generally, pairwise or triplet losses not only suffer from expensive training costs but also lack sufficient semantic information. In this paper, we propose a novel deep supervised hashing model to learn more compact class-level similarity-preserving binary codes. Our model is motivated by deep metric learning that directly takes semantic labels as supervised information in training and generates corresponding discriminant hashing code. Specifically, a novel cubic constraint loss function based on Gaussian distribution is proposed, which preserves semantic variations while penalizes the overlapping part of different classes in the embedding space. To address the discrete optimization problem introduced by binary codes, a two-step optimization strategy is proposed to provide efficient training and avoid the problem of gradient vanishing. Extensive experiments on five large-scale benchmark databases show that our model can achieve the state-of-the-art retrieval performance. Xuefei Zhe, Shifeng Chen, Hong Yan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionabstractVideo Recognition has drawn great research interest and great progress has been made. A suitable frame sampling strategy can improve the accuracy and efficiency of recognition. However, mainstream solutions generally adopt hand-crafted frame sampling strategies for recognition. It could degrade the performance, especially in untrimmed videos, due to the variation of frame-level saliency. To this end, we concentrate on improving untrimmed video classification via developing a learning-based frame sampling strategy. We intuitively formulate the frame sampling procedure as multiple parallel Markov decision processes, each of which aims at picking out a frame/clip by gradually adjusting an initial sampling. Then we propose to solve the problems with multi-agent reinforcement learning (MARL). Our MARL framework is composed of a novel RNN-based context-aware observation network which jointly models context information among nearby agents and historical states of a specific agent, a policy network which generates the probability distribution over a predefined action space at each step and a classification network for reward calculation as well as final recognition. Extensive experimental results show that our MARL-based scheme remarkably outperforms hand-crafted strategies with various 2D and 3D baseline methods. Our single RGB model achieves a comparable performance of ActivityNet v1.3 champion submission with multi-modal multi-model fusion and new state-of-the-art results on YouTube Birds and YouTube Cars. Dongliang He, Xiao Tan 0001, Shifeng Chen, Shilei Wen |
ICCV | 4 |
| 2019 | Collaborative Channel Pruning for Deep NetworksabstractDeep networks have achieved impressive performance in various domains, but their applications are largely limited by the prohibitive computational overhead. In this paper, we propose a novel algorithm, namely collaborative channel pruning (CCP), to reduce the computational overhead with negligible performance degradation. The joint impact of pruned/preserved channels on the loss function is quantitatively analyzed, and such interchannel dependency is exploited to determine which channels to be pruned. The channel selection problem is then reformulated as a constrained 0-1 quadratic optimization problem, and the Hessian matrix, which is essential in constructing the above optimization, can be efficiently approximated. Empirical evaluation on two benchmark data sets indicates that our proposed CCP algorithm achieves higher classification accuracy with similar computational complexity than other stateof-the-art channel pruning algorithms Hanyu Peng, Jiaxiang Wu 0001, Shifeng Chen, Junzhou Huang |
ICML | 3 |
| 2019 | Directional statistics-based deep metric learning for image classification and retrieval
Xuefei Zhe, Shifeng Chen, Hong Yan 0001 |
Pattern Recognit. | 2 |
| 2019 | BDNN: Binary convolution neural networks for fast object detection
Hanyu Peng, Shifeng Chen |
Pattern Recognit. Lett. | 2 |
| 2019 | Dual-supervised attention network for deep cross-modal hashing
Hanyu Peng, Junjun He, Shifeng Chen, Yali Wang 0001, Yu Qiao 0001 |
Pattern Recognit. Lett. | 3 |
| 2015 | Sketch-based 3-D modeling for piecewise planar objects in single images
Changqing Zou, Xiaojiang Peng, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu |
Comput. Graph. | 4 |
| 2015 | Detecting Co-Salient Objects in Large Image SetsabstractCo-salient object detection has attracted much more attention recently as it is useful for many problems in vision computing. However, most of existing methods emphasize detecting the common salient objects in a small group of images and the objects of interest in those images have clear borders with respect to the backgrounds. In this work, we propose a novel co-saliency detection method, which aims at discovering the common objects in a large and diverse image set composed of hundreds of images. First, we search a group of similar images for each image in the set. Our method is based on the overlapped groups. We handle each group with an unsupervised random forest to extract the rough contours of the common objects. Then a contrast-based measure is utilized to produce the saliency map for an individual image. For each image in the set, we collect all the maps from the groups that contain the image and fuse them together as the inter-saliency map for the image. The final co-saliency map is computed by combining the inter-saliency map with the single saliency map of this image. Experimental evaluation on an established large dataset demonstrates that our approach attains superior results and outperforms the state-of-the-art methods. Shuze Du, Shifeng Chen |
IEEE Signal Process. Lett. | 2 |
| 2015 | Progressive 3D Reconstruction of Planar-Faced Manifold Objects with DRF-Based Line Drawing DecompositionabstractThis paper presents an approach for reconstructing polyhedral objects from single-view line drawings. Our approach separates a complex line drawing representing a manifold object into a series of simpler line drawings, based on the degree of reconstruction freedom (DRF). We then progressively reconstruct a complete 3D model from these simpler line drawings. Our experiments show that our decomposition algorithm is able to handle complex drawings which are challenging for the state of the art. The advantages of the presented progressive 3D reconstruction method over the existing reconstruction methods in terms of both robustness and efficiency are also demonstrated. Changqing Zou, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | Learning Flexible Binary Code for Linear Projection Based Hashing with Random ForestabstractExisting linear projection based hashing methods have witnessed many progresses in finding the approximate nearest neighbor(s) of a given query. They perform well when using a short code. But their code length depends on the original data dimension, thus their performance can not be further improved with higher number of bits for low dimensional data. In addition, in the case of high dimensional data, it is not a good choice to produce each bit by a sign function. In this paper, we propose a novel random forest based approach to cope with the above shortcomings. The bits are obtained by recording the paths when a point traversing each tree in the forest. Then we propose a new metric to calculate the similarity between any two codes. Experimental results on two large benchmark datasets show that our approach outperforms its counterparts and demonstrate its superiority over the existing state-of-the-art hashing methods for descriptor retrieval. Shuze Du, Wei Zhang 0081, Shifeng Chen, Yafei Wen |
ICPR | 3 |
| 2014 | Sketch-Based 3D Model Retrieval via Multi-feature FusionabstractSketch-based 3D model retrieval provides a convenient way for users to search for 3D models by sketches. Traditionally, this task is converted to a sketch-based 2D shape retrieval problem by projecting 3D models to 2D images. Local invariant features have been widely used to tackle this problem. However, it suffers from the lack of global context and easily fails when images of different 3D models share multiple similar regions. In this paper, we propose a joint description by fusing local statistical structures and global spatial features. Our description is invariant to scale, translate and rotation. An improved bag-of-features retrieval framework is applied to explore semantic visual word representations. Besides, a novel relevance feedback scheme which combines weight balancing and query modification is designed to further improve the retrieval performance. We conduct various experiments on the common sketch-based watertight model benchmark. The comparative results show that our approach significantly outperforms three state-of-the-art methods, demonstrating its effectiveness and robustness for sketch-based 3D model retrieval. Yafei Wen, Changqing Zou, Jianzhuang Liu, Shuze Du, Shifeng Chen |
ICPR | 5 |
| 2014 | Learning Boundary and Appearance for Video Object CutoutabstractThis letter presents an approach to object cutout for arbitrary videos. Our approach is in a “learning and propagating” way. First, the object(s) of interest is (are) cut out in the first frame (or the key frame) in the video clip using some interactive image segmentation tool. Then some statistical features of the object regions, the non-object regions, and the boundaries are learnt from the result of image segmentation. Finally, the result is propagated to the other frames automatically. We formulate the “learning and propagating” step in Markov random fields (MRFs). In this model, we well design the patch term and the boundary term to significantly improve the performance of the algorithm. Experimental results indicate that the algorithm performs excellent. Shifeng Chen, Huijun Ding |
IEEE Signal Process. Lett. | 1 |
| 2014 | Salient Object Detection via Random ForestabstractSalient object detection plays an important role in image pre-processing. Existing approaches often neglect the contours of salient objects, thus resulting in inaccurate detection for large objects. Besides, they mainly focus on detecting only a single object. In this paper, we detect the salient object from the view of the object contour. We propose to exploit the random forest to measure patch rarities and compute similarities among patches. A global rarity map is calculated based on the patch's rareness over the whole image. The approximate contour of the salient object is extracted based on this rarity map by using an active contour model. Next, a local saliency map is obtained by the similarities of patches inside the contour and those outside. Finally, the local map is refined through image segmentation. Our method can detect not only a single object but also multiple objects. Experimental evaluation on the ASD-1000 and SED2 datasets shows that our method outperforms the state-of-the-art methods. Shuze Du, Shifeng Chen |
IEEE Signal Process. Lett. | 2 |
| 2013 | AdVisual: a visual-based advertising systemabstractIn this work, we present a visual-based contextual advertising system, AdVisual. It is designed for content service providers to effectively select relevant ads for online videos. First, it will analyze each video and extract high level semantic visual information including specific objects, people, and significant scenes. Then ads highly related to the visual concepts are associated with the corresponding shots. As AdVisual is a user interaction system, it allows users to select favorite ads relevant to the video. By saliency detection, selected ads will be displayed as an overlay window embedded at the non-intrusive part of the shot. Chao Dong 0005, Shifeng Chen, Xiaoou Tang |
ACM Multimedia | 2 |
| 2013 | Edge preserving image denoising with a closed form solution
Shifeng Chen, Wei Zhang 0081, Jianzhuang Liu |
Pattern Recognit. | 1 |
| 2013 | Style Transfer Via Image Component AnalysisabstractExample-based stylization provides an easy way of making artistic effects for images and videos. However, most existing methods do not consider the content and style separately. In this paper, we propose a style transfer algorithm via a novel component analysis approach, based on various image processing techniques. First, inspired by the steps of drawing a picture, an image is decomposed into three components: draft, paint and edge, which describe the content, main style, and strengthened strokes along the boundaries. Then the style is transferred from the template image to the source image in the paint and edge components. Style transfer is formulated as a global optimization problem by using Markov random fields, and a coarse-to-fine belief propagation algorithm is used to solve the optimization problem. To combine the draft component and the obtained style information, the final artistic result can be achieved via a reconstruction step. Compared to other algorithms, our method not only synthesizes the style, but also preserves the image content well. We also extend our algorithm from single image stylization to video personalization, by maintaining the temporal coherence and identifying faces in video sequences. The results indicate that our approach performs excellently in stylization and personalization for images and videos. Wayne Zhang 0001, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Multim. | 3 |
| 2012 | Locating high-density clusters with noisy queries
Shifeng Chen, Changqing Zou, Jianzhuang Liu |
ICPR | 2 |
| 2012 | Online non-feedback image re-ranking via dominant data selectionabstractImage re-ranking aims at improving the precision of keyword-based image retrieval, mainly by introducing visual features to re-rank. Many existing approaches require offline training for every keyword, which are unsuitable for online image search. Other real-time approaches demand user interaction, which are inappropriate for large-scale image collection. To improve the accuracy of web image retrieval in a practicable manner, we propose a novel re-ranking algorithm to explore the cluster information of the image set. First, we build spectral graph on images that retrieved bysearch engine, and remove isolated nodes as noisy images. Then, we select positive samples from the most dominant cluster in initial top-ranked images, and the samples are used for semi-supervised learning and ranking. Our algorithm is online and non-feedback. Experiments on two public databases demonstrate that our algorithm outperforms the state-of-the-art approaches. Shifeng Chen, Jianzhuang Liu |
ACM Multimedia | 2 |
| 2011 | Automatic motion-guided video stylization and personalizationabstractVideo stylization transfers a source video into an artistic version while maintaining temporal coherence between adjacent frames. In this paper, we formulate the unsupervised example-based video stylization with Markov random field model. In our algorithm, we implement an improved optical flow algorithm to maintain temporal coherence while improve the accuracy of estimation along motion boundaries. We also extend our algorithm to the application of video personalization, in which human faces keep clear and distinguishable. A series of techniques are fused in video personalization, including face detection and alignment, motion flow, skin detection, and illumination blending. Given a source video and a style template image, our algorithm produces the stylized and/or personalized video(s) automatically. Experimental results demonstrate that our algorithm performs excellently in both video stylization and personalization. Shifeng Chen, Wei Zhang 0081, Xiaoou Tang |
ACM Multimedia | 2 |
| 2011 | Edge-preserving single image super-resolutionabstractThis paper proposes a novel approach to single image super-resolution. First, an image up-sampling scheme is proposed which takes the advantages of both bilateral filtering and mean shift image segmentation. Then we use a shock filter to enhance strong edges in the initial up-sampling result and obtain an intermediate high-resolution image. Finally, we enforce a reconstruction constraint on the high-resolution image so that fine details can be inferred by back projection. Since strong edges in the intermediate result are enhanced, ringing artifacts can be suppressed in the back projection step. We compare our algorithm with several state-of-the-art image super-resolution algorithms. Qualitative and quantitative experimental results demonstrate that our approach performs the best. Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |
| 2010 | Continuous MRF based image denoising with a closed form solutionabstractIn this paper, we formulate the problem of image denoising as the maximum a posterior (MAP) estimation problem using Markov random fields (MRFs). Such an MAP estimation for MRFs is equivalent to a maximum likelihood estimation constrained on spatial homogeneity and is generally NP-hard in the discrete domain. To make it tractable, we convert it to a continuous label assignment problem based on a Gaussian MRF model and then obtain a closed form globally optimal solution. Since the Gaussian MRFs tend to over-smooth images and blur edges, we incorporate pre-estimated image edge information into the energy function to better preserve image structures. Patch similarity based pairwise interaction is also involved to better preserve image details and make the algorithm more robust to impulse noise. Both quantitative and qualitative comparative experimental results are given to demonstrate the better performance of our algorithm. Shifeng Chen, Jianzhuang Liu |
ICIP | 2 |
| 2010 | Fast image rearrangement via multi-scale patch copyingabstractIn this paper, we propose a simple interactive way for a novel type of image synthesis called image rearrangement whose goal is to construct a new image based on some objects cropped from source images. The synthesis results are obtained by copying patches from the source images in a globally consistent way. The patch copying problem is formulated with the Markov random field model, and belief propagation is used as the optimization tool. To speed up our algorithm, a two-step belief propagation and a multi-scale patch copying scheme are taken. Experimental results indicate that our algorithm obtains satisfactory results in both performance and efficiency. Jiayao Hu, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |
| 2010 | A scalable intelligent non-content-based spam-filtering framework
Yong Hu 0002, Eric W. T. Ngai, Shifeng Chen |
Expert Syst. Appl. | 5 |
| 2010 | Image Segmentation by MAP-ML EstimationsabstractImage segmentation plays an important role in computer vision and image analysis. In this paper, image segmentation is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs) and a graph cut algorithm is used to find the solution to the MAP estimation. The ML estimation is achieved by computing the means of region features in a Gaussian model. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. Its results match image edges very well and are consistent with human perception. Comparing to six state-of-the-art algorithms, extensive experiments have shown that our algorithm performs the best. Shifeng Chen, Liangliang Cao, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Image Process. | 1 |
| 2009 | Video completion via motion guided spatial-temporal global optimizationabstractIn this paper, a novel global optimization based approach is proposed for video completion whose target is to restore the spatial-temporal missing regions of a video in a visually plausible way. Our algorithm consists of two stages: motion field completion and color completion via global optimization. First, local motions within the missing parts are completed patch-by-patch greedily using pre-computed available motions in the video. Then the missing regions are filled by sampling patches from available parts of the video. We formulate the video completion as a global energy minimization problem by Markov random fields (MRFs). Based on the completed motion field of the video, a well-defined energy function involving both spatial and temporal coherence relationship is constructed. A coarse-to-fine Belief Propagation (BP) is proposed to solve the optimization problem. Experimental results have demonstrated the good performance of our algorithm. Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |
| 2008 | Decision Fusion of Machine Learning Models to Predict Radiotherapy-Induced Lung PneumonitisabstractCombining different machine learning models (decision fusion) has been shown to be an effective method for estimating the underlying physical mechanism by allowing the models to reinforce each other when consensus exists, or, conversely, negate each other when there is no consensus. To be effective, decision fusion requires that the different models provide some degree of complementary information. In this work, we fuse the results of four different machine learning models (Boosted Decision Trees, Neural Networks, Support Vector Machines, Self Organizing Maps) to predict the risk of lung pneumonitis in patients undergoing thoracic radiotherapy. Fusion was achieved by simple averaging of the 10-fold cross validated predictions for each patient from all four models. To reduce prediction dependence on the manner in which the data set was split, 10-fold cross-validation was repeated 100 times for random data splitting. The area under the receiver operating characteristics curve for the fused cross-validated results was 0.79, higher than the individual models and with (generally) lower variance. The fusion extracted three important features as the consensus among all four models in predicting radiation pneumonitis risk: chemotherapy prior to radiotherapy, equivalent Uniform Dose (EUD) for exponent a = 1.2 to 3, and female gender. The results show great promise for machine learning in radiotherapy outcomes modeling. Shiva K. Das, Shifeng Chen, Joseph O. Deasy, Sumin Zhou, Fang-Fang Yin, Lawrence B. Marks |
ICMLA | 2 |
| 2008 | Easytoon: an easy and quick tool to personalize a cartoon storyboard using family photo albumabstractA family photo album based cartoon personalization system, EasyToon, is proposed in this paper. Using state of the art computer vision and graphics technologies and effective UI design, the interactive tool can quickly generate a personalized cartoon storyboard, which naturally blends a real face chosen from the family photo album into a cartoon picture. The personalized cartoon image is easily and quickly obtained in two main steps. First, the best face candidate is selected from the album interactively. Then a personalized cartoon image is automatically synthesized by blending the selected face into the interesting cartoon image. Experiments show that most users express great interest in our system. Without any art background, they can make a personalized cartoon of high quality using the EasyToon within minutes. Shifeng Chen, Yuandong Tian, Fang Wen 0001, Ying-Qing Xu, Xiaoou Tang |
ACM Multimedia | 1 |
| 2008 | Precise object cutout from imagesabstractIn this paper we propose a novel approach to the problem of interactive foreground/background segmentation in images. With user provided strokes which indicate foreground and background seeds, we estimate two Gaussian mixture models, one for foreground and the other for background, and define two quantities to measure the initial probabilities of each pixel belonging to the foreground and the background respectively. An optimization function constructed based on the quantities and the boundary and coherent region information is proposed to solve the segmentation problem. By relaxing the hard binary segmentation to a soft labelling problem in the continuous domain, a closed form global optimal solution can be achieved, which directly results in the final binary segmentation output. Experimental results demonstrate the excellent performance of our algorithm. Shifeng Chen, Jianzhuang Liu |
ACM Multimedia | 2 |
| 2008 | EasyToon: cartoon personalization using face photosabstractIn this demo, we present a family photo album based cartoon personalization system, EasyToon. Using the family photo album as the candidate pool, a personalized cartoon image is obtained in two main steps. First, the best face candidate is selected from the album interactively. Then a personalized cartoon image is automatically synthesized by lending the selected face into the target cartoon image. By integrating state of the art computer vision and graphics technologies and effective UI design EasyToon can generate a personalized cartoon storyboard easily and quickly. Fang Wen 0001, Shifeng Chen, Xiaoou Tang |
ACM Multimedia | 2 |
| 2007 | Iterative MAP and ML Estimations for Image SegmentationabstractImage segmentation plays an important role in computer vision and image analysis. In this paper, the segmentation problem is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum-likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs). A graph-cut algorithm is used to find the solution to the MAP-MRF estimation. The ML estimation is achieved by finding the means of region features. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. In addition, under the same framework, it can be extended to another algorithm that extracts objects of a particular class from a group of images. Extensive experiments have shown the effectiveness of our approach. Shifeng Chen, Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
CVPR | 1 |
| 2007 | Noise Robust Spectral ClusteringabstractThis paper aims to introduce the robustness against noise into the spectral clustering algorithm. First, we propose a warping model to map the data into a new space on the basis of regularization. During the warping, each point spreads smoothly its spatial information to other points. After the warping, empirical studies show that the clusters become relatively compact and well separated, including the noise cluster that is formed by the noise points. In this new space, the number of clusters can be estimated by eigenvalue analysis. We further apply the spectral mapping to the data to obtain a low-dimensional data representation. Finally, the K-means algorithm is used to perform clustering. The proposed method is superior to previous spectral clustering methods in that (i) it is robust against noise because the noise points are grouped into one new cluster; (ii) the number of clusters and the parameters of the algorithm are determined automatically. Experimental results on synthetic and real data have demonstrated this superiority. Zhenguo Li, Jianzhuang Liu, Shifeng Chen, Xiaoou Tang |
ICCV | 3 |
| 2007 | Image matting using linear optimizationabstractAn image can be assumed to be a composite of the foreground and the background. The foreground and the background of each pixel are linearly combined in terms of this pixel's foreground opacity (called alpha). Image matting is the process of estimating the foreground, the background and the alpha for each pixel. In this paper, we transform the ill-posed image matting problem into two over-determined linear optimization problems by introducing two medium variables and imposing smoothness constraints. Closed form solutions can be obtained from the two problems. Extensive experimental results indicate that our algorithm can generate high-quality matting results. Shifeng Chen, Zhenguo Li, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 1 |
| 2007 | Image inpainting by global structure and texture propagationabstractImage inpainting is a technique to repair damaged images or modify images in a non-detectable form. In this paper, a novel global algorithm for region filling is proposed for image inpainting. After removing objects from an image, our approach fills the regions using patches taken from the image. The filling process is formulated as an energy minimization problem by Markov random fields (MRFs) and the belief propagation (BP) is utilized to solve the problem. Our energy function includes structure and texture information obtained from the image. One challenge in using BP is that its computational complexity is the square of the number of label candidates. To reduce the large number of label candidates, we present a coarse-to-fine scheme where two BPs run with much smaller numbers of label candidates instead of one BP running with a large number of label candidates. Experimental results demonstrate the excellent performance of our algorithm over other related algorithms. Huang Ting, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |