EDBT 2026 Demo / reviewers in the wild / expert
Ning Xu 0007
dblp:04/5856-7
· DBLP profile ↗
35ranked-venue papers
5as first author
11since 2021 · last 2024
0000-0001-8910-0937ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 10 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Latent-Guided Exemplar-Based Image Re-ColorizationabstractExemplar-based re-colorization transfers colors from a reference to a colored or grayscale source image, accounting for the semantic correspondences between the two. Existing grayscale colorization methods usually predict only the chromatic aberration while maintaining the source’s luminance. Consequently, the result’s color may diverge from the reference due to such luminance difference. On the other hand, global photorealistic stylization without segmentation cannot handle scenarios where different parts of the scene need different colors. To overcome this issue, we propose a novel and effective method for re-colorization: 1) We first exploit the spatial-adaptive latent space of SpaceEdit in the context of the re-colorization task and achieve re-colorization via latent maps prediction through a proposed network. 2) We then delve into SpaceEdit’s self-reconstruct latent codes and maps to better characterize the global style and local color property, based on which we construct a novel loss to supervise re-colorization. Qualitative and quantitative results show that our method outperforms previous works by generating superior outputs with more consistent colors and global styles based on references. Ning Xu 0007 |
WACV | 2 |
| 2023 | Semantic Layout Manipulation With High-Resolution Sparse AttentionabstractWe tackle the problem of semantic image layout manipulation, which aims to manipulate an input image by editing its semantic label map. A core problem of this task is how to transfer visual details from the input images to the new semantic layout while making the resulting image visually realistic. Recent work on learning cross-domain correspondence has shown promising results for global layout transfer with dense attention-based warping. However, this method tends to lose texture details due to the resolution limitation and the lack of smoothness constraint on correspondence. To adapt this paradigm for the layout manipulation task, we propose a high-resolution sparse attention module that effectively transfers visual details to new layouts at a resolution up to 512x512. To further improve visual quality, we introduce a novel generator architecture consisting of a semantic encoder and a two-stage decoder for coarse-to-fine synthesis. Experiments on the ADE20k and Places365 datasets demonstrate that our proposed approach achieves substantial improvements over the existing inpainting and layout manipulation methods. Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Jianming Zhang 0001, Ning Xu 0007, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color EditingabstractRecently, large pretrained models (e.g., BERT, Style-GAN, CLIP) show great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified model for open-domain image editing focusing on color and tone adjustment of open-domain images while keeping their original content and structure. Our model learns a unified editing space that is more semantic, intu-itive, and easy to manipulate than the operation space (e.g., contrast, brightness, color curve) used in many existing photo editing softwares. Our model belongs to the image-to-image translation framework which consists of an image encoder and decoder, and is trained on pairs of before-and-after edited images to produce multimodal outputs. We show that by inverting image pairs into latent codes of the learned editing space, our model can be leveraged for vari-ous downstream editing tasks such as language-guided image editing, personalized editing, editing-style clustering, retrieval, etc. We extensively study the unique properties of the editing space in experiments and demonstrate superior performance on the aforementioned tasks11Code and supplementary material can be found at the project page https://jshi31.github.io/SpaceEdit. Jing Shi 0005, Ning Xu 0007, Haitian Zheng, Jiebo Luo 0001, Chenliang Xu |
CVPR | 2 |
| 2022 | Image Inpainting with Cascaded Modulation GAN and Object-Aware Training
Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Eli Shechtman, Connelly Barnes, Jianming Zhang 0001, Ning Xu 0007, Sohrab Amirghodsi, Jiebo Luo 0001 |
ECCV (16) | 8 |
| 2022 | Exploring the Semi-Supervised Video Object Segmentation Problem from a Cyclic Perspective
Yuxi Li 0009, Ning Xu 0007, John See, Weiyao Lin |
Int. J. Comput. Vis. | 2 |
| 2022 | Space-Time Memory Networks for Video Object Segmentation With User GuidanceabstractWe propose a novel and unified solution for user-guided video object segmentation tasks. In this work, we consider two scenarios of user-guided segmentation: semi-supervised and interactive segmentation. Due to the nature of the problem, available cues - video frame(s) with object masks (or scribbles) - become richer with the intermediate predictions (or additional user inputs). However, the existing methods make it impossible to fully exploit this rich source of information. We resolve the issue by leveraging memory networks and learning to read relevant information from all available sources. In the semi-supervised scenario, the previous frames with object masks form an external memory, and the current frame as the query is segmented using the information in the memory. Similarly, to work with user interactions, the frames that are given user inputs form the memory that guides segmentation. Internally, the query and the memory are densely matched in the feature space, covering all the space-time pixel locations in a feed-forward fashion. The abundant use of the guidance information allows us to better handle challenges such as appearance changes and occlusions. We validate our method on the latest benchmark sets and achieve state-of-the-art performance along with a fast runtime. Seoung Wug Oh, Joon-Young Lee, Ning Xu 0007, Seon Joo Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Learning by Planning: Language-Guided Global Image EditingabstractRecently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in interpretability. To overcome the collective difficulties, we develop a text-to-operation model to map the vague editing language request into a series of editing operations, e.g., change contrast, brightness, and saturation. Each operation is interpretable and differentiable. Furthermore, the only supervision in the task is the target image, which is insufficient for a stable training of sequential decisions. Hence, we propose a novel operation planning algorithm to generate possible editing sequences from the target image as pseudo ground truth. Comparison experiments on the newly collected MA5k-Req dataset and GIER dataset show the advantages of our methods. Code is available at https://github.com/jshi31/T2ONet. Jing Shi 0005, Ning Xu 0007, Trung Bui, Franck Dernoncourt, Chenliang Xu |
CVPR | 2 |
| 2021 | Mask Guided Matting via Progressive Refinement NetworkabstractWe propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series of guidance mask perturbation operations are also introduced in the training to further enhance its robustness to external guidance. We show that PRN can generalize to unseen types of guidance masks such as trimap and low-quality alpha matte, making it suitable for various application pipelines. In addition, we revisit the foreground color prediction problem for matting and propose a surprisingly simple improvement to address the dataset issue. Evaluation on real and synthetic benchmarks shows that MG Matting achieves state-of-the-art performance using various types of guidance inputs. Code and models are available at https://github.com/yucornetto/MGMatting. Qihang Yu, Jianming Zhang 0001, He Zhang 0004, Yilin Wang 0002, Zhe Lin 0001, Ning Xu 0007, Yutong Bai, Alan L. Yuille |
CVPR | 6 |
| 2021 | A Simple Baseline for Weakly-Supervised Scene Graph GenerationabstractWe investigate the weakly-supervised scene graph generation, which is a challenging task since no correspondence of label and object is provided. The previous work regards such correspondence as a latent variable which is iteratively updated via nested optimization of the scene graph generation objective. However, we further reduce the complexity by decoupling it into an efficient first-order graph matching module optimized via contrastive learning to obtain such correspondence, which is used to train a standard scene graph generation model. The extensive experiments show that such a simple pipeline can significantly surpass the previous state-of-the-art by more than 30% on the Visual Genome dataset, both in terms of graph matching accuracy and scene graph quality. We believe this work serves as a strong baseline for future research. Code is available at https://github.com/jshi31/WS-SGG. Jing Shi 0005, Yiwu Zhong, Ning Xu 0007, Yin Li 0003, Chenliang Xu |
ICCV | 3 |
| 2021 | Language-Guided Global Image Editing via Cross-Modal Cyclic MechanismabstractEditing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing. Existing works suffer from imbalanced and insufficient data distribution of real-world datasets and thus fail to understand language requests well. To handle this issue, we propose to create a cycle with our image generator by creating a novel model called Editing Description Network (EDNet) which predicts an editing embedding given a pair of images. Given the cycle, we propose several free augmentation strategies to help our model understand various editing requests given the imbalanced dataset. In addition, two other novel ideas are proposed: an Image-Request Attention (IRA) module which allows our method to edit an image spatial-adaptively when the image requires different editing degree at different regions, as well as a new evaluation metric for this task which is more semantic and reasonable than conventional pixel losses (e.g. L1). Extensive experiments on two benchmark datasets demonstrate the effectiveness of our method over existing approaches. Ning Xu 0007, Chen Gao 0005, Jing Shi 0005, Zhe Lin 0002, Si Liu 0001 |
ICCV | 2 |
| 2021 | End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksabstractVideo instance segmentation is a challenging task that extends image instance segmentation to the video domain. Existing methods either rely only on single-frame information for the detection and segmentation subproblems or handle tracking as a separate post-processing step, which limit their capability to fully leverage and share useful spatial-temporal information for all the subproblems. In this paper, we propose a novel graph-neural-network (GNN) based method to handle the aforementioned limitation. Specifically, graph nodes representing instance features are used for detection and segmentation while graph edges representing instance relations are used for tracking. Both inter and intra-frame information is effectively propagated and shared via graph updates and all the subproblems (i.e. detection, segmentation and tracking) are jointly optimized in an unified framework. The performance of our method shows great improvement on the YoutubeVIS validation dataset compared to existing methods and achieves 36.5% AP with a ResNet-50 backbone, operating at 22 FPS. Tao Wang 0002, Ning Xu 0007, Kean Chen, Weiyao Lin |
ICCV | 2 |
| 2020 | Finding Action Tubes with a Sparse-to-Dense FrameworkabstractThe task of spatial-temporal action detection has attracted increasing researchers. Existing dominant methods solve this problem by relying on short-term information and dense serial-wise detection on each individual frames or clips. Despite their effectiveness, these methods showed inadequate use of long-term information and are prone to inefficiency. In this paper, we propose for the first time, an efficient framework that generates action tube proposals from video streams with a single forward pass in a sparse-to-dense manner. There are two key characteristics in this framework: (1) Both long-term and short-term sampled information are explicitly utilized in our spatio-temporal network, (2) A new dynamic feature sampling module (DTS) is designed to effectively approximate the tube output while keeping the system tractable. We evaluate the efficacy of our model on the UCF101-24, JHMDB-21 and UCFSports benchmark datasets, achieving promising results that are competitive to state-of-the-art methods. The proposed sparse-to-dense strategy rendered our framework about 7.6 times more efficient than the nearest competitor. Yuxi Li 0009, Weiyao Lin, Tao Wang 0002, John See, Rui Qian 0001, Ning Xu 0007, Limin Wang 0002, Shugong Xu |
AAAI | 6 |
| 2020 | A Benchmark and Baseline for Language-Driven Image Editing
Jing Shi 0005, Ning Xu 0007, Trung Bui, Franck Dernoncourt, Chenliang Xu |
ACCV (6) | 2 |
| 2020 | Spatial Class Distribution Shift in Unsupervised Domain Adaptation: Local Alignment Comes to Rescue
Safa Cicek, Ning Xu 0007, Hailin Jin, Stefano Soatto |
ACCV (3) | 2 |
| 2020 | M2KD: Incremental Learning via Multi-model and Multi-level Knowledge Distillation
Peng Zhou 0009, Long Mai, Jianming Zhang 0001, Ning Xu 0007, Zuxuan Wu, Larry Davis 0001 |
BMVC | 4 |
| 2020 | Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation
Kenan E. Ak, Ning Xu 0007, Zhe Lin 0002, Yilin Wang 0002 |
ECCV (21) | 2 |
| 2020 | CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Yuxi Li 0009, Weiyao Lin, John See, Ning Xu 0007, Shugong Xu, Yan Ke |
ECCV (16) | 4 |
| 2020 | Multiple Sound Sources Localization from Coarse to Fine
Rui Qian 0001, Di Hu 0001, Heinrich Dinkel, Mengyue Wu, Ning Xu 0007, Weiyao Lin |
ECCV (20) | 5 |
| 2020 | Interactive Training And Architecture For Deep Object SelectionabstractInteractive object cutout tools are the cornerstone of the image editing workflow. Algorithms that can reduce the number of interactions are clearly valuable. Recent deep-learning based interactive segmentation algorithms are capable of rough binary selections with a handful of clicks, yet, they tend to plateau once this rough selection has been reached. In this work, we interpret this plateau as an inability of the algorithm to precisely leverage each user interaction.We introduce a novel interactive architecture and a training scheme that are both tailored to better exploit the user input at higher numbers of clicks. Comprehensive experiments support our approach, and our network achieves state of the art performance. Marco Forte, Brian L. Price, Scott Cohen, Ning Xu 0007, François Pitié |
ICME | 4 |
| 2020 | Delving into the Cyclic Mechanism in Semi-supervised Video Object SegmentationabstractIn this paper, we take attempt to incorporate the cyclic mechanism with the vision task of semi-supervised video object segmentation. By resorting to the accurate reference mask of the first frame, we try to mitigate the error propagation problem in most of current video object segmentation pipelines. Firstly, we propose a cyclic scheme for offline training of segmentation networks. Then, we extend the offline pipeline to an online method by introducing a simple gradient correction module while keeping high efficiency as other offline methods. Finally we develop cycle effective receptive field (cycle-ERF) from gradient correction to provide a new perspective for analyzing object-specific regions of interests. We conduct comprehensive experiments on benchmarks of DAVIS17 and Youtube-VOS, demonstrating that our introduced cyclic mechanism is helpful to boost the segmentation quality. Yuxi Li 0009, Ning Xu 0007, Jinlong Peng, John See, Weiyao Lin |
NeurIPS | 2 |
| 2019 | End-To-End Time-Lapse Video Synthesis From a Single Outdoor ImageabstractTime-lapse videos usually contain visually appealing content but are often difficult and costly to create. In this paper, we present an end-to-end solution to synthesize a time-lapse video from a single outdoor image using deep neural networks. Our key idea is to train a conditional generative adversarial network based on existing datasets of time-lapse videos and image sequences. We propose a multi-frame joint conditional generation framework to effectively learn the correlation between the illumination change of an outdoor scene and the time of the day. We further present a multi-domain training scheme for robust training of our generative models from two datasets with different distributions and missing timestamp labels. Compared to alternative time-lapse video synthesis algorithms, our method uses the timestamp as the control variable and does not require a reference video to guide the synthesis of the final output. We conduct ablation studies to validate our algorithm and compare with state-of-the-art techniques both qualitatively and quantitatively. Seonghyeon Nam, Chongyang Ma, Menglei Chai, William Brendel, Ning Xu 0007, Seon Joo Kim |
CVPR | 5 |
| 2019 | Fast User-Guided Video Object Segmentation by Interaction-And-Propagation NetworksabstractWe present a deep learning method for the interactive video object segmentation. Our method is built upon two core operations, interaction and propagation, and each operation is conducted by Convolutional Neural Networks. The two networks are connected both internally and externally so that the networks are trained jointly and interact with each other to solve the complex video object segmentation problem. We propose a new multi-round training scheme for the interactive video object segmentation so that the networks can learn how to understand the user's intention and update incorrect estimations during the training. At the testing time, our method produces high-quality results and also runs fast enough to work with users interactively. We evaluated the proposed method quantitatively on the interactive track benchmark at the DAVIS Challenge 2018. We outperformed other competing methods by a significant margin in both the speed and the accuracy. We also demonstrated that our method works well with real user interactions. Seoung Wug Oh, Joon-Young Lee, Ning Xu 0007, Seon Joo Kim |
CVPR | 3 |
| 2019 | Large-Scale Tag-Based Font Retrieval With Generative Feature LearningabstractFont selection is one of the most important steps in a design workflow. Traditional methods rely on ordered lists which require significant domain knowledge and are often difficult to use even for trained professionals. In this paper, we address the problem of large-scale tag-based font retrieval which aims to bring semantics to the font selection process and enable people without expert knowledge to use fonts effectively. We collect a large-scale font tagging dataset of high-quality professional fonts. The dataset contains nearly 20,000 fonts, 2,000 tags, and hundreds of thousands of font-tag relations. We propose a novel generative feature learning algorithm that leverages the unique characteristics of fonts. The key idea is that font images are synthetic and can therefore be controlled by the learning algorithm. We design an integrated rendering and learning process so that the visual feature from one image can be used to reconstruct another image with different text. The resulting feature captures important font design details while is robust to nuisance factors such as text. We propose a novel attention mechanism to re-weight the visual feature for joint visual-text modeling. We combine the feature and the attention mechanism in a novel recognition-retrieval model. Experimental results show that our method significantly outperforms the state-of-the-art for the important problem of large-scale tag-based font retrieval. Ning Xu 0007, Hailin Jin, Jiebo Luo 0001 |
ICCV | 3 |
| 2019 | Video Object Segmentation Using Space-Time Memory NetworksabstractWe propose a novel solution for semi-supervised video object segmentation. By the nature of the problem, available cues (e.g. video frame(s) with object masks) become richer with the intermediate predictions. However, the existing methods are unable to fully exploit this rich source of information. We resolve the issue by leveraging memory networks and learn to read relevant information from all available sources. In our framework, the past frames with object masks form an external memory, and the current frame as the query is segmented using the mask information in the memory. Specifically, the query and the memory are densely matched in the feature space, covering all the space-time pixel locations in a feed-forward fashion. Contrast to the previous approaches, the abundant use of the guidance information allows us to better handle the challenges such as appearance changes and occlussions. We validate our method on the latest benchmark sets and achieved the state-of-the-art performance (overall score of 79.4 on Youtube-VOS val set, J of 88.7 and 79.2 on DAVIS 2016/2017 val set respectively) while having a fast runtime (0.16 second/frame on DAVIS 2016 val set). Seoung Wug Oh, Joon-Young Lee, Ning Xu 0007, Seon Joo Kim |
ICCV | 3 |
| 2019 | Controllable Artistic Text Style Transfer via Shape-Matching GANabstractArtistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an important open challenge. In this paper, we present the first text style transfer network that allows for real-time control of the crucial stylistic degree of the glyph through an adjustable parameter. Our key contribution is a novel bidirectional shape matching framework to establish an effective glyph-style mapping at various deformation levels without paired ground truth. Based on this idea, we propose a scale-controllable module to empower a single network to continuously characterize the multi-scale shape features of the style image and transfer these features to the target text. The proposed method demonstrates its superiority over previous state-of-the-arts in generating diverse, controllable and high-quality stylized text. Shuai Yang 0001, Zhangyang Wang, Ning Xu 0007, Jiaying Liu 0001, Zongming Guo |
ICCV | 4 |
| 2019 | An Internal Learning Approach to Video InpaintingabstractWe propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP to video we make two important contributions. First, we show that coherent video inpainting is possible without a priori training. We take a generative approach to inpainting based on internal (within-video) learning without reliance upon an external corpus of visual data to train a one-size-fits-all model for the large space of general videos. Second, we show that such a framework can jointly generate both appearance and flow, whilst exploiting these complementary modalities to ensure mutual consistency. We show that leveraging appearance statistics specific to each video achieves visually plausible results whilst handling the challenging problem of long-term consistency. Long Mai, Hailin Jin, Ning Xu 0007, John P. Collomosse |
ICCV | 5 |
| 2019 | Learning to Trace: Expressive Line Drawing Generation from PhotographsabstractAbstract In this paper, we present a new computational method for automatically tracing high‐resolution photographs to create expressive line drawings. We define expressive lines as those that convey important edges, shape contours, and large‐scale texture lines that are necessary to accurately depict the overall structure of objects (similar to those found in technical drawings) while still being sparse and artistically pleasing. Given a photograph, our algorithm extracts expressive edges and creates a clean line drawing using a convolutional neural network (CNN). We employ an end‐to‐end trainable fully‐convolutional CNN to learn the model in a data‐driven manner. The model consists of two networks to cope with two sub‐tasks; extracting coarse lines and refining them to be more clean and expressive. To build a model that is optimal for each domain, we construct two new datasets for face/body and manga background. The experimental results qualitatively and quantitatively demonstrate the effectiveness of our model. We further illustrate two practical applications. Naoto Inoue, Daichi Ito, Ning Xu 0007, Brian L. Price, Toshihiko Yamasaki |
Comput. Graph. Forum | 3 |
| 2018 | YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
Ning Xu 0007, Yuchen Fan 0001, Jianchao Yang, Dingcheng Yue, Brian L. Price, Scott Cohen, Thomas S. Huang |
ECCV (5) | 1 |
| 2017 | Deep GrabCut for Object Selection
Ning Xu 0007, Brian L. Price, Scott Cohen, Jimei Yang, Thomas S. Huang |
BMVC | 1 |
| 2017 | Deep Image MattingabstractImage matting is a fundamental computer vision problem and has many applications. Previous algorithms have poor performance when an image has similar foreground and background colors or complicated textures. The main reasons are prior methods 1) only use low-level features and 2) lack high-level context. In this paper, we propose a novel deep learning based algorithm that can tackle both these problems. Our deep model has two parts. The first part is a deep convolutional encoder-decoder network that takes an image and the corresponding trimap as inputs and predict the alpha matte of the image. The second part is a small convolutional network that refines the alpha matte predictions of the first network to have more accurate alpha values and sharper edges. In addition, we also create a large-scale image matting dataset including 49300 training images and 1000 testing images. We evaluate our algorithm on the image matting benchmark, our testing set, and a wide variety of real images. Experimental results clearly demonstrate the superiority of our algorithm over previous methods. Ning Xu 0007, Brian L. Price, Scott Cohen, Thomas S. Huang |
CVPR | 1 |
| 2016 | Deep Interactive Object SelectionabstractInteractive object selection is a very important research problem and has many applications. Previous algorithms require substantial user interactions to estimate the foreground and background distributions. In this paper, we present a novel deep-learning-based algorithm which has much better understanding of objectness and can reduce user interactions to just a few clicks. Our algorithm transforms user-provided positive and negative clicks into two Euclidean distance maps which are then concatenated with the RGB channels of images to compose (image, user interactions) pairs. We generate many of such pairs by combining several random sampling strategies to model users' click patterns and use them to finetune deep Fully Convolutional Networks (FCNs). Finally the output probability maps of our FCN-8s model is integrated with graph cut optimization to refine the boundary segments. Our model is trained on the PASCAL segmentation dataset and evaluated on other datasets with different object classes. Experimental results on both seen and unseen objects demonstrate that our algorithm has a good generalization ability and is superior to all existing interactive object selection approaches. Ning Xu 0007, Brian L. Price, Scott Cohen, Jimei Yang, Thomas S. Huang |
CVPR | 1 |
| 2013 | Intra-and-Inter-Constraint-Based Video Enhancement Based on Piecewise Tone MappingabstractVideo enhancement plays an important role in various video applications. In this paper, we propose a new intra-and-inter-constraint-based video enhancement approach aiming to: 1) achieve high intraframe quality of the entire picture where multiple regions-of-interest (ROIs) can be adaptively and simultaneously enhanced, and 2) guarantee the interframe quality consistencies among video frames. We first analyze features from different ROIs and create a piecewise tone mapping curve for the entire frame such that the intraframe quality can be enhanced. We further introduce new interframe constraints to improve the temporal quality consistency. Experimental results show that the proposed algorithm obviously outperforms the state-of-the-art algorithms. Yuanzhe Chen, Weiyao Lin, Zhenzhong Chen 0001, Ning Xu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | A new network-based algorithm for multi-camera abnormal activity detectionabstractIn this paper, a new abnormal activity detection algorithm is proposed for multi-camera surveillance applications. The proposed algorithm models the entire scene covered by the multi-camera system as a network. In this network, each node corresponds to a segmentation of the entire scene and each edge represents the activity correlation between the corresponding segmentations. Based on this network, the proposed algorithm further models human activities as the signal transmission process in the network. Thus, abnormal activities can be detected if their 'network transmission energy' is obviously larger than the normal case. Compared with the previous methods, the proposed algorithm is more general and is flexible to handle various multi-camera scenarios and configurations. Experimental results demonstrate the effectiveness of the proposed algorithm. Weiyao Lin, Xiaokang Yang 0001, Hongxiang Li 0001, Ning Xu 0007 |
ISCAS | 5 |
| 2011 | A new Temporal-Constraint-Based algorithm by handling temporal qualities for video enhancementabstractVideo enhancement has played very important roles in many applications. However, most existing enhancement methods only focus on the spatial quality within a frame while the temporal qualities of the enhanced video are often unguaranteed. In this paper, a new algorithm is proposed for video enhancement. The proposed algorithm introduces new temporal constraints and combines them with the spatial constraints such that both the spatial and temporal qualities of the video can be improved. Two strategies are proposed for including the temporal constraints. Experimental results demonstrate the effectiveness of the proposed algorithm. Weiyao Lin, Hongxiang Li 0001, Ning Xu 0007, Lining Zhang |
ISCAS | 4 |
| 2011 | A new global-based video enhancement algorithm by fusing features of multiple region-of-interestsabstractVideo enhancement plays an important role in various video applications. It is desirable to achieve high visual quality of the entire picture where multiple region-of-interests (ROIs) within the frame can be adaptively and simultaneously enhanced. In this paper, a new global-based video enhancement algorithm is proposed. The proposed algorithm first analyzes features from different ROIs. Then, a 'global' tone mapping curve is created for the entire picture which can adaptively enhance different regions at the same time. According to the statistics of ROIs, two fusion strategies, i.e., piecewise-based and factor-based fusions, are proposed for creating the global tone mapping curve. Experimental results show that the proposed algorithm can obtain more appealing perceptual quality than the state-of-the-art algorithms. Ning Xu 0007, Weiyao Lin, Yu Zhou 0015, Yuanzhe Chen, Zhenzhong Chen 0001, Hongxiang Li 0001 |
VCIP | 1 |