VLDB 2026 Research / reviewers in the wild / expert
Bailan Feng
dblp:56/8117
· DBLP profile ↗
41ranked-venue papers
5as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TurboVSR: Fantastic Video Upscalers and Where to Find ThemabstractDiffusion-based generative models have demonstrated exceptional promise in the video super-resolution (VSR) task, achieving a substantial advancement in detail generation relative to prior methods. However, these approaches face significant computational efficiency challenges. For instance, current techniques may require tens of minutes to super-resolve a mere 2-second, 1080p video. In this paper, we present TurboVSR, an ultra-efficient diffusion-based video super-resolution model. Our core design comprises three key aspects: (1) We employ an autoencoder with a high compression ratio of 32$\times$32$\times$8 to reduce the number of tokens. (2) Highly compressed latents pose substantial challenges for training. We introduce factorized conditioning to mitigate the learning complexity: we first learn to super-resolve the initial frame; subsequently, we condition the super-resolution of the remaining frames on the high-resolution initial frame and the low-resolution subsequent frames. (3) We convert the pre-trained diffusion model to a shortcut model to enable fewer sampling steps, further accelerating inference. As a result, TurboVSR performs on par with state-of-the-art VSR methods, while being 100+ times faster, taking only 7 seconds to process a 2-second long 1080p video. TurboVSR also supports image resolution by considering image as a one-frame video. Our efficient design makes SR beyond 1080p possible, results on 4K (3648$\times$2048) image SR show surprising fine details. Zhongdao Wang, Guodongfang Zhao, Bailan Feng, Wenbo Li 0002 |
ICCV | 4 |
| 2025 | Efficient 3D Perception on Multi-Sweep Point Cloud with Gumbel Spatial PruningabstractThis paper studies point cloud perception within outdoor environments. Existing methods face limitations in recognizing objects located at a distance or occluded, due to the sparse nature of outdoor point clouds. In this work, we observe a significant mitigation of this problem by accumulating multiple temporally consecutive LiDAR sweeps, resulting in a remarkable improvement in perception accuracy. However, the computation cost also increases, hindering previous approaches from utilizing a large number of LiDAR sweeps. To tackle this challenge, we find that a considerable portion of points in the accumulated point cloud is redundant, and discarding these points has minimal impact on perception accuracy. We introduce a simple yet effective Gumbel Spatial Pruning (GSP) layer that dynamically prunes points based on a learned end-toend sampling. The GSP layer is decoupled from other network components and thus can be seamlessly integrated into existing point cloud network architectures. Extensive experiments show that our pruning strategy improves several perception algorithms in multiple tasks. Xueqian Zhang, Zhongdao Wang, Bailan Feng, Ke Xu 0001 |
ICRA | 5 |
| 2025 | Reliable and Calibrated Semantic Occupancy Prediction by Hybrid Uncertainty LearningabstractVision-centric semantic occupancy prediction plays a crucial role in autonomous driving, which requires accurate and reliable predictions from low-cost sensors. Although having notably narrowed the accuracy gap with LiDAR, there is still few research effort to explore the reliability and calibration in predicting semantic occupancy from camera. In this paper, we conduct a comprehensive evaluation of existing semantic occupancy prediction models from a reliability perspective for the first time. Despite the gradual alignment of camera-based models with LiDAR in terms of accuracy, a significant reliability gap still persists. To address this concern, we propose ReliOcc, a method designed to enhance the reliability of camera-based occupancy networks. ReliOcc provides a plug-and-play scheme for existing models, which integrates hybrid uncertainty from individual voxels with sampling-based noise and relative voxels through mix-up learning. Besides, an uncertainty-aware calibration strategy is devised to further improve model reliability in offline mode. Extensive experiments under various settings demonstrate that ReliOcc significantly enhances the reliability of learned model while maintaining the accuracy for both geometric and semantic predictions. Notably, our proposed approach exhibits robustness to sensor failures and out of domain noises during inference. Song Wang 0019, Zhongdao Wang, Wentong Li 0001, Bailan Feng, Junbo Chen, Jianke Zhu |
IJCAI | 5 |
| 2024 | OctOcc: High-Resolution 3D Occupancy Prediction with Octreeabstract3D semantic occupancy has garnered considerable attention due to its abundant structural information encompassing the entire scene in autonomous driving. However, existing 3D occupancy prediction methods contend with the constraint of low-resolution 3D voxel features arising from the limitation of computational memory. To address this limitation and achieve a more fine-grained representation of 3D scenes, we propose OctOcc, a novel octree-based approach for 3D semantic occupancy prediction. OctOcc is conceptually rooted in the observation that the vast majority of 3D space is left unoccupied. Capitalizing on this insight, we endeavor to cultivate memory-efficient high-resolution 3D occupancy predictions by mitigating superfluous cross-attentions. Specifically, we devise a hierarchical octree structure that selectively generates finer-grained cross-attentions solely in potentially occupied regions. Extending our inquiry beyond 3D space, we identify analogous redundancies within another side of cross attentions, 2D images. Consequently, a 2D image feature filtering network is conceived to expunge extraneous regions. Experimental results demonstrate that the proposed OctOcc significantly outperforms existing methods on nuScenes and SemanticKITTI datasets with limited memory consumption. Wenzhe Ouyang, Bailan Feng, Zenglin Xu |
AAAI | 3 |
| 2024 | SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionabstractVision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. However, operating on dense latent spaces introduces a cubic time and space complexity, which limits scalability in terms of perception range or spatial resolution. Existing approaches compress the dense representation using projections like Bird's Eye View (BEV) or Tri-Perspective View (TPV). Although efficient, these projections result in information loss, especially for tasks like semantic occupancy prediction. To address this, we propose SparseOcc, an efficient occupancy network inspired by sparse point cloud processing. It utilizes a lossless sparse latent representation with three key innovations. Firstly, a 3D sparse diffuser performs latent completion using spatially decomposed 3D sparse convolutional kernels. Secondly, a feature pyramid and sparse interpolation enhance scales with information from others. Finally, the transformer head is redesigned as a sparse variant. SparseOcc achieves a remarkable 74.9% reduction on FLOPs over the dense baseline. Interestingly, it also improves accuracy, from 12.8% to 14.1% mIOU, which in part can be attributed to the sparse representation's ability to avoid hallucinations on empty voxels. Pin Tang, Zhongdao Wang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma 0004 |
CVPR | 6 |
| 2024 | Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
Zhongdao Wang, Enze Xie, Bailan Feng, Ze Yuan, Ke Xu 0001, Ping Luo 0002 |
ECCV (84) | 5 |
| 2024 | OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
Zhongdao Wang, Pin Tang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma 0004 |
ECCV (20) | 6 |
| 2024 | VEON: Vocabulary-Enhanced Occupancy Prediction
Jilai Zheng, Pin Tang, Zhongdao Wang, Xiangxuan Ren, Bailan Feng, Chao Ma 0004 |
ECCV (54) | 6 |
| 2023 | AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) learns to segment objects using a single point within the object extent as supervision. Challenged by the non-negligible semantic variance between object parts, however, the single supervision point causes semantic bias and false segmentation. In this study, we propose an AttentionShift method, to solve the semantic bias issue by iteratively decomposing the instance attention map to parts and estimating fine-grained semantics of each part. AttentionShift consists of two modules plugged on the vision transformer backbone: (i) token querying for pointly supervised attention map generation, and (ii) key-point shift, which re-estimates part-based attention maps by key-point filtering in the feature space. These two steps are iteratively performed so that the part-based attention maps are optimized spatially as well as in the feature space to cover full object extent. Experiments on PASCAL VOC and MS COCO 2017 datasets show that AttentionShift respectively improves the state-of-the-art of by 7.7% and 4.8% under [email protected], setting a solid PSIS baseline using vision transformer. Mingxiang Liao, Zonghao Guo, Yuze Wang 0004, Bailan Feng, Fang Wan 0001 |
CVPR | 5 |
| 2022 | CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionabstractThe recently proposed DEtection TRansformer (DETR) achieves promising performance for end-to-end object detection. However, it has relatively lower detection performance on small objects and suffers from slow convergence. This paper observed that DETR performs surprisingly well even on small objects when measuring Average Precision (AP) at decreased Intersection-over-Union (IoU) thresholds. Motivated by this observation, we propose a simple way to improve DETR by refining the coarse features and predicted locations. Specifically, we propose a novel Coarse-to-Fine (CF) decoder layer constituted of a coarse layer and a carefully designed fine layer. Within each CF decoder layer, the extracted local information (region of interest feature) is introduced into the flow of global context information from the coarse layer to refine and enrich the object query features via the fine layer. In the fine layer, the multi-scale information can be fully explored and exploited via the Adaptive Scale Fusion(ASF) module and Local Cross-Attention (LCA) module. The multi-scale information can also be enhanced by another proposed Transformer Enhanced FPN (TEF) module to further improve the performance. With our proposed framework (named CF-DETR), the localization accuracy of objects (especially for small objects) can be largely improved. As a byproduct, the slow convergence issue of DETR can also be addressed. The effectiveness of CF-DETR is validated via extensive experiments on the coco benchmark. CF-DETR achieves state-of-the-art performance among end-to-end detectors, e.g., achieving 47.8 AP using ResNet-50 with 36 epochs in the standard 3x training schedule. Xipeng Cao, Bailan Feng, Kun Niu |
AAAI | 3 |
| 2022 | Unbiased IoU for Spherical Image Object DetectionabstractAs one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while it is challenging for spherical ones. Some existing methods utilize planar bounding boxes to represent spherical objects. However, they are biased due to the distortions of spherical objects. Others use spherical rectangles as unbiased representations, but they adopt excessive approximate algorithms when computing the IoU. In this paper, we propose an unbiased IoU as a novel evaluation criterion for spherical image object detection, which is based on the unbiased representations and utilize unbiased analytical method for IoU calculation. This is the first time that the absolutely accurate IoU calculation is applied to the evaluation criterion, thus object detection algorithms can be correctly evaluated for spherical images. With the unbiased representation and calculation, we also present Spherical CenterNet, an anchor free object detection algorithm for spherical images. The experiments show that our unbiased IoU gives accurate results and the proposed Spherical CenterNet achieves better performance on one real-world and two synthetic spherical object detection datasets than existing methods. Bin Chen 0021, Yike Ma, Bailan Feng, Chenggang Yan 0001, Qiang Zhao 0005 |
AAAI | 6 |
| 2022 | End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution
Mingxiang Liao, Fang Wan 0001, Zhenjun Han, Jialing Zou, Yuze Wang 0004, Bailan Feng, Qixiang Ye |
ECCV (9) | 7 |
| 2022 | PANDORA: A Panoramic Detection Dataset for Object with Orientation
Qiang Zhao 0005, Yike Ma, Bailan Feng, Chenggang Yan 0001 |
ECCV (8) | 6 |
| 2022 | SDTP: Semantic-Aware Decoupled Transformer Pyramid for Dense Image PredictionabstractAlthough transformer has achieved great progress on computer vision tasks, the scale variation in dense image prediction is still the key challenge. Few effective multi-scale techniques are applied in transformer and there are two main limitations in the current methods. On the one hand, self-attention module in vanilla transformer fails to sufficiently exploit the diversity of semantic information because of its rigid mechanism. On the other hand, it is difficult to build attention and interaction among different levels due to the heavy computational burden. To alleviate this problem, we first revisit multi-scale problem in dense prediction, verifying the significance of diverse semantic representation and multi-scale interaction, and exploring the adaptation of transformer to pyramidal structure. Inspired by these findings, we propose a novel Semantic-aware Decoupled Transformer Pyramid (SDTP) for dense image prediction, consisting of Intra-level Semantic Promotion (ISP), Cross-level Decoupled Interaction (CDI) and Attention Refinement Function (ARF). ISP explores the semantic diversity in different receptive space through more flexible self-attention strategy. CDI builds the global attention and interaction among different levels in decoupled space which also solves the problem of heavy computation. Besides, ARF is further added to refine the attention in transformer. Experimental results demonstrate the validity and generality of the proposed method, which outperforms the state-of-the-art by a significant margin in dense image prediction tasks. Furthermore, the proposed components are all plug-and-play, which can be embedded in other methods. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Bailan Feng, Kebin Wu, Chengwei Peng, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Image stitching via deep homography estimation
Qiang Zhao 0005, Yike Ma, Chunfeng Yao, Bailan Feng |
Neurocomputing | 5 |
| 2019 | Towards Facial De-Expression and Expression Recognition in the WildabstractPrevious work has shown that a facial expression is able to be decomposed into an expressive component and an neutral component. Many researchers have utilized this principle to implement the facial expression classification with paired laboratory-controlled datasets where each expression image has a corresponding neutral face image of the same identity. However, existing approaches cannot deal with unpaired in-the-wild datasets where most of the expression images lack their corresponding neutral counterparts. In this paper, we propose a new de-expression method based on domain transfer techniques that applies the above principle to unpaired in-the-wild datasets. Different expressions are treated as different image domains so that domain transfer techniques (such as generative adversarial networks) are able to be adopted to generate neutral face images for in-the-wild datasets. In such way, we obtain better expression recognition accuracy and alleviate the problem of lacking paired expression images. Experimental results conducted on three well-known public datasets, CK+, RAF and FER2013, validate the effectiveness of our proposed method. The results shows that our proposed approach outperforms the state-of-the-art methods on both laboratory-controlled and in-the-wild datasets. Bailan Feng |
ACII | 4 |
| 2019 | Regularization and Iterative Initialization of Softmax for Fast Training of Convolutional Neural NetworksabstractA softmax regularizer is proposed, a simple and elegant constraint on softmax weight distribution in the training process. Since the direct estimation of feature centers is neither memory efficient nor robust, the proposed regularizer utilizes the relations between feature centers and the classifier weights by adding constraints on the distances between softmax weight vectors. This apparently enlarges the distances between softmax weights to benefit the separation of different classes, and provides extra gradients for the optimization of softmax in order to speed up the training process.Furthermore, we argue that the massive amount of softmax parameters is the main cause that makes the network converge slowly, especial in the class classification tasks with large class number such as face recognition. Motivated by the analysis of the relations between deep features and softmax weights, a fast training process is presented, which splits the training into multiple stages and alternates training and initializing softmax weights for fast convergence when the class number is large. Since the softmax weights can be initialized with estimated deep feature centers, the scale of training data can be gradually increased along the stages. By this procedure, the total training computation cost can be reduced. To validate its effectiveness, our approach is applied on both face recognition and image classification tasks. It obtains comparable performance with the state-of-the-art methods while boasting a faster training process. Qiang Rao, Bailan Feng |
IJCNN | 4 |
| 2018 | A Deep Learning Based No-Reference Image Quality Assessment Model for Single-Image Super-ResolutionabstractSingle-image super-resolution (SISR) is a very important and classic problem of the computer vision community. Although a lot of SISR methods have been proposed, few studies have been conducted to address the quality assessment of SISR methods. In this paper, we proposed a deep learning based no-reference image quality assessment (NR-IQA) model for SISR. We took small patches from images to form our training set and labeled them with different scores. With the aid of well-designed architecture and training strategy, our method achieved a performance leap than state-of-the-art methods. Experimental results proved the generalizability and the effectiveness of the proposed model. Bahetiyaer Bare, Ke Li 0010, Bo Yan 0001, Bailan Feng, Chunfeng Yao |
ICASSP | 4 |
| 2018 | Face Hallucination Based on Key Parts EnhancementabstractFace hallucination aims to generate a high resolution face from a low resolution one. Generic super resolution methods can not solve this problem well, because human face has a strong structure. With the rapid development of the deep learning technique, some convolutional neural networks (CNNs) models for face hallucination emerged and achieved state-of-the-art performance. In this paper, we proposed a five-branch network based on five key parts of human face. Each branch of this network aims to generate a high resolution key part. The final high resolution face is the combination of the five branches' output. In addition, we designed a gated enhance unit (GEU) and cascade it to form our network architecture. Experimental results confirm that our method can generate pleasing high resolution faces. Ke Li 0010, Bahetiyaer Bare, Bo Yan 0001, Bailan Feng, Chunfeng Yao |
ICASSP | 4 |
| 2018 | HNSR: Highway Networks Based Deep Convolutional Neural Networks Model for Single Image Super-ResolutionabstractConvolutional neural networks (CNNs) have been widely used in computer vision community. Single image super-resolution (SISR) is a classic computer vision problem, which aims to output a high-resolution image from a low-resolution one. In recent years, CNNs-based SISR methods emerged and achieved a performance leap. In this paper, we present a highly accurate deep CNNs model for SISR. Inspired by the ideas in highway networks, we propose a highway unit and cascade highway units to ensemble our model. Furthermore, we employ structural similarity index (SSIM) as a part of loss function to enhance the accuracy of trained deep CNNs model. Experimental results show that our proposed model outperforms other state-of-the-art methods. Ke Li 0010, Bahetiyaer Bare, Bo Yan 0001, Bailan Feng, Chunfeng Yao |
ICASSP | 4 |
| 2018 | Knot Magnify Loss for Face RecognitionabstractDeep Convolutional Neural Netowrks (DCNN) have significantly improved the performance of face recognition in recent years. Softmax loss is the most widely used loss function for training the DCNN-based face recognition system. It gives the same weights to easy and hard samples in one batch, which would leads to performance gap on the quality imbalanced data. In this paper, we discover that the rare hard samples in the training dataset has become a main obstacle for training a robust face recognition model. We propose to address this problem by a new supervisor signal that pays more attention to the rare hard samples and reduces the effects of the easy samples relatively. Our proposed novel Knot Magnify (KM) loss modulates the classical softmax loss to suppress the influence of easy samples and up-weight the loss of hard samples during training. Our results show that after training with KM loss, face recognition model is able to get competing accuracy on the well-known face recognition benchmark LFW dataset and the challenging CFP dataset. Qiang Rao, Bailan Feng |
ICIP | 4 |
| 2017 | Micro-Expression Recognition by Aggregating Local Spatio-Temporal Patterns
Bailan Feng, Zhineng Chen, Xiangsheng Huang |
MMM (1) | 2 |
| 2015 | Improving Automatic Name-Face Association using Celebrity Images on the WebabstractThis paper investigates the task of automatically associating faces appearing in images (or videos) with their names. Our novelty lies in the use of celebrity Web images to facilitate the task. Specifically, we first propose a method named Image Matching (IM), which uses the faces in images returned from name queries over an image search engine as the gallery set of the names, and a probe face is classified as one of the names, or none of them, according to their matching scores and compatibility characterized by a proposed Assigning-Thresholding (AT) pipeline. Noting IM could provide guidance for association for the well-established Graph-based Association (GA), we further propose two methods that jointly utilize the two kinds of complementary cues. They are: the early fusion of IM and GA (EF-IMGA) that takes the IM score as an additional information source to help the association in GA, and the late fusion of IM and GA (LF-IMGA) that combines the scores from both IM and GA obtained individually to make the association. Evaluations on datasets of captioned news images and Web videos both show the proposed methods, especially the two fused ones, provide significant improvements over GA. Zhineng Chen, Bailan Feng, Chong-Wah Ngo, Caiyan Jia, Xiangsheng Huang |
ICMR | 2 |
| 2014 | Chinese Image Character Recognition Using DNN and Machine Simulated Training Samples
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002 |
ICANN | 3 |
| 2014 | Chinese Image Text Recognition on grayscale pixelsabstractThis paper presents a novel scheme for Chinese text recognition in images and videos. It's different from traditional paradigms that binarize text images, fed the binarized text to an OCR engine and get the recognized results. The proposed scheme, named grayscale based Chinese Image Text Recognition (gCITR), implements the recognition directly on grayscale pixels via the following steps: image text over-segmentation, building recognition graph, Chinese character recognition and beam search determination. The advantages of gCITR lie in: (1) it does not heavily rely on the performance of binarization, which is not robust in practical and thus severely affects the performance of OCR, (2) grayscale image retains more information of the text thus facilitates the recognition. Experimental results on text from 13 TV news videos demonstrate the effectiveness of the proposed gCITR, from which significant performance gains are observed. Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002 |
ICASSP | 3 |
| 2014 | Image character recognition using deep convolutional neural network learned from different languagesabstractThis paper proposes a shared-hidden-layer deep convolutional neural network (SHL-CNN) for image character recognition. In SHL-CNN, the hidden layers are made common across characters from different languages, performing a universal feature extraction process that aims at learning common character traits existed in different languages such as strokes, while the final softmax layer is made language dependent, trained based on characters from the destination language only. This paper is the first attempt to introduce the SHL-CNN framework to image character recognition. Under the SHL-CNN framework, we discuss several issues including architecture of the network, training of the network, from which a suitable SHL-CNN model for image character recognition is empirically learned. The effectiveness of the learned SHL-CNN is verified on both English and Chinese image character recognition tasks, showing that the SHL-CNN can reduce recognition errors by 16–30% relatively compared with models trained by characters of only one language using conventional CNN, and by 35.7% relatively compared with state-of-the-art methods. In addition, the shared hidden layers learned are also useful for unseen image character recognition tasks. Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002 |
ICIP | 3 |
| 2014 | CeleLabel: an interactive system for annotating celebrities in web videosabstractManual annotation of celebrities in Web videos is an essential task in many people-related Web services. The task, however, poses a significant challenge even to skillful annotators, mainly due to the large quantity of unfamiliar and greatly varied celebrities, and the lack of a customized system for it. This work develops CeleLabel, an interactive system for manually annotating celebrities in the Web video domain. The peculiarity of CeleLabel is to exploit and display multiple types of information that could assist the annotation, including video content, context surrounding and within a video, celebrity images on the Web, and human factors. Using the system, annotators can interactively switch between two views, i.e., merging similar faces and labeling faces with names, to approach the annotation. User studies show that the CeleLabel leads to a much better labeling efficiency and satisfaction. Zhineng Chen, Jinfeng Bai, Chong-Wah Ngo, Bailan Feng, Bo Xu 0002 |
ACM Multimedia | 4 |
| 2014 | Spatial Similarity Measure of Visual Phrases for Image Retrieval
Jiansong Chen, Bailan Feng, Bo Xu 0002 |
MMM (2) | 2 |
| 2014 | Video to Article Hyperlinking by Multiple Tag Property Exploration
Zhineng Chen, Bailan Feng, Hongtao Xie 0001, Rong Zheng 0005, Bo Xu 0002 |
MMM (1) | 2 |
| 2014 | Multiple style exploration for story unit segmentation of broadcast news video
Bailan Feng, Zhineng Chen, Rong Zheng 0005, Bo Xu 0002 |
Multim. Syst. | 1 |
| 2013 | Multi-modal topic unit segmentation in videos using conditional random fieldsabstractIn this paper a novel approach of video segmentation into topic units is presented. This approach is built upon the design in which topic unit segmentation is transformed into label identification problem by defining four types of shots that reveal semantic structure of it. To implement our algorithm, four middle-level features including shot difference signal, scene transition graph, shot theme and audio type are extracted to depict the label properties of each shot, and then CRFs model is employed to identify the labels sequence. CRFs model integrates context information, so it produces accurate results in topic unit segmentation. The proposed approach is verified by two types of data: documentary and news. Experiments on testing data set yield average 86% F-measure, which illustrates that the proposed method can accurately detect most topic units in different genres of programs. Bailan Feng, Bo Xu 0002 |
ICASSP | 2 |
| 2013 | A general Framework of video segmentation to logical unit based on conditional random fieldsabstractSegmenting video into logical units like scenes in movies and topic units in News videos is an essential prerequisite for a wide range of video related applications. In this paper, a novel approach for logical unit segmentation based on conditional random fields (CRFs) is presented. In comparison with previous approaches that handle scenes and topic units separately, the proposed approach deals with them in a general framework. Specifically, four types of shots are defined and represented by four middle-level features, i.e., shot difference, scene transition, shot theme and audio type. Then, the problem of logical unit segmentation is novelly formulated as a problem of identifying the type of shot based on the extracted features, by leveraging the CRFs model. The proposed framework effectively integrate visual, audio and contextual features, and it is able to produce ideal result for both scene and topic unit segmentation. The effectiveness of the proposed approach is verified on seven mainstream types of videos, from which average F-measures of 88% and 86% on scenes and topic units are reported respectively, illustrating that the proposed method can accurately segment logical units in different genres of videos. Bailan Feng, Zhineng Chen, Bo Xu 0002 |
ICMR | 2 |
| 2013 | Temporal Video Segmentation to Scene Based on Conditional Random Fileds
Bailan Feng, Bo Xu 0002 |
MMM (2) | 2 |
| 2013 | Fusion of Audio-Visual Features and Statistical Property for Commercial Segmentation
Bailan Feng, Bo Xu 0002 |
MMM (1) | 2 |
| 2012 | Multi-modal information fusion for news story segmentation in broadcast videoabstractWith the fast development of high-speed network and digital video recording technologies, broadcast video has been playing a more and more important role in our daily life. In this paper, we propose a novel news story segmentation scheme which can segment broadcast video into story units with multi-modal information fusion (MMIF) strategy. Compared with traditional methods, the proposed scheme extracts a wealth of semantic-level features including anchor person, topic caption, face, silence, acoustic change, audio keywords and textual content. Parallel to this, we make use of a multi-modal information fusion strategy for news story boundary characterization by joining these visual, audio and textual cues. Encouraging experimental results on News Vision dataset demonstrate the effectiveness of the proposed scheme. Bailan Feng, Peng Ding 0003, Jiansong Chen, Jinfeng Bai, Bo Xu 0002 |
ICASSP | 1 |
| 2012 | Graph-based multi-modal scene detection for movie and teleplayabstractAutomatic scene detection is a fundamental step for efficient video searching and browsing. This paper presents our current work on scene detection that integrates three effective strategies into a single framework. For each video, firstly, a coherence signal is constructed by graph modal obtained from the similarity matrix in a temporal interval. Secondly, the signal is optimized by scene transition graph (STG) analysis and audio classification, in which scene clues hidden in multimedia are discovered from the video. Finally, the scene boundaries are identified by window function. In experiments, we compare the proposed scene detection method with three typical algorithms on teleplay and movies, and the results of our method, yielding an average 0.85 F-measure, is the best one. Bailan Feng, Peng Ding 0003, Bo Xu 0002 |
ICASSP | 2 |
| 2012 | Effective near-duplicate image retrieval with image-specific visual phrase selectionabstractNear-duplicate image retrieval (NDIR) is an important topic for many applications such as multimedia content management, copyright infringement identification et al. In this work we propose a novel NDIR framework based on visual phrase. Compared with previous researches, this paper first introduces a spatial visual phrase (SVP) model enabling to capture relative geometry information between visual words. Then, it proposes an image-specific strategy to select descriptive SVPs. The strategy can not only handle the phrase sparseness problem which occurs in traditional selection strategy but also allow to select visual phrases according to the characteristic of each image. Experiments are carried out over Ukbench dataset and TRECVID dataset respectively, and encouraging experimental results demonstrate that both the SVP model and the selection strategy significantly improve the overall performance. Jiansong Chen, Bailan Feng, Peng Ding 0003, Bo Xu 0002 |
ICIP | 2 |
| 2011 | A Robust Approach to Mining Repeated Sequence in Audio Stream
Jiansong Chen, Bailan Feng, Peng Ding 0003, Bo Xu 0002 |
INTERSPEECH | 3 |
| 2011 | Graph-based multi-space semantic correlation propagation for video retrieval
Bailan Feng, Juan Cao 0001, Xiuguo Bao, Yongdong Zhang 0001, Shouxun Lin, Xiao-chun Yun |
Vis. Comput. | 1 |
| 2010 | Multi-modal query expansion for web video searchabstractQuery expansion is an effective method to improve the usability of multimedia search. Most existing multimedia search engines are able to automatically expand a list of textual query terms based on text search techniques, which can be called textual query expansion (TQE). However, the annotations (title and tag) around web videos are generally noisier for text-only query expansion and search matching. In this paper, we propose a novel multi-modal query expansion (MMQE) framework for web video search to solve the issue. Compared with traditional methods, MMQE provides a more intuitive query suggestion by transforming tex-tual query to visual presentation based on visual clustering. Paral-lel to this, MMQE can enhance the process of search matching with strong pertinence of intent-specific query by joining textual, visual and social cues from both metadata and content of videos. Experimental results on real web videos from YouTube demon-strate the effectiveness of the proposed method. Bailan Feng, Juan Cao 0001, Zhineng Chen, Yongdong Zhang 0001, Shouxun Lin |
SIGIR | 1 |
| 2009 | Motion region-based trajectory analysis and re-ranking for video retrievalabstractEvent-related query is playing a more and more important role in video retrieval. However, it is still a challenge to the existing video retrieval engines for lacking the effective motion analysis. In this paper, we propose a novel re-ranking scheme for video retrieval based on motion region trajectory analysis. By focusing on the changes of the primary moving regions, we construct an intuitive motion region-based trajectory descriptor (MRTD) to depict the shot activities. In the re-ranking phase, the proposed approach takes the MRTD as a motion cue and re-ranks the baseline results by motion-related query selection and MRTD-based weighting. We evaluate our method in TRECVID2007 and 2008 datasets, and observe consistent improvement over all the baselines, leading to a greatest performance gain of 42.9%, and an average gain of 17%. The experiments also show that the motion descriptor of MRTD is fruitful for a variety of features. Bailan Feng, Juan Cao 0001, Shouxun Lin, Yongdong Zhang 0001, Kun Tao |
ICME | 1 |