EDBT 2026 Demo / reviewers in the wild / expert
Feng Su
dblp:65/5133
· DBLP profile ↗
64ranked-venue papers
14as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 27 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 1 since 2021Systems, architecture and hardware · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scene Text Image Super-Resolution with Visual Text Cues Transfer and EnhancementabstractScene text image super-resolution (STISR) aims to improve the visual clarity of the text in low-resolution scene images. Due to the intrinsic lack of detailed text appearance information in the low-resolution input image, the text images generated by most STISR methods often contain varying degrees of distortion or loss of text details. To improve the quality of the super-resolved text image, we propose a novel Visual Text Cues Transfer-based scene text image Super-Resolution Network (VTCTSRN). The network introduces synthetic text prototypes to provide high-resolution, supplementary visual cues of the text to be reconstructed, and leverages a novel visual text cues transfer mechanism to adaptively fuse complementary visual characteristics of the prototype and the input text to help recover clear and accurate text details. Additionally, we propose dynamic attentional sequential recurrent blocks for the text reconstruction pipeline and introduce effective loss function term to enhance text representation, leading to a marked improvement in the quality of generated text images. Experimental results demonstrate that the proposed network significantly enhances the readability of low-resolution text images and establishes new state-of-the-art performance on the mainstream STISR benchmark. Zeming Zhuang, Feng Su |
ICME | 3 |
| 2024 | MPGTSRN: Scene Text Image Super-Resolution Guided by Multiple Visual-Semantic Prompts
Zeming Zhuang, Feng Su |
ICPR (32) | 4 |
| 2024 | Arbitrary-Shaped Scene Text Recognition with Deformable Ensemble Attention
Zeming Zhuang, Feng Su |
ICPR (31) | 4 |
| 2024 | AdaGC: A Novel Adaptive Optimization Algorithm with Gradient Bias CorrectionabstractA proper optimization algorithm is very important for solving the parameters of neural networks . People always want to train neural networks as fast as possible and obtain optimal parameters, while existing optimization algorithms are still far from perfect in both efficiency and convergence. In this paper, we propose an Adaptive optimization algorithm with Gradient bias Correction (AdaGC) for training neural networks . In this algorithm, the iterative direction is improved utilizing the gradient deviation and momentum. The step size is adaptively revised using the second-order moment of gradient deviation. Intuitively, when the iterative vector changes significantly, the iterative vector is corrected by enlarging the effect of iterative deviation and reducing the effect of momentum. At the same time, the step size will decrease correspondingly based on the second-order moment of gradient deviation. When the iterative vector changes slightly, the gradient deviation also decreases, and the influence of momentum will increase at this time. The second-order moment of gradient deviation is correspondingly reduced, resulting in a large step size. In this case, a large iterative step is probably to reduce the iterative time and enlarge the rate of convergence . Furthermore, rigorous theoretical analysis is provided to prove the convergence of the AdaGC algorithm in both convex and non-convex optimization problems. The proposed algorithm is also validated through extensive experimentation. It has better performance and faster convergence for different networks and applications. Code is available at https://github.com/breeze7-opt/optimizer . Qi Wang 0101, Feng Su, Shipeng Dai |
Expert Syst. Appl. | 2 |
| 2023 | Learning and Fusing Multi-Scale Representations for Accurate Arbitrary-Shaped Scene Text RecognitionabstractScene text in natural images carries a wealth of valuable semantic information, while due to the largely varied appearance of the text, accurately recognizing scene text is a challenging task. In this work, we propose an arbitrary-shaped scene text recognition method based on learning and fusing multiple representations of text in the scale space with attention mechanisms. Specifically, as distinctive visual features of text often appear at different scales, given an input text image, we generate a family of multi-scale representations that capture complementary appearance characteristics of the text through multiple encoder branches with progressively increasing scale parameters. We further introduce edge map features as a supplementary high-frequency representation with useful text cues. We then refine the multi-scale representations with in-scale and cross-scale attention mechanisms and adaptively aggregate them into an enhanced representation of the text, which effectively improves the text recognition accuracy. The proposed text recognition method achieves competitive results on several scene text benchmarks, demonstrating its effectiveness in recognizing text of various shapes. Feng Su |
ICMR | 3 |
| 2023 | Masked Swin Transformer Unet for Industrial Anomaly DetectionabstractThe intelligent detection process for industrial anomalies employs artificial intelligence methods to classify images that deviate from a normal appearance. Traditional convolutional neural network (CNN)-based anomaly detection algorithms mainly use the network to restructure abnormal areas and detect anomalies by calculating the errors between the original image and reconstructed image. However, the traditional CNNs struggle to extract global context information, resulting in poor anomaly detection performance. Thus, a masked Swin Transformer Unet (MSTUnet) for anomaly detection is proposed. To solve the problem of insufficient abnormal samples in the training phase, an anomaly simulation and mask strategy is first applied on anomaly-free samples to generate a simulated anomaly and, then, the Swin Transformer's powerful global learning ability is used to inpaint the masked area. Finally, a convolution-based Unet network is used for end-to-end anomaly detection. Experimental results on industrial dataset MVTec AD show that MSTUnet achieves superior anomaly detection and localization performance. Jielin Jiang, Muhammad Bilal 0003, Yan Cui 0007, Neeraj Kumar 0001, Ruihan Dou, Feng Su, Xiaolong Xu 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2022 | Reading Arbitrary-Shaped Scene Text from Images Through Spline Regression and Rectification
Feng Su, Ye Qian |
ACCV (5) | 2 |
| 2022 | Towards Robust Video Text Detection with Spatio-Temporal Attention Modeling and Text Cues FusionabstractInformation carried by video text is of great value to various video applications. However, detecting text in videos of- ten faces great challenges due to the widely varied appearance of text and the complicated, dynamic video context. In this paper, we propose a robust video text detection network that adaptively combines relevant text cues in multiple frames with spatio-temporal attention and fusion mechanisms, which effectively enhance the accuracy and robustness of video text detection compared to single-frame detection. The network first localizes text region proposals and propagates them across frames with an R-CNN based framework. Then, a Transformer-based cross-frame feature fusion model is employed to attentively select and combine relevant text features, yielding an enhanced representation of text region integrating complementary text cues for robust text candidate prediction. The network achieves competitive text detection performance on standard video text benchmarks, demonstrating the effectiveness of the proposed method. Feng Su |
ICME | 2 |
| 2021 | An Adaptive Rectification Model for Arbitrary-Shaped Scene Text Recognition
Ye Qian, Feng Su |
BMVC | 3 |
| 2021 | Robust Video Text Detection Through Parametric Shape Regression, Propagation and FusionabstractScene text in one video carries valuable semantic information for various video applications. The varied and complex text appearance and video context, however, make reliable detection of scene text in the video a challenging task. In this paper, we propose a novel end-to-end video text detection frame-work with an effective proposal-level text information propagation and fusion mechanism for robust detection of video text. Specifically, on the basis of a parametric shape representation and regression model for intra-frame text detection and an integrated cross-frame text region propagation mechanism, we correlate corresponding text candidates in adjacent frames and accordingly propagate a text candidate’s shape parameters and features in the previous frame to the current frame as supplementary text cues, which are then attentively fused with those of the current frame for improved text detection results. Experiment results on standard benchmarks demonstrate the effectiveness of our method for robustly detecting video text with widely varied appearances. Feng Su |
ICME | 3 |
| 2020 | Accurate Arbitrary-Shaped Scene Text Detection via Iterative Polynomial Parameter Regression
Feng Su |
ACCV (3) | 3 |
| 2020 | Robust Scene Text Recognition Through Adaptive Image Enhancement
Ye Qian, Feng Su |
BMVC | 3 |
| 2020 | A Deep Convolutional Deblurring and Detection Neural Network for Localizing Text in Videos
Ye Qian, Feng Su |
MMM (2) | 4 |
| 2019 | Text-Attentional Conditional Generative Adversarial Network for Super-Resolution of Text ImagesabstractText in natural scene images are often faced with low-resolution problem, which brings significant difficulties to many text-related tasks such as text detection and recognition. In this paper, we propose a novel text-attentional Conditional Generative Adversarial Network (cGAN) model for text image super-resolution (SR). The model enhances the original cGAN by introducing effective channel and spatial attention mechanisms based on the proposed Residual Dense Channel Attention Block and text/non-text segmentation information, which focus the model on the text regions instead of the background of the image to learn more effective representations of text and achieve better text super-resolution result. The proposed model achieves state-of-the-art performances on public text image super-resolution dataset. Feng Su, Ye Qian |
ICME | 2 |
| 2019 | Video Text Detection with Fully Convolutional Network and TrackingabstractScene text in videos carries rich semantic information that is of great value in various content-based video applications. In this paper, we propose an effective fully convolutional network model for detecting text in videos based on a novel refine block structure. The model hierarchically exploits low-level features from earlier convolutions to refine high-level semantic features, thereby fusing multi-resolution features extracted from the frame to generate high-resolution semantic feature maps for better capturing widely varied appearances of video text. We further complement the individual-frame detection with an efficient correlation filter based text tracking mechanism, and enhance the overall detection performance by matching and combining detection and tracking results. Experiments on public scene text video datasets demonstrate the state-of-the-art performance of the proposed method. Feng Su |
ICME | 3 |
| 2019 | Procedural Sound Generation for Soft Bodies in Video GamesabstractWe propose a real-time, data-driven method to synthesize the sound of typical soft bodies (such as cloth and rope) in video games. For a given soft body, we perform a geometric analysis of its shape at each frame. These data are used to calculate a variety of motion events at run time that would produce sounds. We then record a database of sounds, which contains sequences of segmented sound units. Finally, we use a concatenative synthesis method to select and synthesize the actual soft-body sounds according to the extracted motion signals. Our approach is more computationally efficient compared to existing soft-body sound synthesis methods and is compatible with any particle-based soft-body physics in video games. We implement our method in Unreal Engine 4 and demonstrate its efficiency with several examples. Feng Su, Chris Joslin |
MIG | 1 |
| 2019 | A Hierarchical Attentive Deep Neural Network Model for Semantic Music Annotation Integrating Multiple Music RepresentationsabstractAutomatically assigning a group of appropriate semantic tags to one music piece provides an effective way for people to efficiently utilize the massive and ever increasing on-line and off-line music data. In this paper, we propose a novel content-based automatic music annotation model that hierarchically combines attentive convolutional networks and recurrent networks for music representation learning, structure modelling and tag prediction. The model first exploits two separate attentive convolutional networks composed of multiple gated linear units (GLUs) to learn effective representations from both 1-D raw waveform signals and 2-D Mel-spectrogram of the music, which better captures informative features of the music for the annotation task than exploiting any single representation channel. The model then exploits bidirectional Long Short-Term Memory (LSTM) networks to depict the time-varying structures embedded in the description sequences of the music, and further introduces a dual-state LSTM network to encode temporal correlations between two representation channels, which effectively enriches the descriptions of the music. Finally, the model adaptively aggregates music descriptions generated at every time step with a self-attentive multi-weighting mechanism for music tag prediction. The proposed model achieves state-of-the-art results on the public MagnaTagATune music dataset, demonstrating its effectiveness on music annotation. Feng Su |
ICMR | 2 |
| 2019 | Video Text Detection by Attentive Spatiotemporal Fusion of Deep Convolutional FeaturesabstractScene text in videos carries rich semantic information and plays an important role in various content-based video applications. Compared to text in static images, scene text in videos exhibits some distinct characteristics such as motion blur and temporal redundancy, which bring additional difficulties as well as exploitable clues to the text detection task. In this paper, we propose a novel end-to-end deep neural network for detecting scene text in the video, which combines complementary text features from multiple related frames to enhance the overall detection performance relative to single-frame detection schemes. Specifically, we first extract descriptive features from each video frame using a hierarchical convolutional neural network. Next, we spatiotemporally sample and warp supplementary features from adjacent frames surrounding the current frame using a multi-scale deformable convolution structure. We then aggregate the sampled features with an attention mechanism to adaptively focus on and augment relevant features and generate an enhanced feature representation of the current frame, which is further fed to the prediction network for localizing text candidates. The proposed model achieves state-of-the-art text detection performance on public scene text video datasets, demonstrating the superiority of the proposed multi-frame feature fusion based video text detection scheme to most single-frame and tracking-based detection schemes. Feng Su |
ACM Multimedia | 4 |
| 2019 | Efficient implementation of convolutional neural networks in the data processing of two-photon in vivo imagingabstractMOTIVATION: Functional imaging at single-neuron resolution offers a highly efficient tool for studying the functional connectomics in the brain. However, mainstream neuron-detection methods focus on either the morphologies or activities of neurons, which may lead to the extraction of incomplete information and which may heavily rely on the experience of the experimenters. RESULTS: We developed a convolutional neural networks and fluctuation method-based toolbox (ImageCN) to increase the processing power of calcium imaging data. To evaluate the performance of ImageCN, nine different imaging datasets were recorded from awake mouse brains. ImageCN demonstrated superior neuron-detection performance when compared with other algorithms. Furthermore, ImageCN does not require sophisticated training for users. AVAILABILITY AND IMPLEMENTATION: ImageCN is implemented in MATLAB. The source code and documentation are available at https://github.com/ZhangChenLab/ImageCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yangzhen Wang, Feng Su, Chaojuan Yang, Yonglu Tian, Peijiang Yuan |
Bioinform. | 2 |
| 2018 | Semantic Music Annotation by Label-Specific Conditional Random FieldsabstractMusic annotation is the task automatically assigning a set of semantically meaningful text labels to a music piece, which is of great value to many variant music applications such as music searching, indexing, recommendation and management. In this paper, we propose a novel music annotation method that integrates feature-to-label correspondence, label smoothness and local-to-global annotation consistency in a conditional random field (CRF) model with label-specific feature learning. For one music piece to be annotated, we first divide the music into a set of acoustically homogeneous segments and infer the relevant labels of every music segment using the CRF models corresponding to respective labels. These local annotations are then aggregated to obtain the holistic annotation of the music. Experiments on the public CAL500 music annotation dataset demonstrate the effectiveness of the proposed method. Feng Su |
ICPR | 3 |
| 2018 | Scene Text Detection and Tracking in Video with Background CuesabstractTo detect scene text in the video is valuable to many content-based video applications. In this paper, we present a novel scene text detection and tracking method for videos, which effectively exploits the cues of the background regions of the text. Specifically, we first extract text candidates and potential background regions of text from the video frame. Then, we exploit the spatial, shape and motional correlations between the text and its background region with a bipartite graph model and the random walk algorithm to refine the text candidates for improved accuracy. We also present an effective tracking framework for text in the video, making use of the temporal correlation of text cues across successive frames, which contributes to enhancing both the precision and the recall of the final text detection result. Experiments on public scene text video datasets demonstrate the state-of-the-art performance of the proposed method. Susu Shan, Feng Su |
ICMR | 4 |
| 2017 | Text Proposals Based on Windowed Maximally Stable Extremal Region for Scene Text DetectionabstractThe generation of text proposals (i.e. local candidate regions most likely containing textual components) is one critical and prerequisite step in scene text detection task. As one popular text proposal algorithm, the Maximally Stable Extremal Region (MSER), has been exploited by many successful text detection methods, while on the other hand has difficulties in handling complicated scene text involving touching characters and characters composed of multiple unconnected parts (e.g. Chinese characters and text in dot matrix fonts). In this paper, we propose a novel text proposal method for localizing text in natural images, which integrates the MSER algorithm with the multi-scale sliding window framework and efficiently extracts Windowed Maximally Stable Extremal Regions (WMSERs) as text proposals. We further present effective proposal filtering and grouping algorithms for exploiting WMSER-based proposals in text detection task. Experiments on public scene text datasets demonstrate the promising aspects of the proposed method in dealing with complicated scene text. Feng Su, Wenjun Ding, Susu Shan, Hailiang Xu |
ICDAR | 1 |
| 2017 | Text detection in natural scene images by hierarchical localization and growing of textual componentsabstractText embedded in natural scene images provide rich semantic information about the scene, which is of great value for content-based image applications. Due to the variety of text appearance and the complexity of scene context, however, text detection in natural images remains a challenging task. In this paper, we propose a robust text detection method that hierarchically and progressively localizes textual components at pixel, intra-character and inter-character levels. For each level, a seed growing mechanism is adopted, which starts by detecting well-conditioned seed textual components and then grows from the seeds to localize related degraded components, exploiting the cues captured by the seeds. We further propose a random walk with restart algorithm to robustly aggregate character candidates into text lines. The experiment on public scene text datasets demonstrates the state-of-the-art performance of the proposed method. Wenjun Ding, Susu Shan, Feng Su |
ICME | 3 |
| 2017 | Automatic music mood classification by learning cross-media relevance between audio and lyricsabstractAutomatic analysis of the mood of a piece of music is of great value in music searching, understanding, recommendation and some other music-related applications. Different from most of previous methods that adopted a discriminative mood classification scheme, in this paper, we propose a generative multimodal method for automatically classifying the mood of a piece of music based on effective learning of the relevance (i.e. the joint distribution) between the audio and the lyrics modalities of music. We present effective algorithms for computing the word-to-audio and word-to-word relations in the music as well as a priori probability of lyrics words, which altogether form the multimodal joint distribution that distinctively captures the intrinsic characteristic of one specific music mood. A piece of music is then classified to the mood category that maximizes this joint probability of different modalities of music data. The experiment results demonstrated the effectiveness of the proposed method. Feng Su |
ICME | 2 |
| 2017 | Robust Scene Text Detection for Multi-script Languages Using Deep Learning
Ruo-Ze Liu, Xin Sun 0009, Hailiang Xu, Palaiahnakote Shivakumara, Feng Su, Tong Lu 0002, Ruoyu Yang |
MMM (1) | 5 |
| 2017 | Graph-Based Multimodal Music Mood Classification in Discriminative Latent Space
Feng Su |
MMM (1) | 1 |
| 2016 | A new method for spatiotemporal textual saliency detection in videoabstractTo detect salient image regions containing textual patterns in the video is valuable to many content-based video applications such as video retrieval, abstraction, classification and analysis. In this paper, we present an effective textual saliency detection method for natural scene videos. We first compute text-alike confidence values of local image regions, which capture the basic visual cues of textual components in the video frames, using an efficient cascaded prediction model. Next, we construct patch features depicting the statistical and spatial distribution of confidence values and combine them with general visual features like colors. We then employ a saliency detection model based on random walk with restart on the graph of local video regions, which effectively integrates both the spatial and the temporal saliency maps. The experiment result demonstrates the effectiveness of the proposed method. Susu Shan, Hailiang Xu, Feng Su |
ICPR | 3 |
| 2015 | Diving Head-First into Virtual Reality: Evaluating HMD Control Schemes for VR Games
Erin Martel, Feng Su, Jesse Gerroir, Audrey Girouard, Kasia Muldner |
FDG | 2 |
| 2015 | Robust seed-based stroke width transform for text detection in natural imagesabstractText detection in natural scene images is challenging due to the significant variations of the appearance of the text itself and its interaction with the context. The popular stroke width transform (SWT) algorithm is highly efficient but sensitive to the defects of the edges extracted from the input image when searching for the matching edge pixels for computing potential stroke width. In this paper, we propose a novel seed-based variant of SWT that enhances significantly the robustness of the original algorithm to complicated image contextual interference and varied text appearance. We first search for the seed segment of strokes, which is defined as a consecutive sequence of neighbouring rays (pairs of edge pixels) with regular length and satisfying certain constraints, and grow from them to localize more stroke segments. We then exploit the principal width and direction information of stroke captured by the stroke segments detected to rectify inaccurate stroke width and recover missed stroke parts, which are resulted from erroneous and noisy edges in complex natural images. The stroke segments detected are also exploited to improve the accuracy of candidate character localization. The experimental results on public datasets demonstrated the effectiveness of the proposed method. Feng Su, Hailiang Xu |
ICDAR | 1 |
| 2015 | A robust hierarchical detection method for scene text based on convolutional neural networksabstractDetecting the text in natural scene images is often challenging due to the complexity and variety of text's appearance and its interaction with the scene context. In this paper, we present a novel hierarchical text detection method exploiting textual characteristics at both character and text line scales for improved accuracy. First, seed candidate characters are detected with discriminative deep convolutional features learned within the maximally stable extremal regions extracted from the image, and are further grown to localize other degraded candidate characters. Next, as a finer filtering of text in the richer text line context, the random forest classifier is exploited on statistical features of text line characterizing the geometrical and conformability properties of constituent character components, to predict the text and non-text label. The effectiveness of the proposed method is demonstrated by the state-of-the-art results achieved on the public datasets. Hailiang Xu, Feng Su |
ICME | 2 |
| 2015 | Graph Learning on K Nearest Neighbours for Automatic Image AnnotationabstractImage annotation is an open and challenging task, especially with large label vocabulary. In this paper, we propose a novel graph learning based method for image annotation, which takes both advantages of the nearest neighbour based and the graph-based methods, by exploiting the graph learning method to propagate the labels on the graph corresponding to the K nearest neighbours of a test image. To acquire more effective graph weights for computing score for each label, besides the similarity of visual features, our method also considers the similarity of two label sets, which is computed based on the label correlation that captures the semantic information between two labels. In addition, we combine the image-to-label distance with the graph learning based score to compute the final decision value for labelling. The proposed method is evaluated on three benchmark datasets for image annotation. The result shows our method substantially outperforms the previous graph learning based methods, and our result matches the current state-of-the-art results in annotation quality. Feng Su, Like Xue |
ICMR | 1 |
| 2015 | Robust Seed Localization and Growing with Deep Convolutional Features for Scene Text DetectionabstractText detection in natural scene images is an open and challenging problem due to the significant variations of the appearance of the text itself and its interaction with the context. In this paper, we present a novel text detection method based on robust localization and adaptive growing of seed text components. The method consists of two main ingredients. First, convolutional neural network is exploited to localize seed candidate characters from the maximally stable extremal regions of the image with learned discriminative deep convolutional features. Next, an iterative and adaptive growing algorithm is employed to grow from seed characters to search for other degraded text components in same text line based on their conformity to the seed, and an associative quality is learned to measure the conformity combining both the geometric and appearance constraints between two neighbouring text components. The effectiveness of the proposed method is demonstrated by the state-of-the-art results achieved on the public datasets. Hailiang Xu, Feng Su |
ICMR | 2 |
| 2015 | Auditory Scene Classification with Deep Belief Network
Like Xue, Feng Su |
MMM (1) | 2 |
| 2015 | Multimodal Music Mood Classification by Fusion of Audio and Lyrics
Like Xue, Feng Su |
MMM (2) | 3 |
| 2015 | Context-based environmental audio event recognition for scene understanding
Gongyou Wang, Feng Su |
Multim. Syst. | 3 |
| 2015 | Content-oriented multimedia document understanding through cross-media correlation
Tong Lu 0002, Yukang Jin, Feng Su, Palaiahnakote Shivakumara, Chew Lim Tan |
Multim. Tools Appl. | 3 |
| 2014 | Scene Text Detection Based on Robust Stroke Width Transform and Deep Belief Network
Hailiang Xu, Like Xue, Feng Su |
ACCV (2) | 3 |
| 2014 | Auditory scene analysis and recognition with LDA topic modelabstractAnalysis and recognition of auditory scenes play an important role in content-based multimedia processing and context-aware applications. In this paper, we propose an auditory scene recognition scheme that integrates the analysis of the audio data of scene with LDA topic model to discover latent structures (i.e. contextual correlations) of audio words, and generation of intermediate contextual descriptions of audio data on basis of the topics learnt by LDA. We further combine the piecewise low-level audio feature and the contextual feature, and discriminatively classify an audio clip of an unknown scene that is represented as a set of these features using the Hough forest model. The experimental results demonstrate the effectiveness of the proposed scheme, which combines the unsupervised topic modeling by LDA and the supervised classification of auditory scene by Hough forest. Feng Su |
ICME | 1 |
| 2013 | Discriminative Weighting and Subspace Learning for Ensemble Symbol RecognitionabstractMulti-scale parts-based models are particularly effective for recognizing non-segmented graphic symbols, i.e. the symbols interfered by other connecting or intersecting objects in the context. However, treating every symbol part and it features on every scale equally, despite some of them contributing little to the recognition, may lead to unnecessarily high dimensional representation of the symbol and affects the overall performance. In this paper, we propose a discriminant subspace learning and part weighting scheme for the parts-based ensemble symbol recognition model. The probabilistic vote on symbol category by each shape point of the symbol is adaptively weighted based on its surrounding context, and a more compact and representative description of the symbol is exploited based on the codebook learnt from the multi-scale shape context pyramids by K-SVD. The experiments demonstrate the effectiveness and adaptability of the proposed method on non-segmented intersecting symbols. Feng Su |
ICDAR | 1 |
| 2013 | A Novel Multi-view Object Class Detection Framework for Document Image Content AnalysisabstractRecognition of objects from arbitrary viewpoints embedded document images is a new challenge in content-oriented document image analysis. In this paper, we propose a novel framework for detecting generic objects from arbitrary viewpoints described by varied object appearances. We first model the annotated objects from different viewpoints, and then build an explicit correspondence across multi-view detectors. As a result, multi-view objects from untrained viewpoints can be detected by combining the outputs of the adjacent view detectors. Our experiments on several public datasets give promising results for the experimental object classes. Weichong Yin, Feng Su |
ICDAR | 3 |
| 2012 | Moving object tracking using an adaptive colour filterabstractMoving object identification and tracking by computer vision plays an important role in surveillance using mobile robots. In this paper, a new method for moving object tracking using an adaptive colour filter is introduced. This method is capable of identifying the most salient colour feature in the moving object and using this colour feature to track the object. This method is also capable of adapting this selected colour feature when the surrounding condition is changed. Experimental results have shown that the proposed method can perform robustly in tracking a moving object using a robot mounted camera in a crowded environment. Feng Su, Gu Fang 0001 |
ICARCV | 1 |
| 2012 | Auditory context classification using random forestsabstractHigh-level semantic information can be extracted from audio materials to facilitate various content-based analysis and context-awareness applications. In this paper, we propose a novel automatic auditory context classification method, which combines the characterization of audio events and the inference of auditory context category in a single ensemble analysis framework. In the proposed framework, key audio events in the context are characterized by composite features from discriminative representation models (local discriminant bases, pseudo-semantic and bag-of-audio-words) learned from samples. A random forest based ensemble learning and classification model is employed for auditory contexts, in which individual segments of audio stream are classified and aggregated by Hough voting or bagging to form the final context category. The effectiveness of the proposed approach is demonstrated by the experimental results. Feng Su |
ICASSP | 2 |
| 2012 | Anomaly detection with spatio-temporal context using depth images
Xiaolin Ma, Feiming Xu, Feng Su |
ICPR | 4 |
| 2012 | Ensemble symbol recognition with Hough forest
Feng Su |
ICPR | 1 |
| 2012 | Movie Keyframe Retrieval Based on Cross-Media Correlation Detection and Context Model
Yukang Jin, Feng Su |
IEA/AIE | 3 |
| 2012 | A Novel Multi-modal Integration and Propagation Model for Cross-Media Information Retrieval
Wanxia Lin, Feng Su |
MMM | 3 |
| 2011 | Symbol Recognition by Multiresolution Shape Context MatchingabstractWe present a multi resolution scheme for symbol representation and recognition based on statistical shape features. We define a symbol as a set of shape points, each of which is then described by a pyramid of shape context features. The pyramid is constructed by successively partitioning the image surrounding one shape point into increasingly finer sub-regions and computing the local shape context descriptor inside each sub-region. To recognize a symbol, we compute the optimal matching between symbol prototypes and the image region, based on the weighted distance measurements across various scales. We also define an adaptive surround suppression measure that assigns different weights to the shape point depending on the complexity of its surrounding context, so as to reduce the effect of local intersections to shape matching. The experimental results show the effectiveness of the proposed shape context pyramid matching method as well as its promising aspects in handling intersecting symbols. Feng Su, Ruoyu Yang |
ICDAR | 1 |
| 2011 | Environmental sound classification for scene recognition using local discriminant bases and HMMabstractAnalysis and classification of auditory scenes or contexts play important roles in content-based indexing and retrieval of multimedia databases and context-aware applications. In this paper, we propose an environmental sound and auditory scene recognition scheme that focuses on efficient feature representation and classfication of the unstructured composition of a scene (for example, restaurant, street, beach, etc.). We propose to use the local discriminant bases (LDB) technique to identify the discriminatory time-frequency subspace for environmental sounds and then use it for corresponding feature extraction. Based on LDB, we present two recognition models, with or without explicit sound event modeling, for auditory scenes, in which the hidden Markov model (HMM) is used to depict the characteristics and correlations among various events that constitute the scene. The experimental results demonstrate the effectiveness of the proposed approach for auditory scene classification. Feng Su, Gongyou Wang |
ACM Multimedia | 1 |
| 2011 | Design and analysis of on-chip charge pumps for micro-power energy harvesting applicationsabstractCharge balance law based on conservation of charge is stated and employed to analyze charge pumps. For micro-power on-chip implementations, both the positive-plate and the negative-plate parasitic capacitors have to be considered. A first iteration approximation analysis is proposed to analyze charge pumps with parasitic capacitors. Using a 0.35µm CMOS process, 8X linear, Fibonacci and exponential charge pumps are designed and their performances are compared and confirmed by simulations. Wing-Hung Ki, Yan Lu 0002, Feng Su, Chi-Ying Tsui |
VLSI-SoC | 3 |
| 2011 | Two adaptive detectors for range-spread targets in non-Gaussian clutter
Tao Jian, Feng Su, Dianfa Ping |
Sci. China Inf. Sci. | 3 |
| 2010 | Symbol Recognition Combining Vectorial and Pixel-Level Features for Line DrawingsabstractIn this paper, we present an approach for symbol representation and recognition in line drawings, integrating both the vector-based structural description and pixel-level statistical features of the symbol. For the former, a vectorial template is defined on the basis of the vectorization model and exploited in segmenting symbols from the line network. For the latter, a Radon-transform-based signature is employed to characterize shapes on the symbol and the components level. Experimental results on real technical drawings are presented to show the promising aspect of our approach. Feng Su, Ruoyu Yang |
ICPR | 1 |
| 2010 | CFAR assessment of covariance matrix estimators for non-Gaussian clutter
Tao Jian, Feng Su, Changwen Qu, Dianfa Ping |
Sci. China Inf. Sci. | 3 |
| 2009 | Semi-automatic Roof ReconstructionabstractA semi-automatic 3D roof reconstruction method is proposed in this paper. It consists of two components: automatic recognition of 2D plane drawings and interactively “pulling” or “pushing” the recognized results. Only a limited number of reconstruction operations are needed to generate various types of 3D roofs, making the method efficient. Feng Su, Zhengxing Sun |
ICDAR | 3 |
| 2008 | Integrated single-inductor dual-input dual-output boost converter for energy harvesting applicationsabstractAn integrated single-inductor dual-input dual-output (SI DIDO) boost converter for energy harvesting applications was designed in a 0.35μm CMOS process. It provides two regulated output voltages for the load and the charge storage device, and two sources, the energy harvesting source and the charge storage device, are multiplexed to serve as the input. The implementation has several special features. (1) The input power MUX is driven by an internal charge pump for a larger gate drive to save area. (2) The power stage is implemented with an active diode core to eliminate gate drive circuitry. (3) A 1:7 timeslot scheduling with a fixed peak inductor current is adopted to deliver energy to the two outputs with a large difference in load currents. The proposed converter could operate at 1V with up to 85% efficiency at 200mW. Ngok-Man Sze, Feng Su, Yat-Hei Lam, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 2 |
| 2008 | An energy-adaptive MPPT power management unit for micro-power vibration energy harvestingabstractA batteryless power management unit (PMU) that manages harvested low-level vibration energy from a piezoelectric device for a wireless sensor node is presented. An energy-adaptive maximum power point tracking (EA-MPPT) scheme is proposed that allows the PMU to activate different operation modes according to the available power level. The harvested energy is processed by an ac-dc voltage doubler followed by on-chip charge pumps with variable up/down conversion ratios for higher efficiency. Interleaving technique is employed for the high-power output to reduce both current and voltage ripples. The PMU is designed using a 0.35μm CMOS process, and simulation results are presented to demonstrate its functions. Feng Su, Yat-Hei Lam, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 2 |
| 2006 | Ultra-low voltage power management circuit and computation methodology for energy harvesting applicationsabstractA power management and computation methodology is proposed for ultra-low power energy harvesting applications. An integrated exponential charge pump that accepts an input voltage of around 150mV and provides an unregulated output voltage of more than 1.5V serves as the power supply. To cater with the fluctuated energy source and unregulated power supply, a supply side charge-based computation methodology is proposed, of which the computation activity tracks with the fluctuation of the available energy. The idea is demonstrated in a test chip fabricated using a 0.35 mum technology Chi-Ying Tsui, Wing-Hung Ki, Feng Su |
ASP-DAC | 4 |
| 2006 | High efficiency cross-coupled doubler with no reversion lossabstractReversion loss is analyzed in details. A 2/spl times/ cross-coupled charge pump in a 0.35/spl mu/ CMOS process is designed using a gate control strategy that eliminates reversion loss. It switches at 1MHz, delivers a current of 30mA, and reaches an efficiency of 95% for an input voltage of 1V to 1.6V. Feng Su, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 1 |
| 2005 | Design and applications of a parameterized language for modeling architecture componentsabstractThis paper presents a novel parameterized language for modeling architecture components and structures. Aiming at the capability and flexibility of describing various information of a large set of components, this language integrates the power of parameterized 3D modeling and easy-to-use 2D interfaces. This language is used in a 3D architecture design software - Easy Structure. We present the design objectives and the language syntax structure, and illustrate its practical applications used in staircase modeling. Lu Lei, Feng Su, Shijie Cai |
CAD/Graphics | 2 |
| 2005 | A new recognition model for electronic architectural drawings
Tong Lu 0002, Chiew-Lan Tai, Feng Su, Shijie Cai |
Comput. Aided Des. | 3 |
| 2004 | Detection of Weak Targets with Wavelet and Neural Network
Changwen Qu, Feng Su, Yong Huang 0007 |
ISNN (2) | 3 |
| 2002 | An Object-Oriented Progressive-Simplification-Based Vectorization System for Engineering Drawings: Model, Algorithm, and PerformanceabstractExisting vectorization systems for engineering drawings usually take a two-phase workflow: convert a raster image to raw vectors and recognize graphic objects from the raw vectors. The first phase usually separates aground truth graphic object that intersects or touches other graphic objects into several parts, thus, the second phase faces the difficulty of searching for and merging raw vectors belonging to the same object. These operations slow down vectorization and degrade the recognition quality. Imitating the way humans read engineering drawings, we propose an efficient one-phase object-oriented vectorization model that recognizes each class of graphic objects from their natural characteristics. Each ground truth graphic object is recognized directly in its entirety at the pixel level. The raster image is progressively simplified by erasing recognized graphic objects to eliminate their interference with subsequent recognition. To evaluate the performance of the proposed model, we present experimental results on real-life drawings and quantitative analysis using third party protocols. The evaluation results show significant improvement in speed and recognition rate. Jiqiang Song, Feng Su, Chiew-Lan Tai, Shijie Cai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Dimension Recognition and Geometry Reconstruction in Vectorization of Engineering DrawingsabstractThis paper presents a novel approach for recognizing and interpreting dimensions in engineering drawings. It starts by detecting potential dimension frames, each comprising only the line and text components of a dimension, then verifies them by detecting the dimension symbols. By removing the prerequisite of symbol recognition from detection of dimension sets, our method is capable of handling low quality drawings. We also propose a reconstruction algorithm for rebuilding the drawing entities based on the recognized dimension annotations. A coordinate grid structure is introduced to represent and analyze two-dimensional spatial constraints between entities; this simplifies and unifies the process of rectifying deviations of entity dimensions induced during scanning and vectorization. Feng Su, Jiqiang Song, Chiew-Lan Tai, Shijie Cai |
CVPR (1) | 1 |
| 2000 | Line Net Global Vectorization: an Algorithm and Its Performance EvaluationabstractIn this paper, an efficient global algorithm for vectorizing line drawings is presented. It first extracts a seed segment of a graphic entity from a raster image to obtain its direction and width, then tracks the pixels under the guidance of the direction so that the tracking can track through functions and is not affected by noise and degradation of image quality. Thus, an entity will be vectorized in one step without postprocessing. The relations among lines are also used to realize the continuous vectorization of a line net. The speed and quality of vectorization are greatly improved with this algorithm. The performance evaluation is carried out both by theoretical analysis and by experiments. Comparisons with other vectorization algorithms are also made. Jiqiang Song, Feng Su, Jibing Chen, Chiew-Lan Tai, Shijie Cai |
CVPR | 2 |
| 2000 | A Knowledge-Aided Line Network Oriented Vectorisation Method for Engineering Drawings
Jiqiang Song, Feng Su, Jibing Chen, Shijie Cai |
Pattern Anal. Appl. | 2 |