Kai Chen 0006

dblp:c/KaiChen6 · DBLP profile ↗
← Back
46ranked-venue papers
3as first author
15since 2021 · last 2025
0009-0001-1321-5217ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 1 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorComputer networks · 2
YearPublicationVenuePosition
2025 Robot Navigation in Unknown and Cluttered Workspace with Dynamical System Modulation in Starshaped Roadmap
abstract
Compared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and control framework for navigating robots in unknown and cluttered environments using a dynamically constructed starshaped roadmap. Our approach generates a starshaped representation of the surrounding free space from real-time sensor data using piece-wise polynomials. Additionally, an incremental roadmap maintaining the connectivity information is constructed, and a searching algorithm efficiently selects short-term goals on this roadmap. Importantly, this framework addresses dead-end situations with a graph updating mechanism. To ensure safe and efficient movement within the starshaped roadmap, we propose a reactive controller based on Dynamic System Modulation (DSM). This controller facilitates smooth motion within starshaped regions and their intersections, avoiding conservative and short-sighted behaviors and allowing the system to handle intricate obstacle configurations in unknown and cluttered environments. Comprehensive evaluations in both simulations and real-world experiments show that the proposed method achieves higher success rates and reduced travel times compared to other methods. It effectively manages intricate obstacle configurations, avoiding conservative and myopic behaviors. The source code will be released on website11Available at: github.com/kkkkkaiai/starshaped_roadmap.
Kai Chen 0006, Haichao Liu 0003, Yulin Li 0001, Jianghua Duan, Lei Zhu 0003, Jun Ma 0008
ICRA1
2025 Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller
abstract
Interactive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in narrow or constrained spaces where the robot’s dimensions restrict its manipulation capabilities. This paper introduces a novel interactive navigation framework for legged manipulators, featuring an active arm-pushing mechanism that enables the robot to reposition movable obstacles in space-constrained environments. To this end, we develop a reinforcement learning-based arm-pushing controller with a two-stage reward strategy for object manipulation. Specifically, this strategy first directs the manipulator to a designated pushing zone to achieve a kinematically feasible contact configuration. Then, the end effector is guided to maintain its position at appropriate contact points for stable object displacement while preventing toppling. The simulations validate the robustness of the arm-pushing controller, showing that the two-stage reward strategy improves policy convergence and long-term performance. Real-World experiments further demonstrate the effectiveness of the proposed navigation framework, which achieves shorter paths and reduced traversal time. The open-source project can be found at https://zhihaibi.github.io/interactive-push.github.io/.
Zhihai Bi, Kai Chen 0006, Chunxin Zheng, Yulin Li 0001, Haoang Li, Jun Ma 0008
IROS2
2025 UDMC: Unified Decision-Making and Control Framework for Urban Autonomous Driving With Motion Prediction of Traffic Participants
abstract
Current autonomous driving systems often struggle to balance decision-making and motion control while ensuring safety and traffic rule compliance, especially in complex urban environments. Existing methods may fall short due to separate handling of these functionalities, leading to inefficiencies and safety compromises. To address these challenges, we introduce UDMC, an interpretable and unified Level 4 autonomous driving framework. UDMC integrates decision-making and motion control into a single optimal control problem (OCP), considering the dynamic interactions with surrounding vehicles, pedestrians, road lanes, and traffic signals. By employing innovative potential functions to model traffic participants and regulations, and incorporating a specialized motion prediction module, our framework enhances on-road safety and rule adherence. The integrated design allows for real-time execution of flexible maneuvers suited to diverse driving scenarios. High-fidelity simulations conducted in CARLA exemplify the framework’s computational efficiency, robustness, and safety, resulting in superior driving performance when compared against various baseline models. Our open-source project is available athttps://github.com/henryhcliu/udmc_carla.git.
Haichao Liu 0003, Kai Chen 0006, Yulin Li 0001, Zhenmin Huang, Ming Liu 0001, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.2
2024 MCGMapper: Light-Weight Incremental Structure from Motion and Visual Localization with Planar Markers and Camera Groups
abstract
Structure from Motion (SfM) and visual localization in indoor texture-less scenes and industrial scenarios present prevalent yet challenging research topics. Existing SfM methods designed for natural scenes typically yield low accuracy or map-building failures due to insufficient robust feature extraction in such settings. Visual markers, with their artificially designed features, can effectively address these issues. Nonetheless, existing marker-assisted SfM methods encounter problems like slow running speed and difficulties in convergence; and also, they are governed by the strong assumption of unique marker size. In this paper, we propose a novel SfM framework that utilizes planar markers and multiple cameras with known extrinsics to capture the surrounding environment and reconstruct the marker map. In our algorithm, the initial poses of markers and cameras are calculated with Perspective-n-Points (PnP) in the front-end, while bundle adjustment methods customized for markers and camera groups are designed in the back-end to optimize the 6-DOF pose directly. Our algorithm facilitates the reconstruction of large scenes with different marker sizes, and its accuracy and speed of map building are shown to surpass existing methods. Our approach is suitable for a wide range of scenarios, including laboratories, basements, warehouses, and other industrial settings. Furthermore, we incorporate representative scenarios into simulations and also supply our datasets with pose labels to address the scarcity of quantitative ground-truth datasets in this research field. The datasets and source code are available on GitHub1.
Yusen Xie, Zhenmin Huang, Kai Chen 0006, Lei Zhu 0003, Jun Ma 0008
IROS3
2024 Incremental Learning-Based Real-Time Trajectory Prediction for Autonomous Driving via Sparse Gaussian Process Regression
abstract
In the context of spatial-temporal autonomous driving, the accurate and real-time trajectory prediction of the surrounding vehicle (SV) is crucial. This paper aims to design an efficient, accurate, and interpretable unimodal trajectory prediction approach. To achieve this objective, we employ Sparse Gaussian Process Regression (SGPR), which enables large dataset learning and efficient inference of future trajectories. This approach ensures accurate predictions while maintaining high computational efficiency. To further enhance the robustness of the prediction module, we propose the translation and rotation transformation strategy, which effectively simplifies the prediction problem. Additionally, we utilize an instant evaluation algorithm to assess the prediction performance and maintain a streaming dataset for incremental learning, capable of adapting to dynamic driving environments. In our experimental evaluation, we compare our proposed trajectory prediction approach with a series of existing methods. The results demonstrate that our work achieves superior prediction accuracy while requiring less inference time. It is noteworthy that, the proposed SGPR-based trajectory prediction approach with rotation equivalence is able to swiftly infer and incrementally learn from dynamic environments, which makes it a promising tool for enhancing safety and efficiency in autonomous driving systems.
Haichao Liu 0003, Kai Chen 0006, Jun Ma 0008
IV2
2023 Boosting Point Clouds Rendering via Radiance Mapping
abstract
Recent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive, which enables its deployment in mobile computing device. In this work, we focus on boosting the image quality of point clouds rendering with a compact model design. We first analyze the adaption of the volume rendering formulation on point clouds. Based on the analysis, we simplify the NeRF representation to a spatial mapping function which only requires single evaluation per pixel. Further, motivated by ray marching, we rectify the the noisy raw point clouds to the estimated intersection between rays and surfaces as queried coordinates, which could avoid spatial frequency collapse and neighbor point disturbance. Composed of rasterization, spatial mapping and the refinement stages, our method achieves the state-of-the-art performance on point clouds rendering, outperforming prior works by notable margins, with a smaller model size. We obtain a PSNR of 31.74 on NeRF-Synthetic, 25.88 on ScanNet and 30.81 on DTU. Code and data are publicly available in https://github.com/seanywang0408/RadianceMapping.
Bingbing Ni, Teng Li 0001, Kai Chen 0006, Wenjun Zhang 0001
AAAI5
2023 GenTC: Generative Transformer via Contrastive Learning for Receipt Information Extraction
Xinrui Deng, Kefan Ma, Kai Chen 0006, Jie Guo 0011, Weidong Qiu
ICANN (6)4
2023 RRecT: Chinese Text Recognition with Radical-Enhanced Recognition Transformer
Xinrui Deng, Kefan Ma, Kai Chen 0006, Jie Guo 0011, Weidong Qiu
ICANN (6)4
2023 Learning Shape Primitives via Implicit Convexity Regularization
abstract
Shape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view images by introducing implicit surface rendering. It is challenging since implicit shapes have a high degree of freedom, which violates the simplicity property of shape primitives. In this work, a novel regularization term named Implicit Convexity Regularization (ICR) imposed on implicit primitive learning is proposed to tackle this problem. We start with the convexity definition of general 3D shapes, and then derive the equivalent expression for implicit shapes represented by signed distance functions (SDFs). Further, instead of directly constraining the output SDF values which cause unstable optimization, we alternatively impose constraint on second order directional derivatives on line segments inside the shapes, which proves to be a tighter condition for 3D convexity. Implicit primitives constrained by the proposed ICR are combined into a whole object via softmax-weighted-sum operation over all primitive SDFs. Experiments on synthetic and real-world datasets show that our method is able to decompose objects into simple and reasonable shape primitives without the need of segmentation labels or 3D data. Code and data is publicly available in https://github.com/seanywang0408/ICR.
Kai Chen 0006, Teng Li 0001, Wenjun Zhang 0001, Bingbing Ni
ICCV3
2023 TDAE: Text Detection with Affinity Areas and Evolution Strategies
Kefan Ma, Kai Chen 0006, Jie Guo 0011, Weidong Qiu
ICDAR (6)4
2023 Relational Contrastive Learning for Scene Text Recognition
abstract
Context-aware methods achieved great success in supervised scene text recognition via incorporating semantic priors from words. We argue that such prior contextual information can be interpreted as the relations of textual primitives due to the heterogeneous text and background, which can provide effective self-supervised labels for representation learning. However, textual relations are restricted to the finite size of dataset due to lexical dependencies, which causes the problem of over-fitting and compromises representation robustness. To this end, we propose to enrich the textual relations via rearrangement, hierarchy and interaction, and design a unified framework called RCLSTR: Relational Contrastive Learning for Scene Text Recognition. Based on causality, we theoretically explain that three modules suppress the bias caused by the contextual prior and thus guarantee representation robustness. Experiments on representation quality show that our method outperforms state-of-the-art self-supervised STR methods. Code is available at https://github.com/ThunderVVV/RCLSTR.
Jinglei Zhang 0003, Tiancheng Lin 0001, Yi Xu 0001, Kai Chen 0006, Rui Zhang 0052
ACM Multimedia4
2023 A Learning-Based Object Tracking Strategy Using Visual Sensors and Intelligent Robot Arm
abstract
This paper focuses on addressing the visual tracking problem using learning-based methods for object tracking tasks. This problem contains a major difficulty, i.e., how to acquire a satisfactory generalization ability of the developed system? In this paper, firstly, the object state tracking system, including a camera-in-hand, a 3D camera and a Rethink Baxter robot, is introduced. The problem formulation is also presented. Secondly, we propose a Kalman-based estimation strategy to acquire the object’s state. In addition, a learning-based tracking controller is developed using the Gaussian mixture models (GMM) method to steer the robot end-effector to track the mobile object. Thirdly, to guarantee system stability (i.e., the position and velocity errors between the object and end-effector will always converge to zeros), the controller parameter constraints are derived, which is a theoretical contribution of this paper. The controller parameter adjustment is avoided by the proposed training process. Thus, the proposed method becomes easy to implement, which is a practical contribution. Finally, the effectiveness of the proposed method is demonstrated by simulation and experimental examples, and the proposed method has satisfactory generalization ability. Note to Practitioners—This paper studies object tracking problems for different practical applications, such as industrial cutting, grasping and dynamic monitoring. Different trajectory tracking methods have been widely applied in the industrial area. However, users always complain that when the object or trajectory is changed, the tracking controller more or less needs to be re-adjusted. This re-adjust process always requires professional knowledge and programming experience, and thus a factory must employ some professional engineers. In addition, since the objects may be diverse in shape, color and size, to acquire the accurate object position and velocity, an appropriate solution is necessary. Motivated by the above introductions, this paper aims to develop a learning-based controller to track different complex trajectories without frequent and specific parameter adjustment processes. Firstly, a visual measurement system is developed to quickly find and estimate the position and velocity of an object. Secondly, with the object’s information, the learning from demonstration method (GMM method) is applied for the control policy design. Thirdly, the detailed system stability analysis is presented, and the corresponding controller parameter constraints are derived and considered in the proposed control policy. Subsequently, with the demonstration data, the packaged learning algorithm will automatically compute the controller parameters, and users can change the controller performance only by providing the desired demonstrations. In summary, this paper proposes a systematic object tracking solution, and it may bring a new idea to develop a practical object tracking system, using both the learning-based methods to improve the ability of generalization for tracking different objects.
Sheng Xu 0004, Kai Chen 0006, Yongsheng Ou, Chenguang Yang 0001
IEEE Trans Autom. Sci. Eng.2
2021 Spatial Aggregation for Scene Text Recognition
Yi-Li Huang, Chengyu Gu, Shi-Lin Wang, Kai Chen 0006
BMVC5
2021 UMLE: Unsupervised Multi-discriminator Network for Low Light Enhancement
abstract
Low-light image enhancement is a complex and vital task including, recovering color and texture details from low-light images. For automated driving, low-light scenarios will have severe implications for vision-based applications. To address this problem, we propose a real-time unsupervised generative adversarial network (GAN) with multiple discriminators. It includes a multi-scale discriminator, a texture discriminator, and a color discriminator to evaluate images from different perspectives. Furthermore, considering the uneven illumination distribution of images and the different information contained in the channels, we adopte a feature fusion attention module to combine channel attention with pixel attention to extract image features. Experiments show that our method outperforms state-of-the-art methods in qualitative and quantitative evaluation and provides visible improvements in SLAM localization effects.
Yangyang Qu, Kai Chen 0006, Chao Liu 0056, Yongsheng Ou
ICRA2
2021 Gliding Vertex on the Horizontal Bounding Box for Multi-Oriented Object Detection
abstract
Object detection has recently experienced substantial progress. Yet, the widely adopted horizontal bounding box representation is not appropriate for ubiquitous oriented objects such as objects in aerial images and scene texts. In this paper, we propose a simple yet effective framework to detect multi-oriented objects. Instead of directly regressing the four vertices, we glide the vertex of the horizontal bounding box on each corresponding side to accurately describe a multi-oriented object. Specifically, We regress four length ratios characterizing the relative gliding offset on each corresponding side. This may facilitate the offset learning and avoid the confusion issue of sequential label points for oriented objects. To further remedy the confusion issue for nearly horizontal objects, we also introduce an obliquity factor based on area ratio between the object and its horizontal bounding box, guiding the selection of horizontal or oriented detection for each object. We add these five extra target variables to the regression head of faster R-CNN, which requires ignorable extra computation time. Extensive experimental results demonstrate that without bells and whistles, the proposed method achieves superior performances on multiple multi-oriented object detection benchmarks including object detection in aerial images, scene text detection, pedestrian detection in fisheye images.
Yongchao Xu, Mingtao Fu, Qimeng Wang, Yukang Wang, Kai Chen 0006, Gui-Song Xia, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Real-Time Scene Text Detection with Differentiable Binarization
abstract
Recently, segmentation-based methods are quite popular in scene text detection, as the segmentation results can more accurately describe scene text of various shapes such as curve text. However, the post-processing of binarization is essential for segmentation-based detection, which converts probability maps produced by a segmentation method into bounding boxes/regions of text. In this paper, we propose a module named Differentiable Binarization (DB), which can perform the binarization process in a segmentation network. Optimized along with a DB module, a segmentation network can adaptively set the thresholds for binarization, which not only simplifies the post-processing but also enhances the performance of text detection. Based on a simple segmentation network, we validate the performance improvements of DB on five benchmark datasets, which consistently achieves state-of-the-art results, in terms of both detection accuracy and speed. In particular, with a light-weight backbone, the performance improvements by DB are significant so that we can look for an ideal tradeoff between detection accuracy and efficiency. Specifically, with a backbone of ResNet-18, our detector achieves an F-measure of 82.8, running at 62 FPS, on the MSRA-TD500 dataset. Code is available at: https://github.com/MhLiao/DB.
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 0006, Xiang Bai
AAAI4
2020 Weakly Supervised Attention Rectification for Scene Text Recognition
abstract
Scene text recognition has become a hot topic in recent years due to its booming real-life applications. Attention-based encoder-decoder framework has become one of the most popular frameworks especially in the irregular text scenario. However, the “attention drift” problem hinders the recognition performance for most existing attention-based scene text recognition methods. To solve this problem, we propose an auxiliary supervision branch along with the attention-based encoder-decoder framework. A new loss function is designed to refine the feature map and to help the attention region align the target character area. Compared with existing attention rectification mechanisms, our method does not require character-level annotations or introduce any additional trainable parameter. Furthermore, our method can improve the performance for both RNN-Attention and Scaled Dot-Product Attention. The experiment results on various benchmarks have demonstrated that the proposed approach outperforms the state-of-the-art methods in both regular and irregular text recognition scenarios.
Chengyu Gu, Shi-Lin Wang, Yiwei Zhu, Kai Chen 0006
ICPR5
2020 RLST: A Reinforcement Learning Approach to Scene Text Detection Refinement
abstract
Within the research of scene text detection, some previous work has already achieved significant accuracy and efficiency. However, most of the work was generally done without considering about the implicit relationship between detection and eye movements. In this paper, we propose a new method for scene text detection especially for its refinement based on reinforcement learning. The idea of this method is inspired by Saccadic Eye Movements and Peripheral Vision. A saccade makes it possible for humans to orient the gaze to the location where a visual object has appeared. Peripheral vision gathers visual information of surroundings which provides supplement to foveal vision during gazing. We propose a simple pipeline, imitating the way human eyes do a saccade and collect peripheral information, to locate scene text roughly and to refine multi-scale vision field iteratively using reinforcement learning. For both training and evaluation, we use ICDAR2015 Challenge 4 dataset as a base and design several criteria to measure the feasibility of our work.
Kai Chen 0006, Jie Guo 0011, Weidong Qiu
ICPR3
2020 Image-based Table Cell Detection: a Novel Table Structure Decomposition Method with New Dataset
abstract
Recently deep learning has been applied to decompose table structure with the main ideas of detecting table lines and then forming table cells. However, the existing methods face problems in dealing with tables with rotation or no internal table lines. To tackle these problems, we propose a novel table structure decomposition method, which directly detects table cells as objects and creates table structure. Extensions to the existing object detection models including effective table projection module are proposed to adapt to the table cell detection. To support the training of the enhanced models, we create a large image-based table dataset TableCell with cell level annotations. A novel and efficient semi-supervised method is proposed to annotate this new dataset. Experiments demonstrate that our proposed table structure decomposition method is simple, effective and robust to the tables without table lines or with rotation. Our dataset and code will be made available11https://github.com/weidafeng/TableCell.
Dafeng Wei, Hongtao Lu 0001, Yi Zhou 0003, Kai Chen 0006
ICPR4
2020 Enhanced Object Detection With Deep Convolutional Neural Networks for Advanced Driving Assistance
abstract
Object detection is a critical problem for advanced driving assistance systems (ADAS). Recently, convolutional neural networks (CNN) achieved large successes on object detection, with performance improvement over traditional approaches, which use hand-engineered features. However, due to the challenging driving environment (e.g., large object scale variation, object occlusion, and bad light conditions), popular CNN detectors do not achieve very good object detection accuracy over the KITTI autonomous driving benchmark dataset. In this paper, we propose three enhancements for CNN-based visual object detection for ADAS. To address the large object scale variation challenge, deconvolution and fusion of CNN feature maps are proposed to add context and deeper features for better object detection at low feature map scales. In addition, soft non-maximal suppression (NMS) is applied across object proposals at different feature scales to address the object occlusion challenge. As the cars and pedestrians have distinct aspect ratio features, we measure their aspect ratio statistics and exploit them to set anchor boxes properly for better object matching and localization. The proposed CNN enhancements are evaluated with various image input sizes by experiments over KITTI dataset. The experimental results demonstrate the effectiveness of the proposed enhancements with good detection performance over KITTI test set.
Jianhua He 0001, Yi Zhou 0003, Kai Chen 0006, Zuoyin Tang, Zhiliang Xiong
IEEE Trans. Intell. Transp. Syst.4
2019 ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
abstract
The ICDAR 2019 Challenge on "Scanned receipts OCR and key information extraction" (SROIE) covers important aspects related to the automated analysis of scanned receipts. The SROIE tasks play a key role in many document analysis systems and hold significant commercial potential. Although a lot of work has been published over the years on administrative document analysis, the community has advanced relatively slowly, as most datasets have been kept private. One of the key contributions of SROIE to the document analysis community is to offer a first, standardized dataset of 1000 whole scanned receipt images and annotations, as well as an evaluation procedure for such tasks. The Challenge is structured around three tasks, namely Scanned Receipt Text Localization (Task 1), Scanned Receipt OCR (Task 2) and Key Information Extraction from Scanned Receipts (Task 3). The competition opened on 10th February, 2019 and closed on 5th May, 2019. We received 29, 24 and 18 valid submissions received for the three competition tasks, respectively. This report presents the competition datasets, define the tasks and the evaluation protocols, offer detailed submission statistics, as well as an analysis of the submitted performance. While the tasks of text localization and recognition seem to be relatively easy to tackle, it is interesting to observe the variety of ideas and approaches proposed for the information extraction task. According to the submissions' performance we believe there is still margin for improving information extraction performance, although the current dataset would have to grow substantially in following editions. Given the success of the SROIE competition evidenced by the wide interest generated and the healthy number of submissions from academic, research institutes and industry over different countries, we consider that the SROIE competition can evolve into a useful resource for the community, drawing further attention and promoting research and development efforts in this field.
Kai Chen 0006, Jianhua He 0001, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar
ICDAR2
2019 A New Approach for Integrated Recognition and Correction of Texts from Images
abstract
Automatic recognition and error correction of texts from images are critical for many commercial applications such as receipt recognition, which have very high accuracy requirements. In this paper we propose an integrated image based text recognition and correction approach to improve accuracy. There are two levels of text recognition and correction integration in the proposed approach. Firstly, a beam search strategy is designed to generate a set of text candidates, based on the probability distribution of text prediction outcomes from a deep learning recognition model. Then a word-level lexicon check is applied to select only one from the candidate text sentences, which has the highest prediction probability among those with all words present in the lexicon. Jointly the beam search and lexicon check can effectively correct some recognition errors. Secondly, an encoder-decoder language model based corrector is developed to correct potential recognition errors in the selected output texts that fail the lexicon check. Training samples for the corrector are created from the recognition outcomes and can be expanded by associating multiple text candidates with one image label. We conduct experiments on ICDAR'13 and CH10K datasets to evaluate the proposed approach and the impact of these two levels of integration on accuracy. Experiment results show that the proposed approach outperforms the existing one with higher recall and much higher recognition accuracy through effective exploitation of joint recognition and correction design.
Kai Chen 0006, Jianhua He 0001, Yunrui Lian, Yi Zhou 0003
ICDAR2
2019 Wacnet: Word Segmentation Guided Characters Aggregation Net for Scene Text Spotting With Arbitrary Shapes
abstract
In this paper, we propose an end-to-end trainable framework for scene text spotting which can handle text with arbitrary shapes. The proposed framework is called Word Segmentation Guided Characters Aggregation Net (WAC-Net), which consists of a shared convolutional backbone and two task-specific subnetworks. One subnetwork does word-level instance-aware segmentation (WSN) and the other does char-level detection and recognition (CDRN). The entire framework segments each word instance while detects and recognizes each character in one single forward pass. These two subnetworks are jointly trained by multi-task learning. At the inference stage, characters are aggregated into words guided by word instance segmentation results. Experiments are conducted on two datasets with arbitrary shapes, and the results demonstrate the effectiveness of the proposed method.
Yuchen Dai, Kai Chen 0006, Jie Guo 0011, Weidong Qiu
ICIP4
2019 Text Recognition in Images Based on Transformer with Hierarchical Attention
abstract
Recognizing text in images has been a hot research topic in computer vision for decades due to its various application. However, the variations in text appearance in term of perspective distortion, text line curvature, text styles, etc., cause great trouble in text recognition. Inspired by the Transformer structure [1] that achieved outstanding performance in many natural language processing related applications, we propose a new Transformer-like structure for text recognition in images, which is referred to as the Hierarchical Attention Transformer Network (HATN). The entire network can be trained end-to-end by using only images and sentence-level annotations. A new hierarchical attention mechanism is proposed to lean the character-level, word-level and sentence-level contexts more efficiently and sufficiently. Extensive experiments on seven public datasets with regular and irregular text arrangements have demonstrated that the proposed HATN can achieve accurate recognition results with high efficiency.
Yiwei Zhu, Shi-Lin Wang, Kai Chen 0006
ICIP4
2019 DSAN: Double Supervised Network with Attention Mechanism for Scene Text Recognition
abstract
In this paper, we propose Double Supervised Network with Attention Mechanism (DSAN), a novel end-to-end trainable framework for scene text recognition. It incorporates one text attention module during feature extraction which enforces the model to focus on text regions and the whole framework is supervised by two branches. One supervision comes from context-level modelling branch and another comes from one extra supervision enhancement branch which aims at tackling inexplicit semantic information at character level. These two supervisions can benefit each other and yield better performance. The proposed approach can recognize text in arbitrary length and does not need any predefined lexicon. Our method achieves the current state-of-the-art results on three text recognition benchmarks: IIIT5K, ICDAR2013 and SVT reaching accuracy 88.6%, 92.3% and 84.1% respectively which suggests the effectiveness of the proposed method.
Yuchen Dai, Kai Chen 0006, Jie Guo 0011
VCIP5
2018 Fused Text Segmentation Networks for Multi-oriented Scene Text Detection
abstract
In this paper, we introduce a novel end-end framework for multi-oriented scene text detection from an instance-aware semantic segmentation perspective. We present Fused Text Segmentation Networks, which combine multi-level features during the feature extracting as text instance may rely on finer feature expression compared to general objects. It detects and segments the text instance jointly and simultaneously, leveraging merits from both semantic segmentation task and region proposal based object detection task. Not involving any extra pipelines, our approach surpasses the current state of the art on multi-oriented scene text detection benchmarks: ICDAR2015 Incidental Scene Text and MSRA-TD500 reaching Hmean 84.1 % and 82.0 % respectively. Morever, we report a baseline on total-text containing curved text which suggests effectiveness of the proposed approach.
Yuchen Dai, Youxuan Xu, Kai Chen 0006, Jie Guo 0011, Weidong Qiu
ICPR5
2018 Multitier Fog Computing With Large-Scale IoT Data Analytics for Smart Cities
abstract
Analysis of Internet of Things (IoT) sensor data is a key for achieving city smartness. In this paper a multitier fog computing model with large-scale data analytics service is proposed for smart cities applications. The multitier fog is consisted of ad-hoc fogs and dedicated fogs with opportunistic and dedicated computing resources, respectively. The proposed new fog computing model with clear functional modules is able to mitigate the potential problems of dedicated computing infrastructure and slow response in cloud computing. We run analytics benchmark experiments over fogs formed by Rapsberry Pi computers with a distributed computing engine to measure computing performance of various analytics tasks, and create easy-to-use workload models. Quality of services (QoS) aware admission control, offloading, and resource allocation schemes are designed to support data analytics services, and maximize analytics service utilities. Availability and cost models of networking and computing resources are taken into account in QoS scheme design. A scalable system level simulator is developed to evaluate the fog-based analytics service and the QoS management schemes. Experiment results demonstrate the efficiency of analytics services over multitier fogs and the effectiveness of the proposed QoS schemes. Fogs can largely improve the performance of smart city analytics services than cloud only model in terms of job blocking probability and service utility.
Jianhua He 0001, Kai Chen 0006, Zuoyin Tang, Yi Zhou 0003, Yan Zhang 0002
IEEE Internet Things J.3
2017 Collaborative filtering and deep learning based recommendation system for cold start items
Jianhua He 0001, Kai Chen 0006, Yi Zhou 0003, Zuoyin Tang
Expert Syst. Appl.3
2017 A Convolutional Neural Network-Based Chinese Text Detection Algorithm via Text Structure Modeling
abstract
Text detection in a natural environment plays an important role in many computer vision applications. While existing text detection methods are focused on English characters, there are strong application demands on text detection in other languages, such as Chinese. In this paper, we present a novel text detection algorithm for Chinese characters based on a specific designed convolutional neural network (CNN). The CNN contains a text structure component detector layer, a spatial pyramid layer, and a multi-input-layer deep belief network (DBN). The CNN is pre-trained via a convolutional sparse auto-encoder, specifically designed for extracting complex features from Chinese characters. In particular, the text structure component detectors enhance the accuracy and uniqueness of feature descriptors by extracting multiple text structure components in various ways. The spatial pyramid layer enhances the scale invariability of the CNN for detecting texts in multiple scales. Finally, the multi-input-layer DBN replaces the fully connected layers in the CNN to ensure features from multiple scales are comparable. A multilingual text detection dataset, in which texts in Chinese, English, and digits are labeled separately, is set up to evaluate the proposed text detection algorithm. The proposed algorithm shows a significant performance improvement over the baseline CNN algorithms. In addition the proposed algorithm is evaluated over a public multilingual benchmark and achieves state-of-the-art result under multiple languages. Furthermore, a simplified version of the proposed algorithm with only general components is evaluated on the ICDAR 2011 and 2013 datasets, showing comparable detection performance to the existing general text detection algorithms.
Xiaohang Ren, Yi Zhou 0003, Jianhua He 0001, Kai Chen 0006, Xiaokang Yang 0001, Jun Sun 0005
IEEE Trans. Multim.4
2017 Cost-Effective Online Trending Topic Detection and Popularity Prediction in Microblogging
abstract
Identifying topic trends on microblogging services such as Twitter and estimating those topics’ future popularity have great academic and business value, especially when the operations can be done in real time. For any third party, however, capturing and processing such huge volumes of real-time data in microblogs are almost infeasible tasks, as there always exist API (Application Program Interface) request limits, monitoring and computing budgets, as well as timeliness requirements. To deal with these challenges, we propose a cost-effective system framework with algorithms that can automatically select a subset of representative users in microblogging networks in offline, under given cost constraints. Then the proposed system can online monitor and utilize only these selected users’ real-time microposts to detect the overall trending topics and predict their future popularity among the whole microblogging network. Therefore, our proposed system framework is practical for real-time usage as it avoids the high cost in capturing and processing full real-time data, while not compromising detection and prediction performance under given cost constraints. Experiments with real microblogs dataset show that by tracking only 500 users out of 0.6 million users and processing no more than 30,000 microposts daily, about 92% trending topics could be detected and predicted by the proposed system and, on average, more than 10 hours earlier than they appear in official trends lists.
Zhongchen Miao, Kai Chen 0006, Yi Fang 0008, Jianhua He 0001, Yi Zhou 0003, Wenjun Zhang 0001, Hongyuan Zha
ACM Trans. Inf. Syst.2
2016 A novel text structure feature extractor for Chinese scene text detection and recognition
abstract
Scene text information extraction plays an important role in many computer vision applications. Unlike most existing text extraction algorithms for English texts, in this paper, we focus on Chinese texts, which are more complex in stroke and structure. To tackle this challenging problem, we propose a novel convolutional neural network (CNN) based text structure feature extractor for Chinese texts. Each Chinese character contains its specific types and combination of text structure components, which is rarely seen in backgrounds. Thus, different from the features only applicable to one text extraction stage (text detection or text recognition), the text structure component feature is suitable for both Chinese text detection and recognition. A text structure component detector (TSCD) layer is designed to detect the large amount of component types, which is the most challenging part of extracting text structure component features. Through statistical classification various types of text structure component are detected by their specially designed convolutional units in the TSCD layer. With the TSCD layer, the CNN has improvements in the accuracy and uniqueness of text feature description. In the evaluation, both text detection and recognition algorithms based on the proposed text structure feature extractor achieve state-of-the-art results in two datasets.
Xiaohang Ren, Kai Chen 0006, Xiaokang Yang 0001, Yi Zhou 0003, Jianhua He 0001, Jun Sun 0005
ICPR2
2016 A novel scene text detection algorithm based on convolutional neural network
abstract
Candidate text region extraction plays a critical role in convolutional neural network (CNN) based text detection from natural images. In this paper, we propose a CNN based scene text detection algorithm with a new text region extractor. The so called candidate text region extractor I-MSER is based on Maximally Stable Extremal Region (MSER), which can improve the independency and completeness of the extracted candidate text regions. Design of I-MSER is motivated by the observation that text MSERs have high similarity and are close to each other. The independency of candidate text regions obtained by I-MSER is guaranteed by selecting the most representative regions from a MSER tree which is generated according to the spatial overlapping relationship among the MSERs. A multi-layer CNN model is trained to score the confidence value of the extracted regions extracted by the I-MSER for text detection. The new text detection algorithm based on I-MSER is evaluated with wide-used ICDAR 2011 and 2013 datasets and shows improved detection performance compared to the existing algorithms.
Xiaohang Ren, Kai Chen 0006, Xiaokang Yang 0001, Yi Zhou 0003, Jianhua He 0001, Jun Sun 0005
VCIP2
2015 A LSTM-based method for stock returns prediction: A case study of China stock market
abstract
The presented paper modeled and predicted China stock returns using LSTM. The historical data of China stock market were transformed into 30-days-long sequences with 10 learning features and 3-day earning rate labeling. The model was fitted by training on 900000 sequences and tested using the other 311361 sequences. Compared with random prediction method, our LSTM model improved the accuracy of stock returns prediction from 14.3% to 27.2%. The efforts demonstrated the power of LSTM in stock market prediction in China, which is mechanical yet much more unpredictable.
Kai Chen 0006, Yi Zhou 0003, Fangyan Dai
IEEE BigData1
2015 Online trendy topics detection in microblogs with selective user monitoring under cost constraints
abstract
As microblog services such as Twitter become a fast and convenient communication approach, identification of trendy topics in microblog services has great academic and business value. However detecting trendy topics is very challenging due to huge number of users and short-text posts in microblog diffusion networks. In this paper we introduce a trendy topics detection system under computation and communication resource constraints. In stark contrast to retrieving and processing the whole microblog contents, we develop an idea of selecting a small set of microblog users and processing their posts to achieve an overall acceptable trendy topic coverage, without exceeding resource budget for detection. We formulate the selection operation of these subset users as mixed-integer optimization problems, and develop heuristic algorithms to compute their approximate solutions. The proposed system is evaluated with real-time test data retrieved from Sina Weibo, the dominant microblog service provider in China. It's shown that by monitoring 500 out of 1.6 million microblog users and tracking their microposts (about 15,000 daily) with our system, nearly 65% trendy topics can be detected, while on average 5 hours earlier before they appear in Sina Weibo official trends.
Zhongchen Miao, Kai Chen 0006, Yi Zhou 0003, Hongyuan Zha, Jianhua He 0001, Xiaokang Yang 0001, Wenjun Zhang 0001
ICC2
2014 An improved memory management scheme for large scale graph computing engine GraphChi
abstract
GraphChi is the first reported disk-based graph engine that can handle billion-scale graphs on a single PC efficiently. GraphChi is able to execute several advanced data mining, graph mining and machine learning algorithms on very large graphs. With the novel technique of parallel sliding windows (PSW) to load subgraph from disk to memory for vertices and edges updating, it can achieve data processing performance close to and even better than those of mainstream distributed graph engines. GraphChi mentioned that its memory is not effectively utilized with large dataset, which leads to suboptimal computation performances. In this paper we are motivated by the concepts of “pin ” from TurboGraph and “ghost” from GraphLab to propose a new memory utilization mode for GraphChi, which is called Part-in-memory mode, to improve the GraphChi algorithm performance. The main idea is to pin a fixed part of data inside the memory during the whole computing process. Part-in-memory mode is successfully implemented with only about 40 additional lines of code to the original GraphChi engine. Extensive experiments are performed with large real datasets (including Twitter graph with 1.4 billion edges). The preliminary results show that Part-in-memory mode memory management approach effectively reduces the GraphChi running time by up to 60% in PageRank algorithm. Interestingly it is found that a larger portion of data pinned in memory does not always lead to better performance in the case that the whole dataset cannot be fitted in memory. There exists an optimal portion of data which should be kept in the memory to achieve the best computational performance.
Yifang Jiang, Diao Zhang, Kai Chen 0006, Qu Zhou, Yi Zhou 0003, Jianhua He 0001
IEEE BigData3
2013 A probability based subnet selection method for hot event detection in Sina Weibo microblogging
abstract
Microblogging has become a popular means of communication and information diffusion. Due to the huge amount of microblogs generated daily, the communication and computing costs required for real hot event detection is a big challenge. Choosing a small subnet of nodes to detect events has received increasing research interests in recent years. But the previous methods manage to select nodes to cover all the events including less popular events in sample datasets under the limited subnet size, which cause a big difference of event detection ratio between sample events and online real events in microblogs. In this paper we propose a new subnet nodes selection scheme based on the event detection ratio and nodes' events participation probabilities. Under the requirement of average event detection ratio, we prefer to choose the nodes who are active in propagating hot events than the nodes who participate in the less popular events. And we take dynamic programming to accelerate the computing. The experimental results show that our proposed method has a better performance.
Pei Shen, Yi Zhou 0003, Kai Chen 0006
ASONAM3
2013 Observation of Matthew Effects in Sina Weibo microblogger
abstract
This paper researches on Matthew Effect in Sina Weibo microblogger. We choose the microblogs in the ranking list of Hot Microblog App in Sina Weibo microblogger as target of our study. The differences of repost number of microblogs in the ranking list between before and after the time when it enter the ranking list of Hot Microblog app are analyzed. And we compare the spread features of the microblogs in the ranking list with those hot microblogs not in the list and those ordinary microblogs of users who have some microblog in the ranking list before. Our study proves the existence of Matthew Effect in social network.
Yi Zhou 0003, Qu Zhou, Kai Chen 0006, Jianhua He 0001, Xiaokang Yang 0001
IEEE BigData4
2012 Feature Analysis of Spammers in Social Networks with Active Honeypots: A Case Study of Chinese Microblogging Networks
abstract
In this poster we report our study on the microblog spammers with samples attracted by 50 honeyspots from two popular Chinese microblogging networks: Sina Weibo (weibo.com), and Ten cent Weibo (t.QQ.com) in seven months. We studied their features such as social information, activity, account age and spamming strategy. Several distinguishing characteristics of spammers on these two social network communities are observed, which can be helpful to the further study on automatic detection of microblog spammers. To our best knowledge our work is the first of its kind on the analysis of features of Chinese micloblog spammers.
Yi Zhou 0003, Kai Chen 0006, Li Song 0001, Xiaokang Yang 0001, Jianhua He 0001
ASONAM2
2012 Image super-resolution based on a novel edge sharpness prior
Yi Xu 0001, Xiaokang Yang 0001, Kai Chen 0006
ICPR4
2011 Building Artificial Identities in Social Network Using Semantic Information
abstract
As the popularity of social networking sites increase, so does their attractiveness for criminals. In this work, we show how an adversary can build artificial identities using semantic information in social network. Our method make the identities look more like real people, therefore can be used to support many kinds of attacks, such as ASE, profile cloning. A prototype of this method is implemented, includes following stages: Firstly, categories of virtual identity are predefined, and each category has multiple properties, such as geographical region, hobby, education, age, interested topic/keywords, etc. Secondly, based on category information, each identity will foster its own "life" semantically, such as edit profile and update status, find hot related news/topic from Google then post to wall, find related groups/networks then request to add in, and find/like/create/comment pages/posts, etc. Thirdly, artificial identity will evolve to multiple stages according to its status (for example, number of friends of real people), single identity with different evolutionary stages is linked together to a group that will help to ensure the number of attack edges.
Kai Chen 0006, Yi Zhou 0003, Li Song 0001, Xiaokang Yang 0001
ASONAM1
2011 Enhanced Slotted Aloha Protocols for Underwater Sensor Networks with Large Propagation Delay
abstract
Recently underwater sensor networks (UWSN) attracted large research interests. Medium access control (MAC) is one of the major challenges faced by UWSN due to the large propagation delay and narrow channel bandwidth of acoustic communications used for UWSN. Widely used slotted aloha (S-Aloha) protocol suffers large performance loss in UWSNs, which can only achieve performance close to pure aloha (PAloha). In this paper we theoretically model the performances of S-Aloha and P-Aloha protocols and analyze the adverse impact of propagation delay. According to the observation on the performances of S-Aloha protocol we propose two enhanced S-Aloha protocols in order to minimize the adverse impact of propagation delay on S-Aloha protocol. The first enhancement is a synchronized arrival S-Aloha (SA-Aloha) protocol, in which frames are transmitted at carefully calculated time to align the frame arrival time with the start of time slots. Propagation delay is taken into consideration in the calculation of transmit time. As estimation error on propagation delay may exist and can affect network performance, an improved SA-Aloha (denoted by ISAAloha) is proposed, which adjusts the slot size according to the range of delay estimation errors. Simulation results show that both SA-Aloha and ISA-Aloha perform remarkably better than S-Aloha and P-Aloha for UWSN, and ISA-Aloha is more robust even when the propagation delay estimation error is large.
Yi Zhou 0003, Kai Chen 0006, Jianhua He 0001, Haibing Guan
VTC Spring2
2010 Text Localization and Recognition in Complex Scenes Using Local Features
Kai Chen 0006, Yi Zhou 0003, Congcong Gu, Haibing Guan
ACCV (3)2
2010 Real-time Enhancement for Xen Hypervisor
abstract
System virtualization, which provides good isolation, is now widely used in server consolidation. Meanwhile, one of the hot topics in this field is to extend virtualization for embedded systems. However, current popular virtualization platforms do not support real-time operating systems such as embedded Linux well because the platform is not real-time ware, which will bring low-performance I/O and high scheduling latency. The goal of this paper is to optimize the Xen virtualization platform to be real-time operating system friendly. We improve two aspects of the Xen virtualization platform. First, we improve the xen scheduler to manage the scheduling latency and response time of the real-time operating system. Second, we import multiple real-time operating systems balancing method. Our experiment demonstrates that our enhancement to the Xen virtualization platform support real-time operating system well and the improvement to the real-time performance is about 20%.
Peijie Yu, Mingyuan Xia 0001, Qian Lin 0002, Shang Gao 0009, Zhengwei Qi, Kai Chen 0006, Haibing Guan
EUC7
2010 DistriBit: a distributed dynamic binary translator system for thin client computing
abstract
Although dynamic binary translators (DBT) are gaining popularity in the modern virtual execution environments (VEE), the requirement of DBTs' processing and memory resources has seriously hampered the performance of host platform. In this paper, we propose a distributed DBT system--DistriBit for resource-limited thin clients to overcome these challenges.
Haibing Guan, Yindong Yang, Kai Chen 0006, Yi Ge, Liang Liu 0010, Ying Chen 0004
HPDC3
2009 A Novel System for Robust Text Location and Recognition of Book Covers
Kaiyue Qi, Kai Chen 0006, Haibing Guan
ACCV (2)3
2009 A Hierarchical Localization Scheme for Large Scale Underwater Wireless Sensor Networks
abstract
In this paper, we study the localization problem in large-scale Underwater Wireless Sensor Networks (UWSNs). Unlike in the terrestrial positioning, the global positioning system (GPS) can not work efficiently underwater. The limited bandwidth, the severely impaired channel and the cost of underwater equipment all makes the localization problem very challenging. Most current localization schemes are not well suitable for deep underwater environment. We propose a hierarchical localization scheme to address the challenging problems. The new scheme mainly consists of four types of nodes, which are surface buoys, Detachable Elevator Transceivers (DETs), anchor nodes and ordinary nodes. Surface buoy is assumed to be equipped with GPS on the water surface. A DET is attached to a surface buoy and can rise and down to broadcast its position. The anchor nodes can compute their positions based on the position information from the DETs and the measurements of distance to the DETs. The hierarchical localization scheme is scalable, and can be used to make balances on the cost and localization accuracy. Initial simulation results show the advantages of our proposed scheme.
Yi Zhou 0003, Kai Chen 0006, Jianhua He 0001, Alei Liang
HPCC2