VLDB 2026 Research / reviewers in the wild / expert
Kai Wang 0001
dblp:78/2022-1
· DBLP profile ↗
43ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-5589-7060ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ODGNet: Learning Dynamic Variable Association for Online Time Series ForecastingabstractTime series forecasting is integral to diverse fields, significantly influencing both industrial production and human activities. In the real world, time-series data often exists in the form of streaming data and its distribution often changes over time, which is referred to as concept drift. This results in a gradual decline in the performance of traditional deep learning models as time progresses. Current online learning methods attempt to alleviate this issue by employing online adaptation methods in temporal dimension. However, they overlook the changes and drifts in the inter-variable association within time series. To this end, we propose a novel Online Dynamic Graph Network (ODGNet). ODGNet represents variable associations as a matrix polynomial and acquires polynomial coefficients based on online gradients, which models the evolutionary trends of spatial patterns. Furthermore, we emphasize the lossless information mapping between the adjacency matrix and its corresponding polynomial coefficient. Based on this mapping, a graph memory module with low memory consumption is proposed to avoid catastrophic forgetting when graph drift occurs. And a graph drift awareness mechanism is designed to rapidly detect graph drift. Experimental results demonstrate that ODGNet achieves a significant reduction in forecasting error compared to the existing online learning methods on twelve benchmark datasets. The code is available at https://github.com/LiuYasuo/ODGNet. Yushuo Liu, Kai Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionabstractWe present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabilities of the next-token prediction paradigm to conditional 3D object generation. To achieve this, the 3D VQ-VAE first encodes a wide range of 3D shapes into a compact triplane latent space and utilizes a set of discrete representations from a trainable codebook to reconstruct fine-grained geometries under the supervision of query point occupancy. Then, the 3D GPT, equipped with a custom triplane position embedding called TriPE, predicts the codebook index sequence with prefilling prompt tokens in an autoregressive manner so that the composition of 3D geometries can be modeled part by part. Extensive experiments on ShapeNet and Objaverse demonstrate that TAR3D can achieve superior generation quality over existing methods in text-to-3D and image-to-3D tasks Xuying Zhang, Yangguang Li 0001, Renrui Zhang, Kai Wang 0001, Wanli Ouyang, Zhiwei Xiong, Peng Gao 0007, Qibin Hou, Ming-Ming Cheng |
ICCV | 6 |
| 2025 | AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang 0001, Zhen Li 0031, Shaohui Jiao, Daquan Zhou, Qibin Hou, Ming-Ming Cheng |
ICCV | 3 |
| 2024 | TS-SAM: Fine-Tuning Segment-Anything Model for Downstream TasksabstractAdapter based fine-tuning has been studied for improving the performance of SAM on downstream tasks. However, there is still a significant performance gap between fine-tuned SAMs and domain-specific models. To reduce the gap, we propose Two-Stream SAM (TS-SAM). On the one hand, inspired by the side network in Parameter-Efficient Fine-Tuning (PEFT), we designed a lightweight Convolutional Side Adapter (CSA), which integrates the powerful features from SAM into side network training for comprehensive feature fusion. On the other hand, in line with the characteristics of segmentation tasks, we designed Multi-scale Refinement Module (MRM) and Feature Fusion Decoder (FFD) to keep both the detailed and semantic features. Extensive experiments on ten public datasets from three tasks demonstrate that TS-SAM not only significantly outperforms the recently proposed SAM-Adapter and SSOM, but achieves competitive performance with the SOTA domain-specific models. Our code is available at: https://github.com/maoyangou147/TS-SAM. Kai Wang 0001 |
ICME | 3 |
| 2024 | NKUT: Dataset and Benchmark for Pediatric Mandibular Wisdom Teeth SegmentationabstractGermectomy is a common surgery in pediatric dentistry to prevent the potential dangers caused by impacted mandibular wisdom teeth. Segmentation of mandibular wisdom teeth is a crucial step in surgery planning. However, manually segmenting teeth and bones from 3D volumes is time-consuming and may cause delays in treatment. Deep learning based medical image segmentation methods have demonstrated the potential to reduce the burden of manual annotations, but they still require a lot of well-annotated data for training. In this paper, we initially curated a Cone Beam Computed Tomography (CBCT) dataset, NKUT, for the segmentation of pediatric mandibular wisdom teeth. This marks the first publicly available dataset in this domain. Second, we propose a semantic separation scale-specific feature fusion network named WTNet, which introduces two branches to address the teeth and bones segmentation tasks. In WTNet, We design a Input Enhancement (IE) block and a Teeth-Bones Feature Separation (TBFS) block to solve the feature confusions and semantic-blur problems in our task. Experimental results suggest that WTNet performs better on NKUT compared to previous state-of-the-art segmentation methods (such as TransUnet), with a maximum DSC lead of nearly 16%. Zhenhuan Zhou, Along He, Xitao Que, Kai Wang 0001, Rui Yao 0010, Tao Li 0022 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Bilateral Supervision Network for Semi-Supervised Medical Image SegmentationabstractMassive high-quality annotated data is required by fully-supervised learning, which is difficult to obtain for image segmentation since the pixel-level annotation is expensive, especially for medical image segmentation tasks that need domain knowledge. As an alternative solution, semi-supervised learning (SSL) can effectively alleviate the dependence on the annotated samples by leveraging abundant unlabeled samples. Among the SSL methods, mean-teacher (MT) is the most popular one. However, in MT, teacher model's weights are completely determined by student model's weights, which will lead to the training bottleneck at the late training stages. Besides, only pixel-wise consistency is applied for unlabeled data, which ignores the category information and is susceptible to noise. In this paper, we propose a bilateral supervision network with bilateral exponential moving average (bilateral-EMA), named BSNet to overcome these issues. On the one hand, both the student and teacher models are trained on labeled data, and then their weights are updated with the bilateral-EMA, and thus the two models can learn from each other. On the other hand, pseudo labels are used to perform bilateral supervision for unlabeled data. Moreover, for enhancing the supervision, we adopt adversarial learning to enforce the network generate more reliable pseudo labels for unlabeled data. We conduct extensive experiments on three datasets to evaluate the proposed BSNet, and results show that BSNet can improve the semi-supervised segmentation performance by a large margin and surpass other state-of-the-art SSL methods. Along He, Tao Li 0022, Juncheng Yan, Kai Wang 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 4 |
| 2023 | H2Former: An Efficient Hierarchical Hybrid Transformer for Medical Image SegmentationabstractAccurate medical image segmentation is of great significance for computer aided diagnosis. Although methods based on convolutional neural networks (CNNs) have achieved good results, it is weak to model the long-range dependencies, which is very important for segmentation task to build global context dependencies. The Transformers can establish long-range dependencies among pixels by self-attention, providing a supplement to the local convolution. In addition, multi-scale feature fusion and feature selection are crucial for medical image segmentation tasks, which is ignored by Transformers. However, it is challenging to directly apply self-attention to CNNs due to the quadratic computational complexity for high-resolution feature maps. Therefore, to integrate the merits of CNNs, multi-scale channel attention and Transformers, we propose an efficient hierarchical hybrid vision Transformer (H2Former) for medical image segmentation. With these merits, the model can be data-efficient for limited medical data regime. The experimental results show that our approach exceeds previous Transformer, CNNs and hybrid methods on three 2D and two 3D medical image segmentation tasks. Moreover, it keeps computational efficiency in model parameters, FLOPs and inference time. For example, H2Former outperforms TransUNet by 2.29% in IoU score on KVASIR-SEG dataset with 30.77% parameters and 59.23% FLOPs. Along He, Kai Wang 0001, Tao Li 0022, Chengkun Du, Huazhu Fu |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Long-Term Person Re-identification with Dramatic Appearance Change: Algorithm and BenchmarkabstractFor person re-identification (Re-ID) task, most of previous studies assumed that the pedestrians do not change their appearances. The works on cross-appearance Re-ID, including datasets and algorithms, are still few. Therefore, this paper contributes a cross-season appearance change Re-ID dataset, namely NKUP+, including more than 300 IDs from surveillance videos over 10 months, to support the studies of the cross-appearance Re-ID. In addition, we propose a network named M2Net, which integrates multi-modality features from the RGB images, contour images and human parsing images. By ignoring irrelevant misleading information for cross-appearance retrieval in RGB images, M2Net can learn features that are robust to appearance changes. Meanwhile, we propose a sampling strategy called RAS to contain a variety of appearances in one batch. And appearance loss and multi-appearance loss are designed to guide the network to learn both same-appearance and cross-appearance features. Finally, we evaluated our method on NKUP+/PRCC/DeepChange datasets, and the results showed that, compared with the baseline, our method renders significant improvement, leading to the state-of-the-art performance over other methods. Our dataset is available at https://github.com/nkicsl/NKUP-dataset. Tao Li 0022, Yanfeng Jiang, Kai Wang 0001 |
ACM Multimedia | 5 |
| 2022 | Progressive Multiscale Consistent Network for Multiclass Fundus Lesion SegmentationabstractEffectively integrating multi-scale information is of considerable significance for the challenging multi-class segmentation of fundus lesions because different lesions vary significantly in scales and shapes. Several methods have been proposed to successfully handle the multi-scale object segmentation. However, two issues are not considered in previous studies. The first is the lack of interaction between adjacent feature levels, and this will lead to the deviation of high-level features from low-level features and the loss of detailed cues. The second is the conflict between the low-level and high-level features, this occurs because they learn different scales of features, thereby confusing the model and decreasing the accuracy of the final prediction. In this paper, we propose a progressive multi-scale consistent network (PMCNet) that integrates the proposed progressive feature fusion (PFF) block and dynamic attention block (DAB) to address the aforementioned issues. Specifically, PFF block progressively integrates multi-scale features from adjacent encoding layers, facilitating feature learning of each layer by aggregating fine-grained details and high-level semantics. As features at different scales should be consistent, DAB is designed to dynamically learn the attentive cues from the fused features at different scales, thus aiming to smooth the essential conflicts existing in multi-scale features. The two proposed PFF and DAB blocks can be integrated with the off-the-shelf backbone networks to address the two issues of multi-scale and feature inconsistency in the multi-class segmentation of fundus lesions, which will produce better feature representation in the feature space. Experimental results on three public datasets indicate that the proposed method is more effective than recent state-of-the-art methods. Along He, Kai Wang 0001, Tao Li 0022, Wang Bo, Hong Kang, Huazhu Fu |
IEEE Trans. Medical Imaging | 2 |
| 2021 | OSN: Onion-ring support neighbors for correspondence selection
Ye Lu 0004, Chunying Song, Tao Li 0022, Kai Wang 0001 |
Inf. Sci. | 5 |
| 2021 | PA-Net: Learning local features using by pose attention for short-term person re-identification
Kai Wang 0001, Junhui Yang, Tao Li 0002, Qinghua Hu |
Inf. Sci. | 1 |
| 2021 | Applications of deep learning in fundus images: A review
Tao Li 0022, Wang Bo, Hong Kang, Hanruo Liu, Kai Wang 0001, Huazhu Fu |
Medical Image Anal. | 6 |
| 2021 | CABNet: Category Attention Block for Imbalanced Diabetic Retinopathy GradingabstractDiabetic Retinopathy (DR) grading is challenging due to the presence of intra-class variations, small lesions and imbalanced data distributions. The key for solving fine-grained DR grading is to find more discriminative features corresponding to subtle visual differences, such as microaneurysms, hemorrhages and soft exudates. However, small lesions are quite difficult to identify using traditional convolutional neural networks (CNNs), and an imbalanced DR data distribution will cause the model to pay too much attention to DR grades with more samples, greatly affecting the final grading performance. In this article, we focus on developing an attention module to address these issues. Specifically, for imbalanced DR data distributions, we propose a novel Category Attention Block (CAB), which explores more discriminative region-wise features for each DR grade and treats each category equally. In order to capture more detailed small lesion information, we also propose the Global Attention Block (GAB), which can exploit detailed and class-agnostic global attention feature maps for fundus images. By aggregating the attention blocks with a backbone network, the CABNet is constructed for DR grading. The attention blocks can be applied to a wide range of backbone networks and trained efficiently in an end-to-end manner. Comprehensive experiments are conducted on three publicly available datasets, showing that CABNet produces significant performance improvements for existing state-of-the-art deep architectures with few additional parameters and achieves the state-of-the-art results for DR grading. Code and models will be available at https://github.com/he2016012996/CABnet. Along He, Tao Li 0022, Kai Wang 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 4 |
| 2020 | A benchmark for clothes variation in person re-identificationabstractPerson re-identification (re-ID) has drawn attention significantly in the computer vision society due to its application and research significance. It aims to retrieve a person of interest across different camera views. However, there are still several factors that hinder the applications of person re-ID. In fact, most common data sets either assume that pedestrians do not change their clothing across different camera views or are taken under constrained environments. Those constraints simplify the person re-ID task and contribute to early development of person re-ID, yet a person has a great possibility to change clothes in real life. To facilitate the research toward conquering those issues, this paper mainly introduces a new benchmark data set for person re-identification. To the best of our knowledge, this data set is currently the most diverse for person re-identification. It contains 107 persons with 9,738 images, captured in 15 indoor/outdoor scenes from September 2019 to December 2019, varying according to viewpoints, lighting, resolutions, human pose, seasons, backgrounds, and clothes especially. We hope that this benchmark data set will encourage further research on person re-identification with clothes variation. Moreover, we also perform extensive analyses on this data set using several state-of-the-art methods. Our dataset is available at https://github.com/nkicsl/NKUP-dataset. Kai Wang 0001, Shiyan Chen, Jinni Yang, Keke Zhou, Tao Li 0022 |
Int. J. Intell. Syst. | 1 |
| 2020 | Bin loss for hard exudates segmentation in fundus images
Song Guo 0002, Kai Wang 0001, Hong Kang, Yingqi Gao, Tao Li 0022 |
Neurocomputing | 2 |
| 2020 | Subspace Clustering via Good NeighborsabstractFinding the informative subspaces of high-dimensional datasets is at the core of numerous applications in computer vision, where spectral-based subspace clustering is arguably the most widely studied method due to its strong empirical performance. Such algorithms first compute an affinity matrix to construct a self-representation for each sample using other samples as a dictionary. Sparsity and connectivity of the self-representation play important roles in effective subspace clustering. However, simultaneous optimization of both factors is difficult due to their conflicting nature, and most existing methods are designed to address only one factor. In this paper, we propose a post-processing technique to optimize both sparsity and connectivity by finding good neighbors. Good neighbors induce key connections among samples within a subspace and not only have large affinity coefficients but are also strongly connected to each other. We reassign the coefficients of the good neighbors and eliminate other entries to generate a new coefficient matrix. We show that the few good neighbors can effectively recover the subspace, and the proposed post-processing step of finding good neighbors is complementary to most existing subspace clustering algorithms. Experiments on five benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art methods with negligible additional computation cost. Jufeng Yang, Jie Liang 0007, Kai Wang 0001, Paul L. Rosin, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Random Inception Module and Its Parallel Implementation
Yingqi Gao, Kunpeng Xie, Song Guo 0002, Kai Wang 0001, Hong Kang, Tao Li 0022 |
APPT | 4 |
| 2019 | A Lightweight Neural Network for Hard Exudate Segmentation of Fundus Image
Song Guo 0002, Tao Li 0022, Kai Wang 0001, Chan Zhang, Hong Kang |
ICANN (3) | 3 |
| 2019 | Random Drop Loss for Tiny Object Segmentation: Application to Lesion Segmentation in Fundus Images
Song Guo 0002, Tao Li 0022, Chan Zhang, Hong Kang, Kai Wang 0001 |
ICANN (3) | 6 |
| 2019 | Aggregation Connection Network For Tiny Face DetectionabstractFace detection has been greatly developed in recent years. Despite the remarkable progress, finding tiny faces in the wild is still a challenge due to the vastly scales, blur, occlusion and low resolution. This paper proposes an Aggregation Connection Network (ACN) which robustly solves these problems in tiny face detection. ACN utilizes the features from different convolution layers and performs superiorly on finding multi-scale faces in a single shot, especially for tiny faces. Specially, there are two novel modules in ACN that play significant roles: an aggregation connection module and a context module. First, by integrating efficient aggregation connection module, our ACN can effectively reduce the feature disappearance caused by image scaling. Second, the elaborately designed context module can make full use of the rich contextual cues without adding extra parameters. As a consequence, our ACN achieves state-of-the-art detection performance among several popular face detection benchmarks i.e. WIDER FACE, FDDB and Pascal Face. Chan Zhang, Tao Li 0022, Song Guo 0002, Yingqi Gao, Kai Wang 0001 |
IJCNN | 6 |
| 2019 | L-Seg: An end-to-end unified framework for multi-lesion segmentation of fundus images
Song Guo 0002, Tao Li 0022, Hong Kang, Yujun Zhang 0001, Kai Wang 0001 |
Neurocomputing | 6 |
| 2019 | Diagnostic assessment of deep learning algorithms for diabetic retinopathy screening
Tao Li 0022, Yingqi Gao, Kai Wang 0001, Song Guo 0002, Hanruo Liu, Hong Kang |
Inf. Sci. | 3 |
| 2019 | DCNR: deep cube CNN with random forest for hyperspectral image classification
Tao Li 0022, Jiabing Leng, Lingyan Kong, Song Guo 0002, Gang Bai, Kai Wang 0001 |
Multim. Tools Appl. | 6 |
| 2018 | Automatic Model Selection in Subspace Clustering via Triplet RelationshipsabstractThis paper addresses both the model selection (i.e., estimating the number of clusters K) and subspace clustering problems in a unified model. The real data always distribute on a union of low-dimensional sub-manifolds which are embedded in a high-dimensional ambient space. In this regard, the state-of-the-art subspace clustering approaches firstly learn the affinity among samples, followed by a spectral clustering to generate the segmentation. However, arguably, the intrinsic geometrical structures among samples are rarely considered in the optimization process. In this paper, we propose to simultaneously estimate K and segment the samples according to the local similarity relationships derived from the affinity matrix. Given the correlations among samples, we define a novel data structure termed the Triplet, each of which reflects a high relevance and locality among three samples which are aimed to be segmented into the same subspace. While the traditional pairwise distance can be close between inter-cluster samples lying on the intersection of two subspaces, the wrong assignments can be avoided by the hyper-correlation derived from the proposed triplets due to the complementarity of multiple constraints. Sequentially, we propose to greedily optimize a new model selection reward to estimate K according to the correlations between inter-cluster triplets. We simultaneously optimize a fusion reward based on the similarities between triplets and clusters to generate the final segmentation. Extensive experiments on the benchmark datasets demonstrate the effectiveness and robustness of the proposed approach. Jufeng Yang, Jie Liang 0007, Kai Wang 0001, Yongliang Yang 0002, Ming-Ming Cheng |
AAAI | 3 |
| 2018 | Sub-GAN: An Unsupervised Generative Model via Subspaces
Jie Liang 0007, Jufeng Yang, Hsin-Ying Lee 0001, Kai Wang 0001, Ming-Hsuan Yang 0001 |
ECCV (11) | 4 |
| 2018 | Learning hybrid convolutional features for edge detection
Xiaowei Hu 0003, Yun Liu 0011, Kai Wang 0001, Bo Ren 0003 |
Neurocomputing | 3 |
| 2018 | Dynamic Match Kernel With Deep Convolutional Features for Image RetrievalabstractFor image retrieval methods based on bag of visual words, much attention has been paid to enhancing the discriminative powers of the local features. Although retrieved images are usually similar to a query in minutiae, they may be significantly different from a semantic perspective, which can be effectively distinguished by convolutional neural networks (CNN). Such images should not be considered as relevant pairs. To tackle this problem, we propose to construct a dynamic match kernel by adaptively calculating the matching thresholds between query and candidate images based on the pairwise distance among deep CNN features. In contrast to the typical static match kernel which is independent to the global appearance of retrieved images, the dynamic one leverages the semantical similarity as a constraint for determining the matches. Accordingly, we propose a semantic-constrained retrieval framework by incorporating the dynamic match kernel, which focuses on matched patches between relevant images and filters out the ones for irrelevant pairs. Furthermore, we demonstrate that the proposed kernel complements recent methods, such as hamming embedding, multiple assignment, local descriptors aggregation, and graph-based re-ranking, while it outperforms the static one under various settings on off-the-shelf evaluation metrics. We also propose to evaluate the matched patches both quantitatively and qualitatively. Extensive experiments on five benchmark data sets and large-scale distractors validate the merits of the proposed method against the state-of-the-art methods for image retrieval. Jufeng Yang, Jie Liang 0007, Kai Wang 0001, Paul L. Rosin, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Richer Convolutional Features for Edge DetectionabstractIn this paper, we propose an accurate edge detector using richer convolutional features (RCF). Since objects in natural images possess various scales and aspect ratios, learning the rich hierarchical representations is very critical for edge detection. CNNs have been proved to be effective for this task. In addition, the convolutional features in CNNs gradually become coarser with the increase of the receptive fields. According to these observations, we attempt to adopt richer convolutional features in such a challenging vision task. The proposed network fully exploits multiscale and multilevel information of objects to perform the image-to-image prediction by combining all the meaningful convolutional features in a holistic manner. Using VGG16 network, we achieve state-of-the-art performance on several available datasets. When evaluating on the well-known BSDS500 benchmark, we achieve ODS F-measure of 0.811 while retaining a fast speed (8 FPS). Besides, our fast version of RCF achieves ODS F-measure of 0.806 with 30 FPS. Yun Liu 0011, Ming-Ming Cheng, Xiaowei Hu 0003, Kai Wang 0001, Xiang Bai |
CVPR | 4 |
| 2017 | Multi-scale energy optimization for object proposal generation
Congchao Wang, Jufeng Yang, Kai Wang 0001, Shang-Hong Lai |
Multim. Tools Appl. | 3 |
| 2017 | An intelligent character recognition method to filter spam images on cloud
Jufeng Yang, Tao Li 0022, Kai Wang 0001 |
Soft Comput. | 6 |
| 2016 | A Benchmark for Automatic Visual Classification of Clinical Skin Disease Images
Xiaoxiao Sun 0002, Jufeng Yang, Ming Sun 0006, Kai Wang 0001 |
ECCV (6) | 4 |
| 2016 | Shape-guided segmentation for fine-grained visual categorizationabstractIn this paper, we propose a shape-guided segmentation algorithm for fine-grained visual classification(FGVC). First, edge information is extracted from the query image and compared with each sample of training set, which can help us retrieve a subset of candidate proposals. These proposals are used to learn prior shape knowledge by separately estimating the foreground probabilities of corresponding pixels in the query image. Then, a redefined energy function is introduced to translate the minimum of energy to a good segmentation, with which we can dynamically pick out the most preferable proposal. After that, we obtain the label map of the image at the pixel level. Finally, the high-quality segmentation is used to aid locating semantic parts. We fine-tune one global model and two part models on Caffe to extract deep features and use a learned SVM classifier for categorization. We test three aspects in our experiment, including foreground segmentation, part localization and final classification. The results show that our method outperforms the state-of-the-art approaches on the famous Caltech-UCSD Birds 200-2011 dataset. Ming Sun 0006, Jufeng Yang, Kai Wang 0001 |
ICME | 4 |
| 2016 | Discovering affective regions in deep convolutional neural networks for visual sentiment predictionabstractIn this paper, we address the problem of automatically recognizing emotions in still images. While most of current work focus on improving whole-image representations using CNNs, we argue that discovering affective regions and supplementing local features will boost the performance, which is inspired by the observation that both global distributions and salient objects carry massive sentiments. We propose an algorithm to discover affective regions via deep framework, in which we use an off-the-shelf tool to generate N object proposals from a query image and rank these proposals with their objectness scores. Then, each proposal's sentiment score is computed using a pre-trained and fine-tuned CNN model. We combine both scores and select top K regions from the N candidates. These K regions are regarded as the most affective ones of the input image. Finally, we extract deep features from the whole-image and the selected regions, respectively, and sentiment label is predicted. The experiments show that our method is able to detect the affective local regions and achieve state-of-the-art performances on several popular datasets. Ming Sun 0006, Jufeng Yang, Kai Wang 0001 |
ICME | 3 |
| 2016 | HPSVM: Heterogeneous Parallel SVM with Factorization Based IPM Algorithm on CPU-GPU ClusterabstractSupport vector machine (SVM) is a supervised method widely used in the statistical classification and regression analysis. SVM training can be solved via the interior point method (IPM) with the advantages of low storage, fast convergence and easy parallelization. However, it is still confronted with the challenges of training speed and memory use. In this paper, we propose a parallel primal-dual IPM algorithm based on the incomplete Cholesky factorization (ICF) for efficiently training large-scale SVMs, named HPSVM, on CPU-GPU cluster. Our approach is distinguished from earlier work in that it is specifically designed to take maximal advantage of the CPU-GPU collaborative computation with the dual buffers 3-stage pipeline mechanism, and efficiently handles large-scale training datasets. In HPSVM, the heterogeneous hierarchical memory is fully explored to alleviate the bottleneck for optimizing data transfer, and the programming paradigm is presented to build an efficient collaboration mechanism between CPU and GPU. Comprehensive experiments show that HPSVM is up to 11 times faster than the CPU version on real datasets. Tao Li 0022, Xuechen Liu 0002, Qiankun Dong, Wenjing Ma, Kai Wang 0001 |
PDP | 5 |
| 2015 | Machine-readable region identification from partially blurred document imagesabstractPartial blur sometimes occurs in the document images captured by a camera, which will influence the performance of OCR on the non-blurred text region. A real-time method, named MRRI, is proposed in this paper to identify the machine-readable region from partially blurred document images. Firstly, a reference image is generated by low-pass filtering on the given document image. Secondly, a weight matrix is generated by calculating the structural similarity for each patch. Thirdly, a cost function is minimized to identify the maximum machine-readable region that can be well-recognized by OCR. In experiments, two applications are considered with the identified machine-readable region. On one hand, Tesseract-OCR is used for the word recognition to build index for a given document image. Compared with the results by applying OCR on the whole image, more words are correctly recognized by applying OCR on the identified region. On the other hand, the identified machine-readable region is used to assess the quality of a document image. Compared with other two image quality assessment methods, the machine-readable region based method shows a better performance. Also, MRRI is light and time-saving, which can meet the requirement of real-time applications. Qinwen Wang, Yixue Wang, Jufeng Yang, Tao Li 0022, Kai Wang 0001 |
ICDAR | 6 |
| 2015 | OCR with Adaptive Dictionary
Yanhong Xie, Kai Wang 0001, Tao Li 0022 |
ICIG (2) | 3 |
| 2012 | An improved binarization method using inter- and intra-block features for natural imagesabstractBinarization of natural images is important for text location and content-based analysis. In this work, a new adaptive method is introduced. It is able to improve the binarization results on the degraded images, such as the complex background, the non-uniform illumination, the variations of text font, size, color, and line orientation. The presented method contains three main stages. Firstly, original threshold of each pixel is calculated to produce some candidate blocks. Secondly, the new inter- and intra-block features are extracted from the candidates based on the characteristics of text. Finally, each block is scored from 0 to s using the mentioned features. The blocks with low scores are considered as subcomponents of background. After extensive experiments, our method demonstrated superior performance against two well-known techniques on the ICDAR 2005 competition dataset. Jufeng Yang, Kai Wang 0001, Jing Xu 0008 |
ICIP | 2 |
| 2012 | A fast adaptive binarization method for complex scene imagesabstractA novel adaptive binarization method based on wavelet filter is proposed in this paper, which shows comparable performance to other similar methods and processes faster, so that it is more suitable for real-time processing and applicable for mobile devices. The proposed method is evaluated on complex scene images of ICDAR 2005 Robust Reading Competition, and experimental results provide a support for our work. Jufeng Yang, Kai Wang 0001, Jiaofeng Li, Jing Xu 0008 |
ICIP | 2 |
| 2010 | A SVM-HMM Based Online Classifier for Handwritten Chemical SymbolsabstractThis paper presents a novel double-stage classifier for handwritten chemical symbols recognition task. The first stage is rough classification, SVM method is used to distinguish non-ring structure (NRS) and organic ring structure (ORS) symbols, while HMM method is used for fine recognition at second stage. A point-sequence-reordering algorithm is proposed to improve the recognition accuracy of ORS symbols. Our test data set contains 101 chemical symbols, 9090 training samples and 3232 test samples. Finally, we obtained top-1 accuracy of 93.10% and top-3 accuracy of 98.08% based on the test data set. Yang Zhang 0019, Guangshun Shi, Kai Wang 0001 |
ICPR | 3 |
| 2009 | High Performance Chinese/English Mixed OCR with Character Level Language IdentificationabstractCurrently, there have been several high performance OCR products for Chinese or for English. However, no one OCR technique can be simultaneously fit for both the English and the Chinese due to the large differences between Chinese and English. On the other hand, Chinese/English mixed document increases drastically with the globalization, so it is rather important to study the Chinese/English mixed document processing. Obviously, the key problem to resolve is how to split the mixed document into two parts: Chinese part and English part, so that the different OCR techniques can be applied to different parts. To further improve the previous system performance, a novel Chinese/English split algorithm based on global information is proposed and a rule for language identification is achieved by Bayesian formula. Experiment shows, the system error rate drops from 1.52% to 0.87% on magazine samples and from 1.32% to 0.75% on book samples, more than 2/5 of errors are excluded, which provides an experimental support for our research work. Kai Wang 0001, Jianming Jin, Qingren Wang |
ICDAR | 1 |
| 2009 | The asymptotic optimization of pre-edited ANN classifier
Kai Wang 0001, Jufeng Yang, Guangshun Shi, Qingren Wang |
Soft Comput. | 1 |
| 2008 | A study of on-line handwritten chemical expressions recognitionabstractIn this paper, we study the major modules of on-line handwritten chemical expressions recognition. We propose a novel two-level algorithm to recognize expressions. In the first level, structural information is used to distinguish different parts and recognize substances. Then the algorithm segments expressions fatherly and recognizes isolated symbols. To meet the demand of actual applications, the paper also designs an XML-based system to help users save, modify and search the recognition result. The experiment shows that the presented algorithm is reliable. Jufeng Yang, Guangshun Shi, Kai Wang 0001, Qian Geng, Qingren Wang |
ICPR | 3 |
| 2007 | A Margin Maximization Training Algorithm for BP Network
Kai Wang 0001, Qingren Wang |
ISNN (2) | 1 |