VLDB 2026 Research / reviewers in the wild / expert
Donglin Cao
dblp:04/915
· DBLP profile ↗
44ranked-venue papers
6as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Short video rumor detection based on causal graph
Donglin Cao, Xiong Tang, Yanghao Lin, Dazhen Lin |
Inf. Sci. | 1 |
| 2024 | Interpretable Short Video Rumor Detection Based on Modality TamperingabstractWith the rapid development of social media and short video applications in recent years, browsing short videos has become the norm. Due to its large user base and unique appeal, spreading rumors via short videos has become a severe social problem. Many methods simply fuse multimodal features for rumor detection, which lack interpretability. For short video rumors, rumor makers create rumors by modifying and/or splicing different modal information, so we should consider how to detect rumors from the perspective of modality tampering. Inspired by cross-modal contrastive learning, we propose a novel short video rumor detection framework by designing two pretraining tasks: modality tampering detection and inter-modal matching, imbuing the model with the ability to detect modality tampering and employing it for downstream rumor detection tasks. In addition, we design an interpretability mechanism to make the rumor detection results more reasonable by backtracking the model’s decision-making process. The experimental results show that the method on the short video rumor dataset has an improvement of about 4.6%-12% in macro-F1 compared with other models and can explain whether the short video is a rumor or not through the perspective of modality tampering. Kaixuan Wu, Yanghao Lin, Donglin Cao, Dazhen Lin |
LREC/COLING | 3 |
| 2024 | Leverage Causal Graphs and Rumor-Refuting Texts for Interpretable Rumor AnalysisabstractPrevious rumors detection study mostly ignores causal features in rumor texts and the interpretability of rumor classification results. Based on the phenomenon that rumors can cause changes or even loss of the original causal relationship in the truth, we can leverage causal features to help classify rumors. Meanwhile, rumor-refuting text is of great significance in curbing the spread of rumors and can interpret the results of rumor detection. To address these challenges, we propose a rumor analysis model integrating rumor detection and rumor refuting. It builds the causal graph of rumor texts and uses it to improve the performance of rumor detection. In particular, we embed the rumor-refuting text generation module in the model to realize the integration of rumor detection and rum -cation results’ analytical ability. We conducted experiments on two benchmark datasets and performed better than the state-of-the-art methods. Donglin Cao, Dazhen Lin |
ICASSP | 2 |
| 2024 | Evidence-Aware Multimodal Chinese Social Media Rumor DetectionabstractThe rapid proliferation of the internet has expedited the dissemination of multimedia rumor content on social media. Nevertheless, existing multimodal rumor detection approaches predominantly concentrate on the intrinsic features of the multimodal rumor content, lacking factual evidence support, which constrains their generalizability and rationality. Inspired by the task of Natural Language Inference (NLI), we propose a multimodal rumor-evidence-aware method in this paper. Given a multimodal rumor instance, we predict the extent of evidence support in the given instance for the multimodal rumor content via an evidence-aware model. To prove the effectiveness of this method, we have collected a Chinese Social Media Rumor dataset with Evidence (CSMRE) comprising 12,371 instances. Experimental results on CSMRE datasets show the superior performance of our method. Kaixuan Wu, Donglin Cao |
ICASSP | 2 |
| 2024 | End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image CaptioningabstractThe perspective of captions in fine-grained image captioning crucially impacts people’s perception and understanding of the image. However, existing methods often overlook this aspect, resulting in captions that struggle to accurately convey the image’s hierarchical and spatial information. In this paper, we propose an end-to-end Spatially-Constrained multi-perspective fine-grained Image Captioning (SCIC) model. SCIC initially predicts the optimal perspective for captioning the image and subsequently utilizes this optimal perspective as a constraint condition for caption generation. Furthermore, SCIC is capable of generating multi-perspective image captions based on customized perspectives. Experimental results show that our model effectively improves the state-of-the-art CIDEr score by about 14.38% and can generate multi-perspective fine-grained captions for the same image. Chunzhen Lin, Donglin Cao, Dazhen Lin |
ICASSP | 3 |
| 2024 | An autoencoder-based self-supervised learning for multimodal sentiment analysis
Wenjun Feng, Donglin Cao, Dazhen Lin |
Inf. Sci. | 3 |
| 2023 | Cross-Modality Earth Mover's Distance for Visible Thermal Person Re-identificationabstractVisible thermal person re-identification (VT-ReID) suffers from inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, however, it is usually restricted to the influence of the intra-identity variations. In this paper, we propose the Cross-Modality Earth Mover's Distance (CM-EMD) that can alleviate the impact of the intra-identity variations during modality alignment. CM-EMD selects an optimal transport strategy and assigns high weights to pairs that have a smaller intra-identity variation. In this manner, the model will focus on reducing the inter-modality discrepancy while paying less attention to intra-identity variations, leading to a more effective modality alignment. Moreover, we introduce two techniques to improve the advantage of CM-EMD. First, Cross-Modality Discrimination Learning (CM-DL) is designed to overcome the discrimination degradation problem caused by modality alignment. By reducing the ratio between intra-identity and inter-identity variances, CM-DL leads the model to learn more discriminative representations. Second, we construct the Multi-Granularity Structure (MGS), enabling us to align modalities from both coarse- and fine-grained levels with the proposed CM-EMD. Extensive experiments show the benefits of the proposed CM-EMD and its auxiliary techniques (CM-DL and MGS). Our method achieves state-of-the-art performance on two VT-ReID benchmarks. Yongguo Ling, Zhun Zhong, Zhiming Luo, Fengxiang Yang, Donglin Cao, Yaojin Lin, Shaozi Li, Nicu Sebe |
AAAI | 5 |
| 2023 | Multimodal Rumor Detection with Causal Graph Attention NetworkabstractIn the era of big data, the automatic detection and governance of rumors are of great significance in protecting the privacy and security of users. Many previous works have achieved good results in rumor detection through multimodal feature fusion, background knowledge expansion, etc. Causality in rumors is a potential textual feature representing rumors’ linguistic logic and causality semantics. The rumors’ causality often appears inconsistent with the truth, based on which we can improve the rumor detection model. However, previous work has neglected the use of the feature of causality between entities in rumor texts. To address these challenges, we propose a Causal Graph Attention Network (CGANet) for multimodal rumor detection, which constructs a causal graph for the rumor dataset to achieve better detection performance. CGANet extracts the node and causal relationship features of causal graphs through graph learning and enhances the causal semantics of multimodal features, thereby improving the detection ability of the model on false causal semantic content. Finally, we use a self-attention classification network to obtain the rumor detection results. We conducted experiments on two public benchmark datasets, Pheme and Weibo, and achieved an accuracy of 92.81% and 91.32%, respectively, better performance than the state-of-the-art methods. It proves the advantage of causal graphs in rumor detection tasks. Donglin Cao, Dazhen Lin |
ICPADS | 2 |
| 2023 | Multimodal Causal Relations Enhanced CLIP for Image-to-Text Retrieval
Wenjun Feng, Dazhen Lin, Donglin Cao |
PRCV (1) | 3 |
| 2023 | Text Causal Discovery Based on Sequence Structure Information
Donglin Cao, Dazhen Lin |
PRCV (7) | 2 |
| 2023 | Towards Robust Person Re-Identification by Defending Against Universal AttackersabstractRecent studies show that deep person re-identification (re-ID) models are vulnerable to adversarial examples, so it is critical to improving the robustness of re-ID models against attacks. To achieve this goal, we explore the strengths and weaknesses of existing re-ID models, i.e., designing learning-based attacks and training robust models by defending against the learned attacks. The contributions of this paper are three-fold: First, we build a holistic attack-defense framework to study the relationship between the attack and defense for person re-ID. Second, we introduce a combinatorial adversarial attack that is adaptive to unseen domains and unseen model types. It consists of distortions in pixel and color space (i.e., mimicking camera shifts). Third, we propose a novel virtual-guided meta-learning algorithm for our attack-defense system. We leverage a virtual dataset to conduct experiments under our meta-learning framework, which can explore the cross-domain constraints for enhancing the generalization of the attack and the robustness of the re-ID model. Comprehensive experiments on three large-scale re-ID benchmarks demonstrate that: 1) Our combinatorial attack is effective and highly universal in cross-model and cross-dataset scenarios; 2) Our meta-learning algorithm can be readily applied to different attack and defense approaches, which can reach consistent improvement; 3) The defense model trained on the learning-to-learn framework is robust to recent SOTA attacks that are not even used during training. Fengxiang Yang, Juanjuan Weng, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Donglin Cao, Shaozi Li, Shin'ichi Satoh 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Hypergraph-Based Reinforcement Learning for Stock Portfolio SelectionabstractStock portfolio selection is an important financial planning task that dynamically re-allocates the investments to stock assets to achieve the goals such as maximal profits and minimal risks. In this paper, we propose a hypergraph-based reinforcement learning method for stock portfolio selection, in which the fundamental issue is to learn a policy function generating appropriate trading actions given the current environments. The historical time-series patterns of stocks are firstly captured. Then, different from prior works ignoring or implicitly modeling stock pairwise correlations, we present a HyperGraph Attention Module (HGAM) in the portfolio policy learning, which utilizes the hypergraph structure to explicitly model the group-wise industry-belonging relationships among stocks. The attention mechanism is also introduced in HGAM that quantifies the importance of different neighbors regarding the target node to aggregate the information on the stock hypergraph adaptively. Extensive experiments on the real-world dataset collected from China’s A-share market demonstrate the significant superiority of our method, compared with state-of-the-art methods in portfolio selection, including both online learning-based methods and reinforcement learning-based methods. The data and codes of our work have been released at https://github.com/lixiaojieff/stock-portfolio. Chaoran Cui, Donglin Cao, Chunyun Zhang |
ICASSP | 3 |
| 2022 | A self-regulated generative adversarial network for stock price movement prediction based on the historical price and tweets
Hongfeng Xu, Donglin Cao, Shaozi Li |
Knowl. Based Syst. | 2 |
| 2021 | Grammar guided embedding based Chinese long text sentiment classificationabstractAbstract Although the state‐of‐the‐art sentiment classification approaches, such as LSTM and TextCNN, have achieved a good performance on Chinese short text sentiment analysis, the Chinese long text sentiment classification is still a challenge because of the sentiment change problem and the long text structure problem. Therefore, we propose a grammar guided embedding model (GGE) and a novel Chinese long text sentiment classification framework. First, the part‐of‐speech (POS) tags are introduced as the Chinese long text grammar guided information which can help classification approaches to model the Chinese long text structure and the important structure of sentiment change. Second, we proposed a simple GGE training method which considers the combination representation of word sequence and POS sequence. Finally, we proposed a unified framework which combines our novel GGE with TextCNN. Experiment results show that after using GGE, the model outperforms the state‐of‐the‐art approaches. At the same time, we also found that the GGE achieves the model converge faster, that is, it can achieve better results than without GGE when there is only a small amount of training data. Thus, we believe that the GGE can help machines better understand human language sentiment expression structure. Dazhen Lin, Donglin Cao, Shaozi Li |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Rumor knowledge embedding based data augmentation for imbalanced rumor detection
Xiangyan Chen, Duoduo Zhu, Dazhen Lin, Donglin Cao |
Inf. Sci. | 4 |
| 2019 | Chinese microblog rumor detection based on deep sequence contextabstractSummary Rumor is one of the main problems in social media, which often shows deeply and rapidly undesirable affection on the society. Although many rumor detection models consider content features and social features, all of them are based on the word independence assumption, which lacks the sequence context. Thus, if we use some words that often appear in rumors, our posts will be recognized as a rumor. To solve this problem, we propose a deep sequence context model (DSCM) for Chinese microblog rumor detection. This model considers two important factors of rumors: falsity and influence. Firstly, to learn falsity, we abolish the word independence assumption and use long short‐term memory (LSTM) units to capture bi‐direction sequence context information in content. Secondly, to learn influence, we combine the deep sequence context information with social features to learn the connection between content and social features. In our experiment, our results show that our approach outperforms several state‐of‐the‐art machine learning approaches, including term frequency and inverse document frequency (TFIDF), LSTM, and gated recurrent unit (GRU) in rumor detection. Dazhen Lin, Donglin Cao, Shaozi Li |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | Discriminative parts learning for 3D human action recognition
Min Huang 0004, Guo-Rong Cai, Hongbo Zhang 0002, Sheng Yu 0007, Dong-Ying Gong, Donglin Cao, Shaozi Li, Songzhi Su |
Neurocomputing | 6 |
| 2018 | Multi-modality weakly labeled sentiment learning based on Explicit Emotion Signal for Chinese microblog
Dazhen Lin, Donglin Cao, Yanping Lv, Xiao Ke |
Neurocomputing | 3 |
| 2018 | Chinese Character CAPTCHA Recognition and performance estimation via deep neural network
Dazhen Lin, Fan Lin, Yanping Lv, Feipeng Cai, Donglin Cao |
Neurocomputing | 5 |
| 2018 | Multi-label learning with label-specific features by resolving label correlations
Jia Zhang 0019, Candong Li, Donglin Cao, Yaojin Lin, Songzhi Su, Shaozi Li |
Knowl. Based Syst. | 3 |
| 2018 | Chinese microblog users' sentiment-based traffic condition analysis
Donglin Cao, Shuru Wang, Dazhen Lin |
Soft Comput. | 1 |
| 2018 | Predicting Microblog Sentiments via Weakly Supervised Multimodal Deep LearningabstractPredicting sentiments of multimodal microblogs composed of text, image, and emoticon have attracted ever-increasing research focus recently. The key challenge lies in the difficulty of collecting a sufficient amount of training labels to train a discriminative model for multimodal prediction. One potential solution is to exploit the labels collected from social media users, which is, however, restricted by the negative effect of label noise. Besides, we have quantitatively found that sentiments in different modalities may be independent, which disables the usage of previous multimodal sentiment analysis schemes in our problem. In this paper, we introduce a weakly supervised multimodal deep learning (WS-MDL) scheme toward robust and scalable sentiment prediction. WS-MDL learns convolutional neural networks iteratively and selectively from “weak” emoticon labels, which are cheaply available and noise containing. In particular, to filter out the label noise and to capture the modality dependency, a probabilistic graphical model is introduced to simultaneously learn discriminative multimodal descriptors and infer the confidence of label noise. Extensive evaluations are conducted in a million-scale, real-world microblog sentiment dataset crawled from Sina Weibo. We have validated the merits of the proposed scheme by quantitatively showing its superior performance over several state-of-the-art and alternative approaches. Fuhai Chen, Rongrong Ji, Jinsong Su, Donglin Cao, Yue Gao 0002 |
IEEE Trans. Multim. | 4 |
| 2018 | Multifeature Selection for 3D Human Action RecognitionabstractIn mainstream approaches for 3D human action recognition, depth and skeleton features are combined to improve recognition accuracy. However, this strategy results in high feature dimensions and low discrimination due to redundant feature vectors. To solve this drawback, a multi-feature selection approach for 3D human action recognition is proposed in this paper. First, three novel single-modal features are proposed to describe depth appearance, depth motion, and skeleton motion. Second, a classification entropy of random forest is used to evaluate the discrimination of the depth appearance based features. Finally, one of the three features is selected to recognize the sample according to the discrimination evaluation. Experimental results show that the proposed multi-feature selection approach significantly outperforms other approaches based on single-modal feature and feature fusion. Min Huang 0004, Songzhi Su, Hongbo Zhang 0002, Guo-Rong Cai, Dong-Ying Gong, Donglin Cao, Shaozi Li |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2017 | Re-ranking Person Re-identification with k-Reciprocal EncodingabstractWhen considering person re-identification (re-ID) as a retrieval process, re-ranking is a critical step to improve its accuracy. Yet in the re-ID community, limited effort has been devoted to re-ranking, especially those fully automatic, unsupervised solutions. In this paper, we propose a k-reciprocal encoding method to re-rank the re-ID results. Our hypothesis is that if a gallery image is similar to the probe in the k-reciprocal nearest neighbors, it is more likely to be a true match. Specifically, given an image, a k-reciprocal feature is calculated by encoding its k-reciprocal nearest neighbors into a single vector, which is used for re-ranking under the Jaccard distance. The final distance is computed as the combination of the original distance and the Jaccard distance. Our re-ranking method does not require any human interaction or any labeled data, so it is applicable to large-scale datasets. Experiments on the large-scale Market-1501, CUHK03, MARS, and PRW datasets confirm the effectiveness of our method. Zhun Zhong, Liang Zheng 0001, Donglin Cao, Shaozi Li |
CVPR | 3 |
| 2017 | Meta-action descriptor for action recognition in RGBD videoabstractAction recognition is one of the hottest research topics in computer vision. Recent methods represent actions based on global or local video features. These approaches, however, lack semantic structure and may not provide a deep insight into the essence of an action. In this work, the authors argue that semantic clues, such as joint positions and part‐level motion clustering, help verify actions. To this end, a meta‐action descriptor for action recognition in RGBD video is proposed in this study. Specifically, two discrimination‐based strategies – dynamic and discriminative part clustering – are introduced to improve accuracy. Experiments conducted on the MSR Action 3D dataset show that the proposed method significantly outperforms the methods without joint position semantic. Min Huang 0004, Songzhi Su, Guo-Rong Cai, Hongbo Zhang 0002, Donglin Cao, Shaozi Li |
IET Comput. Vis. | 5 |
| 2017 | Class-specific object proposals re-ranking for object detection in automatic driving
Zhun Zhong, Mingyi Lei, Donglin Cao, Jianping Fan 0001, Shaozi Li |
Neurocomputing | 3 |
| 2017 | Detecting ground control points via convolutional neural network for stereo matching
Zhun Zhong, Songzhi Su, Donglin Cao, Shaozi Li, Zhihan Lyu |
Multim. Tools Appl. | 3 |
| 2016 | A spatial-temporal visual mid-level ontology for GIF sentiment analysisabstractWith the progress of social medias, an increasing number of dynamic multimedia, such as video clips (GIF), were used in Internet. Many of them show the subjective sentiment of users. GIF sentiment analysis is important for political election predication and economic indicator evaluation. Unfortunately, such task is quite challenging because the users' sentiment hinges on spatial-temporal visual concepts. And the relationship between such concepts and overall sentiment polarity remains unknown. In this paper, dedicated to build a bridge to exploring such relationship, we proposed a SentiPair Sequence based spatial-temporal visual sentiment ontology. The ontology serves as a mid-level representation of GIF sentiment analysis. In order to analyze sentiment polarity, we first constructed a Synset Forest to define the semantic tree structure of visual sentiment concepts. Then, through the Synset Forest, we organically select and combine sentiment label elements to form a mid-level visual sentiment representation. Our experiments indicate that SentiPair outperforms Adjective Noun Pairs which is a state-of-art visual sentiment ontology for image. We also released our dataset (GSO-2015) to the research community. GSO-2015 contains more than 6,000 manually labeled GIFs and 40,000 unlabeled GIFs. Each is labeled with both sentiment and SentiPair Sequence. Zheng Cai, Donglin Cao, Dazhen Lin, Rongrong Ji |
CEC | 2 |
| 2016 | Chinese character CAPTCHA recognition based on convolution neural networkabstractCAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) are increasingly used in many applications for machine and human identification. Compared with traditional English and digital characters based CAPTCHAs, Chinese characters contain more complicated characters which greatly enhance difficulty of automatic recognition. To solve that problem, we proposed a Convolution Neural Network (CNN) based approach. This approach greatly improves the recognition accuracy of Chinese Character CAPTCHAs with distortion, rotation and background noise. Our experiment results show that this approach achieves more than 95% accuracy for single character and 84% accuracy for three types of Chinese Character CAPTCHAs with four characters. This encouraging result indicates that deep neural network is useful in complicated structure perception of Chinese Character CAPTCHAs. Yanping Lv, Feipeng Cai, Dazhen Lin, Donglin Cao |
CEC | 4 |
| 2016 | Survey of visual sentiment prediction for social media analysis
Rongrong Ji, Donglin Cao, Yiyi Zhou, Fuhai Chen |
Frontiers Comput. Sci. | 2 |
| 2016 | Detection based object labeling of 3D point cloud for indoor scenes
Wei Liu 0005, Shaozi Li, Donglin Cao, Songzhi Su, Rongrong Ji |
Neurocomputing | 3 |
| 2016 | A cross-media public sentiment analysis system for microblog
Donglin Cao, Rongrong Ji, Dazhen Lin, Shaozi Li |
Multim. Syst. | 1 |
| 2016 | Visual sentiment topic model based microblog image sentiment analysis
Donglin Cao, Rongrong Ji, Dazhen Lin, Shaozi Li |
Multim. Tools Appl. | 1 |
| 2015 | Sentiment analysis of Chinese micro-blog based on multi-modal correlation modelabstractText, emoticons and images, various modalities have been used to express users' feelings on social media, which significantly challenges traditional text-based sentiment analysis approaches. In this paper, we propose a Multi-modal Correlation Model (MCM) for multi-modal sentiment analysis. Compared with other multi-modal methods, MCM models hierarchical correlations among modalities, as well as between modalities and sentiments. Specifically, a probabilistic graphical model (PGM) is subsequently built upon the proposed MCM model, which considers the hierarchical correlations and preserves the classification ability of each modality. In order to compute the posterior probabilities of sentiments in PGM, we optimize the model by Maximum Likelihood Estimation. Experimental results demonstrate: 1) the hierarchical correlations among different modalities and sentiment; 2) the importance of hierarchical correlations to sentiment analysis. Donglin Cao, Shaozi Li, Rongrong Ji |
ICIP | 2 |
| 2015 | Multimodal hypergraph learning for microblog sentiment predictionabstractMicroblog sentiment analysis has attracted extensive research attention in the recent literature. However, most existing works mainly focus on the textual modality, while ignore the contribution of visual information that contributes ever increasing proportion in expressing user emotions. In this paper, we propose to employ a hypergraph structure to formulate textual, visual and emoticon information jointly for sentiment prediction. The constructed hypergraph captures the similarities of tweets on different modalities where each vertex represents a tweet and the hyperedge is formed by the “centroid” vertex and its k-nearest neighbors on each modality. Then, the transductive inference is conducted to learn the relevance score among tweets for sentiment prediction. In this way, both intra- and inter- modality dependencies are taken into consideration in sentiment prediction. Experiments conducted on over 6,000 microblog tweets demonstrate the superiority of our method by 86.77% accuracy and 7% improvement compared to the state-of-the-art methods. Fuhai Chen, Yue Gao 0002, Donglin Cao, Rongrong Ji |
ICME | 3 |
| 2015 | A Cross-media Sentiment Analytics Platform For MicroblogabstractIn this demo, a cross-media public sentiment analysis system is presented. The system presents and visualizes the sentiments of microblog data by organizing the results by region, topic, and content, respectively. Such sentiment is obtained by fusing of sentiment classification scores from both visual and textual channel. In such a way, social multimedia sentiment is shown in a multi-level and user-friendly form. Chao Chen 0026, Fuhai Chen, Donglin Cao, Rongrong Ji |
ACM Multimedia | 3 |
| 2014 | Hacking Chinese Touclick CAPTCHA by Multi-Scale Corner Structure Model with Fast Pattern MatchingabstractIn this paper, we tackle the challenge of hacking CAPTCHA (Completely Automated Public Turing test to tell Computers and Humans Apart), which is widely used to identify machine and human in webpage registration and authorization [17]. More specially, we target at touclick Chinese CAPTCHA that is recently popular in mobile application scenario. Hacking such CAPTCHA is much more challenging and left unexploited in the literature. Our main idea is a multi-scale Corner based Structure Model, termed CSM, with a very efficient pattern matching scheme. CSM can accurately capture the intrinsic statistics of touclick Chinese CAPTCHA against background clutters. We demonstrate the efficiency and effectiveness of the proposed approach by extensive experiments on a Chinese touclick CAPTCHA dataset collected from Internet forums. We report encouraging results with an overall success rate of almost 100% and an averaged detection speed of 170 millisecond. Upon our work, we also provide suggestions on improving the current CAPTCHA-based human-machine identification systems. Yunhang Shen, Rongrong Ji, Donglin Cao |
ACM Multimedia | 3 |
| 2014 | Perspective-Invariant Image Matching Framework with Binary Feature Descriptor and APSOabstractA novel perspective invariant image matching framework is proposed in this paper, noted as Perspective-Invariant Binary Robust Independent Elementary Features (PBRIEF). First, we use the homographic transformation to simulate the distortion between two corresponding patches around the feature points. Then, binary descriptors are constructed by comparing the intensity of sample points surrounding the feature location. We transform the location of the sample points with simulated homographic matrices. This operation is to ensure that the intensities which we compared are the realistic corresponding pixels between two image patches. Since the exact perspective transform matrix is unknown, an Adaptive Particle Swarm Optimization (APSO) algorithm-based iterative procedure is proposed to estimate the real transformation angles. Experimental results obtained on five different datasets show that PBRIEF outperforms significantly the existing methods on images with large viewpoint difference. Moreover, the efficiency of our framework is also improved comparing with Affine-Scale Invariant Feature Transform (ASIFT). Li-Chuan Geng, Songzhi Su, Donglin Cao, Shaozi Li |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2014 | Online semi-supervised compressive coding for robust visual tracking
Si Chen 0002, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | A new camera self-calibration method based on CSAabstractA large number of computer vision applications rely on camera calibration. Camera self-calibration which only depends on the relationship between corresponding points of a pair of images draws much attention for its simplicity. Almost all the camera self-calibration methods rely on the solution of Kruppa equations which are difficult to be directly solved. The state-of-the-art self-calibration algorithms usually convert the solution of these equations to non-linear optimization problem, traditional optimization methods usually have the drawback of convergent to local extreme. Artificial immune system (AIS) has the ability to fast convergent to global extreme. To address this problem, we proposed an artificial immune system based method which can fast convergent to the global optimization solutions. We demonstrate the performance of the proposed method with synthetic and real data. Li-Chuan Geng, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji |
VCIP | 4 |
| 2012 | A two-level model for automatic image annotation
Xiao Ke, Shaozi Li, Donglin Cao |
Multim. Tools Appl. | 3 |
| 2010 | Making intelligent business decisions by mining the implicit relation from bloggers' posts
Dazhen Lin, Shaozi Li, Donglin Cao |
Soft Comput. | 3 |
| 2008 | A Novel Language Model Based on Cognition Attention Attenuation in Web RetrievalabstractLanguage model is widely used in many retrieval systems. Its document representation is based on the bag of words assumption. Hence, each term in document is treated as an equal object and only the term frequency is considered as the evidence of the importance of term. In this paper, we study the problem of cognition attention attenuation in processing documents and present a cognition attention attenuation based language model. This model estimates the document model by attenuation process of term in document. Compared with the classical language model, the advantage of this model is considering about the document structure which is often used in text summarization. From the experiments results, our novel cognition attention attenuation based language model outperformed the classical language model with Dirichlet smoothing in blog page and Web page. Donglin Cao, Shuo Bai, Xueqi Cheng 0001, Shaozi Li |
Web Intelligence | 1 |
| 2006 | A Collaborative Multimedia Editing System Based on Shallow Nature Language Parsing
Donglin Cao, Dazhen Lin, Shaozi Li |
CDVE | 1 |