VLDB 2026 Research / reviewers in the wild / expert
Chao Shao
dblp:12/1176
· DBLP profile ↗
18ranked-venue papers
6as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 67% Language models and text generation · 27% Probabilistic and Bayesian machine learning · 6% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › event analysis
event prediction |
0.3 | 1 | 2017 | What Happens Next? Future Subevent Prediction Using Contextual Hierarchical LSTM · AAAI 2017 |
Natural language and speech › Language models and text generation
text generation |
0.3 | 1 | 2017 | What Happens Next? Future Subevent Prediction Using Contextual Hierarchical LSTM · AAAI 2017 |
Natural language and speech › Information extraction and text analysis › text mining
text clustering |
0.2 | 1 | 2015 | TSDPMM: Incorporating Prior Topic Knowledge into Dirichlet Process Mixture Models for Text Clustering · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 1 | 2015 | TSDPMM: Incorporating Prior Topic Knowledge into Dirichlet Process Mixture Models for Text Clustering · EMNLP 2015 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model |
0.1 | 1 | 2015 | TSDPMM: Incorporating Prior Topic Knowledge into Dirichlet Process Mixture Models for Text Clustering · EMNLP 2015 |
Mathematical optimization
separation algorithms |
0.0 | 1 | 2004 | Some properties of the alternating separation (AS), alternating projection (AP) and ASAP algorithm · Sci. China Ser. F Inf. Sci. 2004 |
Methods — techniques the papers use, named apart from their topics
topic modeling · 0.3hierarchical LSTM · 0.3seeded pólya urn scheme · 0.2dirichlet process mixture model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PSA-Swin Transformer: Image Classification on Small-Scale DatasetsabstractThis paper introduces the PSA-Swin Transformer, a novel framework for image classification on small-scale datasets, highlighting the challenges of training effective models in resource-limited environments. Recognizing the limitations of current deep learning methods that rely heavily on large-scale datasets and extensive pre-training, we propose a method for handling small datasets. Our model can effectively han-dle smaller data volumes without the need for pre-training weights. The key to our approach is in the introduction of an efficient positional embedding (EPE) module, which improves parameter utilization and network expressiveness through a grouped convolutional architecture and shuffling operations for dynamic information exchange. In addition, we integrated the Polarized Self-Attention (PSA) module in Windows Multi-Head Self-Attention (W-MSA) and named the new module PSA-W-MSA; PSA addresses the complexity of learning element-specific attention by combining polarization filtering with augmentation techniques. Through a series of experiments on the Mini-Imagenet dataset, the PSA-Swin Transformer demonstrates notable performance, especially in environments where high-quality annotated data is scarce or costly to acquire. Our research results are expected to make progress in areas that require efficient and accurate image classification using limited resources. Chao Shao, Shaochen Jiang |
SMC | 1 |
| 2024 | Climate change characteristics and population health impact factors using deep neural network and hyperautomation mechanism
Chao Shao, Hairui Zhang |
J. Supercomput. | 1 |
| 2023 | Center Weighted Convolution and GraphSAGE Cooperative Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is one of the basic tasks of remote sensing image processing, which is to predict the label of each HSI pixel. Convolution neural network (CNN) and graph convolution neural network (GCN) have become the current research focus due to their outstanding performance in the field of HSI classification in recent years. However, GCN is a transductive learning method, which needs all nodes to participate in the training process to get the node embedding. Graph sample and aggregation (GraphSAGE) is an important branch of graph neural network, which can flexibly aggregate new neighbor nodes in non-Euclidean data of any structure, and capture long-range contextual relationships. Superpixel-based GraphSAGE can not only integrate the global spatial relationship of data, but also further reduce its computing cost. CNN can extract pixel-level features in a small area, and our center attention module (CAM) and center weighted convolution (CW-Conv) can also improve the feature extraction ability of CNN by enhancing the dominant position of target pixels. In order to make full use of the advantages of CNN and GraphSAGE, we propose a center weighted convolution and GraphSAGE (CW-SAGE) cooperative network for HSI classification. Specifically, graph simple and aggregate branch is constructed by superpixel-based encoder and decoder modules, then pixel-level features are extracted by a central attention convolutional neural network. Finally, the features of the two branches are spliced together for feature fusion. We conduct experiments on three hyperspectral datasets and compare the results with other current state-of-the-art methods. A series of experiments demonstrate the advantages of our method. Chao Shao, Liguo Wang 0001, Shan Gao 0007 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | A neural model for joint event detection and prediction
Linmei Hu, Shuqi Yu, Bin Wu 0001, Chao Shao, Xiaoli Li 0001 |
Neurocomputing | 4 |
| 2020 | Graph neural news recommendation with long-term and short-term interest modeling
Linmei Hu, Chuan Shi 0001, Cheng Yang 0002, Chao Shao |
Inf. Process. Manag. | 5 |
| 2020 | Graph neural entity disambiguation
Linmei Hu, Chuan Shi 0001, Chao Shao |
Knowl. Based Syst. | 4 |
| 2017 | What Happens Next? Future Subevent Prediction Using Contextual Hierarchical LSTMabstractEvents are typically composed of a sequence of subevents. Predicting a future subevent of an event is of great importance for many real-world applications. Most previous work on event prediction relied on hand-crafted features and can only predict events that already exist in the training data. In this paper, we develop an end-to-end model which directly takes the texts describing previous subevents as input and automatically generates a short text describing a possible future subevent. Our model captures the two-level sequential structure of a subevent sequence, namely, the word sequence for each subevent and the temporal order of subevents. In addition, our model incorporates the topics of the past subevents to make context-aware prediction of future subevents. Extensive experiments on a real-world dataset demonstrate the superiority of our model over several state-of-the-art methods. Linmei Hu, Juan-Zi Li, Liqiang Nie, Xiaoli Li 0001, Chao Shao |
AAAI | 5 |
| 2016 | RiMOM-IM: A Novel Iterative Framework for Instance Matching
Chao Shao, Linmei Hu, Juan-Zi Li, Zhichun Wang, Tong Lee Chung, Jun-Bo Xia |
J. Comput. Sci. Technol. | 1 |
| 2016 | Errata and comments on "Errata and comments on Orthogonal moments based on exponent functions: Exponent-Fourier moments"
Haitao Hu, Quan Ju, Chao Shao |
Pattern Recognit. | 3 |
| 2015 | TSDPMM: Incorporating Prior Topic Knowledge into Dirichlet Process Mixture Models for Text ClusteringabstractDirichlet process mixture model (DPM-M) has great potential for detecting the underlying structure of data.Extensive studies have applied it for text clustering in terms of topics.However, due to the unsupervised nature, the topic clusters are always less satisfactory.Considering that people often have some prior knowledge about which potential topics should exist in given data, we aim to incorporate such knowledge into the DPMM to improve text clustering.We propose a novel model TSDPMM based on a new seeded Pólya urn scheme.Experimental results on document clustering across three datasets demonstrate our proposed TSDPMM significantly outperforms stateof-the-art DPMM model and can be applied in a lifelong learning framework. Linmei Hu, Juan-Zi Li, Xiaoli Li 0001, Chao Shao, Xuzhong Wang |
EMNLP | 4 |
| 2015 | o-HETM: An Online Hierarchical Entity Topic Model for News Streams
Linmei Hu, Juan-Zi Li, Jing Zhang 0036, Chao Shao |
PAKDD (1) | 4 |
| 2015 | Incremental learning from news events
Linmei Hu, Chao Shao, Juan-Zi Li, Heng Ji 0001 |
Knowl. Based Syst. | 2 |
| 2014 | Orthogonal moments based on exponent functions: Exponent-Fourier moments
Haitao Hu, Ya-dong Zhang, Chao Shao, Quan Ju |
Pattern Recognit. | 3 |
| 2013 | Incorporating Entities in News Topic Modeling
Linmei Hu, Juan-Zi Li, Chao Shao |
NLPCC | 4 |
| 2007 | Selection of the Suitable Neighborhood Size for the ISOMAP AlgorithmabstractThe success of ISOMAP depends greatly on selecting a suitable neighborhood size; however, it's an open problem how to do this efficiently. When the neighborhood size is unsuitable, shortcut edges can emerge in the neighborhood graph and shorten the involved shortest path lengths greatly, which makes them not approximate the corresponding geodesic distances anymore, that is, there doesn't exist such an approximately monotonically increasing relationship between them anymore. Based on this observation, in the paper, we use costs over the minimal connected neighborhood graph to approximate the corresponding geodesic distances, and then present an efficient method to judge whether a neighborhood size is suitable beforehand, by which a suitable neighborhood size can be selected more efficiently than the straightforward method with the residual variance. Besides, the correctness of the intrinsic dimensionality, estimated by ISOMAP, of the data can also be judged more easily by our method. Chao Shao, Houkuan Huang, Chunhong Wan |
IJCNN | 1 |
| 2006 | Combining Label Information and Neighborhood Graph for Semi-supervised Learning
Lianwei Zhao, Siwei Luo, Chao Shao, Hongliang Ma |
ISNN (1) | 4 |
| 2004 | Improvement of Data Visualization Based on SOM
Chao Shao, Houkuan Huang |
ISNN (2) | 1 |
| 2004 | Some properties of the alternating separation (AS), alternating projection (AP) and ASAP algorithm
Chao Shao, Guangyue Lu, Zheng Bao 0001 |
Sci. China Ser. F Inf. Sci. | 1 |