EDBT 2026 Demo / reviewers in the wild / expert
Shaohua Sun
dblp:03/7009
· DBLP profile ↗
12ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Knowledge graphs · 74% Data integration and cleaning · 26% | |
| Artificial intelligence
3 papers |
Knowledge representation and reasoning · 68% Question answering and dialogue systems · 32% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs
knowledge graph construction |
0.4 | 2 | 2014 | From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 Knowledge vault: a web-scale approach to probabilistic knowledge fusion · KDD 2014 |
Data integration and cleaning
truth discovery |
0.2 | 1 | 2015 | Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources · Proc. VLDB Endow. 2015 |
Data integration and cleaning
data fusion |
0.2 | 1 | 2014 | From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 |
Knowledge graphs
knowledge base integration |
0.2 | 1 | 2014 | From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 |
Knowledge graphs › knowledge graph construction
knowledge extraction |
0.2 | 1 | 2014 | Knowledge vault: a web-scale approach to probabilistic knowledge fusion · KDD 2014 |
Knowledge graphs
link prediction |
0.2 | 1 | 2014 | Knowledge base completion via search-based question answering · WWW 2014 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.1 | 1 | 2015 | Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources · Proc. VLDB Endow. 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge extraction |
0.1 | 1 | 2014 | From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 |
Natural language and speech › Question answering and dialogue systems › open-domain question answering
web question answering |
0.1 | 1 | 2014 | Knowledge base completion via search-based question answering · WWW 2014 |
Methods — techniques the papers use, named apart from their topics
multi-layer probabilistic model · 0.4joint inference · 0.4query learning · 0.4probabilistic aggregation · 0.4data fusion techniques · 0.4supervised machine learning · 0.2probabilistic inference · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Cross-Scale Spatial Feature Aggregation for Remote Sensing Change Detection
Shaohua Sun |
ICIC | 2 |
| 2026 | GS-DMSR: Dynamic Sensitive Multi-scale Manifold Enhancement for Accelerated High-Quality 3D Gaussian Splatting
Nengbo Lu, Minghua Pan, Shaohua Sun, Yizhou Liang |
MMM (1) | 3 |
| 2026 | An enhanced total variation regularization model for phase retrieval
Shaohua Sun, Benxin Zhang |
Signal Process. | 1 |
| 2025 | A Deep Learning Model for Water Quality Prediction Based on Temporal Decomposition and Feature Enhancement Learning
Chuchu Tan, Deke Wang, Liyi Guo, Liangmei Huang, Kamarul Hawari Ghazali, Shaohua Sun, Shiming Shen |
IEEE Big Data | 9 |
| 2024 | SymMatch: Symmetric Bi-Scale Matching with Self-Knowledge Distillation in Semi-Supervised Medical Image SegmentationabstractWith the development of medical image segmentation technology, high-quality automatic segmentation methods, particularly within semi-supervised learning frameworks, have become a research hotspot. This study introduces a new semi-supervised medical image segmentation algorithm called SymMatch. The algorithm effectively leverages limited labeled data along with a large amount of unlabeled data through a symmetrical network structure and knowledge distillation techniques. SymMatch applies a spectrum of perturbations, from weak to strong, at both image and feature levels, effectively leveraging the potential of unlabeled data. Additionally, by incorporating a bi-scale distillation loss, the model’s robustness and accuracy in handling complex medical imaging data are further enhanced. Experimental results show that SymMatch demonstrates superior performance across multiple recognized medical imaging datasets (such as ACDC, LA and PanNuke). Notably, even with very limited labeled data, it maintains high segmentation accuracy. These achievements not only advance the development of semi-supervised medical image segmentation technology but also provide new ideas and methods for future research in related technologies. Code is available at https://github.com/AiEson/SymMatch. Chunshi Wang, Shougan Teng, Shaohua Sun, Bin Zhao 0007 |
BIBM | 3 |
| 2024 | A fast interpretable adaptive meta-learning enhanced deep learning framework for diagnosis of diabetic retinopathy
Maofa Wang, Qizhou Gong, Zhixiong Leng, Yanlin Xu, Bingchen Yan, Hongliang Huang, Shaohua Sun |
Expert Syst. Appl. | 9 |
| 2024 | Meta-learning of feature distribution alignment for enhanced feature sharing
Zhixiong Leng, Maofa Wang, Yanlin Xu, Bingchen Yan, Shaohua Sun |
Knowl. Based Syst. | 6 |
| 2015 | Knowledge-Based Trust: Estimating the Trustworthiness of Web SourcesabstractThe quality of web sources has been traditionally evaluated using exogenous signals such as the hyperlink structure of the graph. We propose a new approach that relies on endogenous signals, namely, the correctness of factual information provided by the source. A source that has few false facts is considered to be trustworthy. The facts are automatically extracted from each source by information extraction methods commonly used to construct knowledge bases. We propose a way to distinguish errors made in the extraction process from factual errors in the web source per se, by using joint inference in a novel multi-layer probabilistic model. We call the trustworthiness score we computed Knowledge-Based Trust (KBT) . On synthetic data, we show that our method can reliably compute the true trustworthiness levels of the sources. We then apply it to a database of 2.8B facts extracted from the web, and thereby estimate the trustworthiness of 119M webpages. Manual evaluation of a subset of the results confirms the effectiveness of the method. Xin Dong 0001, Evgeniy Gabrilovich, Kevin Murphy 0002, Van Dang, Wilko Horn, Camillo Lugaresi, Shaohua Sun, Wei Zhang 0152 |
Proc. VLDB Endow. | 7 |
| 2014 | Knowledge vault: a web-scale approach to probabilistic knowledge fusionabstractRecent years have witnessed a proliferation of large-scale knowledge bases, including Wikipedia, Freebase, YAGO, Microsoft's Satori, and Google's Knowledge Graph. To increase the scale even further, we need to explore automatic methods for constructing knowledge bases. Previous approaches have primarily focused on text-based extraction, which can be very noisy. Here we introduce Knowledge Vault, a Web-scale probabilistic knowledge base that combines extractions from Web content (obtained via analysis of text, tabular data, page structure, and human annotations) with prior knowledge derived from existing knowledge repositories. We employ supervised machine learning methods for fusing these distinct information sources. The Knowledge Vault is substantially bigger than any previously published structured knowledge repository, and features a probabilistic inference system that computes calibrated probabilities of fact correctness. We report the results of multiple studies that explore the relative utility of the different information sources and extraction methods. Xin Dong 0001, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Ni Lao, Kevin Murphy 0002, Thomas Strohmann, Shaohua Sun, Wei Zhang 0152 |
KDD | 8 |
| 2014 | Knowledge base completion via search-based question answeringabstractOver the past few years, massive amounts of world knowledge have been accumulated in publicly available knowledge bases, such as Freebase, NELL, and YAGO. Yet despite their seemingly huge size, these knowledge bases are greatly incomplete. For example, over 70% of people included in Freebase have no known place of birth, and 99% have no known ethnicity. In this paper, we propose a way to leverage existing Web-search-based question-answering technology to fill in the gaps in knowledge bases in a targeted way. In particular, for each entity attribute, we learn the best set of queries to ask, such that the answer snippets returned by the search engine are most likely to contain the correct value for that attribute. For example, if we want to find Frank Zappa's mother, we could ask the query `who is the mother of Frank Zappa'. However, this is likely to return `The Mothers of Invention', which was the name of his band. Our system learns that it should (in this case) add disambiguating terms, such as Zappa's place of birth, in order to make it more likely that the search results contain snippets mentioning his mother. Our system also learns how many different queries to ask for each attribute, since in some cases, asking too many can hurt accuracy (by introducing false positives). We discuss how to aggregate candidate answers across multiple queries, ultimately returning probabilistic predictions for possible values for each attribute. Finally, we evaluate our system and show that it is able to extract a large number of facts with high confidence. Robert West 0001, Evgeniy Gabrilovich, Kevin Murphy 0002, Shaohua Sun, Dekang Lin |
WWW | 4 |
| 2014 | From Data Fusion to Knowledge FusionabstractThe task of data fusion is to identify the true values of data items ( e.g. , the true date of birth for Tom Cruise ) among multiple observed values drawn from different sources ( e.g. , Web sites) of varying (and unknown) reliability. A recent survey [20] has provided a detailed comparison of various fusion methods on Deep Web data. In this paper, we study the applicability and limitations of different fusion techniques on a more challenging problem: knowledge fusion . Knowledge fusion identifies true subject-predicate-object triples extracted by multiple information extractors from multiple information sources. These extractors perform the tasks of entity linkage and schema alignment, thus introducing an additional source of noise that is quite different from that traditionally considered in the data fusion literature, which only focuses on factual errors in the original sources. We adapt state-of-the-art data fusion techniques and apply them to a knowledge base with 1.6B unique knowledge triples extracted by 12 extractors from over 1B Web pages, which is three orders of magnitude larger than the data sets used in previous data fusion papers. We show great promise of the data fusion approaches in solving the knowledge fusion problem, and suggest interesting research directions through a detailed error analysis of the methods. Xin Dong 0001, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Kevin Murphy 0002, Shaohua Sun, Wei Zhang 0152 |
Proc. VLDB Endow. | 6 |
| 2008 | Learning-enhanced simulated annealing: method, evaluation, and application to lung nodule registration
Shaohua Sun, Feng Zhuge, Jarrett Rosenberg, Robert M. Steiner, Geoffrey D. Rubin, Sandy Napel |
Appl. Intell. | 1 |