EDBT 2026 Demo / reviewers in the wild / expert
Congkai Sun
dblp:70/4555
· DBLP profile ↗
3ranked-venue papers
2as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Web and social media mining · 57% Data mining · 43% | |
| Network and information security
1 paper |
Network security · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Web and social media mining › web spam detection
link spam |
0.1 | 1 | 2011 | Let web spammers expose themselves · WSDM 2011 |
Data mining › anomaly detection › spam detection
spammer detection |
0.1 | 1 | 2011 | Let web spammers expose themselves · WSDM 2011 |
Web and social media mining
web spam detection |
0.1 | 1 | 2011 | Let web spammers expose themselves · WSDM 2011 |
Data mining › text mining
topic model |
0.1 | 1 | 2008 | HTM: A Topic Model for Hypertexts · EMNLP 2008 |
Network security
spam |
0.0 | 1 | 2011 | Let web spammers expose themselves · WSDM 2011 |
Methods — techniques the papers use, named apart from their topics
semi-supervised learning · 0.2feature extraction · 0.2topic modeling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Dialogue Act Recognition using Reweighted Speaker Adaptation
Congkai Sun, Louis-Philippe Morency |
SIGDIAL Conference | 1 |
| 2011 | Let web spammers expose themselvesabstractThis paper is concerned with mining link spams (e.g., link farm and link exchange) from search engine optimization (SEO) forums. To provide quality services, it is critical for search engines to address web spam. Several techniques such as TrustRank, BadRank, and SpamRank have been proposed for this purpose. Most of these methods try to downgrade the effects of the spam websites by identifying specific link patterns of them. However, spam websites have appeared to be more and more similar to normal or even good websites in their link structures, by reforming their spam techniques. As a result, it is very challenging to automatically detect link spams from the Web graph. In this paper, we propose a different approach, which detects link spams by looking at how web spammers make link spam happen. We find that web spammers usually ally with each other, and SEO forum is one of the major means for them to form the alliance. We therefore propose mining suspicious link spams directly from the posts in the SEO forums. However, the task is non-trivial because there are also other information and even noises contained in these posts, in addition to useful clues of link spam. To tackle the challenges, we first extract all the URLs contained in the posts of the SEO forums. Second, we extract features for the URLs from their relationships with forum users (potential spammers) and from their link structure in the web graph. Third, we build a semi-supervised learning framework to calculate the spam scores for the URLs, which encodes several heuristics such as spam websites usually linking to each other, and good websites seldom linking to spam websites. We tested our approach on seven major SEO forums. A lot of spam websites were identified, a significant proportion of which cannot be detected by conventional anti-spam methods. It indicates that the proposed approach can be a good complement of existing anti-spam techniques. Zhicong Cheng, Bin Gao 0001, Congkai Sun, Yanbing Jiang, Tie-Yan Liu |
WSDM | 3 |
| 2008 | HTM: A Topic Model for Hypertexts
Congkai Sun, Bin Gao 0001, Zhenfu Cao, Hang Li 0001 |
EMNLP | 1 |