VLDB 2026 Research / reviewers in the wild / expert
Craig Martell
dblp:07/3213
· DBLP profile ↗
4ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 50% Information extraction and text analysis · 50% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.2 | 1 | 2015 | Transfer Learning for Bilingual Content Classification · KDD 2015 |
Natural language and speech › Information extraction and text analysis › text classification
spam detection |
0.2 | 1 | 2015 | Transfer Learning for Bilingual Content Classification · KDD 2015 |
Methods — techniques the papers use, named apart from their topics
transfer learning · 0.2machine translation · 0.2feature generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Transfer Learning for Bilingual Content ClassificationabstractLinkedIn Groups provide a platform on which professionals with similar background, target and specialities can share content, take part in discussions and establish opinions on industry topics. As in most online social communities, spam content in LinkedIn Groups poses great challenges to the user experience and could eventually lead to substantial loss of active users. Building an intelligent and scalable spam detection system is highly desirable but faces difficulties such as lack of labeled training data, particularly for languages other than English. In this paper, we take the spam (Spanish) job posting detection as the target problem and build a generic machine learning pipeline for multi-lingual spam detection. The main components are feature generation and knowledge migration via transfer learning. Specifically, in the feature generation phase, a relatively large labeled data set is generated via machine translation. Together with a large set of unlabeled human written Spanish data, unigram features are generated based on the frequency. In the second phase, machine translated data are properly reweighted to capture the discrepancy from human written ones and classifiers can be built on top of them. To make effective use of a small portion of labeled data available in human written Spanish, an adaptive transfer learning algorithm is proposed to further improve the performance. We evaluate the proposed method on LinkedIn's production data and the promising results verify the efficacy of our proposed algorithm. The pipeline is ready for production. Qian Sun 0002, Mohammad Shafkat Amin, Baoshi Yan, Craig Martell, Vita Markman, Anmol Bhasin, Jieping Ye |
KDD | 4 |
| 2004 | Talkbank: Building an Open Unified Multimodal Database of Communicative Interaction
Brian MacWhinney, Steven Bird, Christopher Cieri, Craig Martell |
LREC | 4 |
| 2002 | FORM: an extensible, kinematically-based gesture annotation schemeabstractAnnotated corpora have played a critical role in speech and natural language research; and, there is an increasing interest in corpora-based research in sign language and gesture as well.As examples, consider the tools Anvil and MediaTagger.These are excellent tools which allow for multi-track annotation of videos of speakers or signers.With tools such as these, researchers can create corpora containing, for example, grammatical information, discourse structure, facial expression, and gesture.The issue, then, is not the ability to create corpora containing gesture and speech information, but the type of information captured when describing gestures.We present a non-semantic, geometrically-based annotation scheme, FORM, which allows an annotator to capture the kinematic information in a gesture just from videos of speakers.In addition, FORM stores this gestural information in Annotation Graph format-allowing for easy integration of gesture information with other types of communication information, e.g., discourse structure, parts of speech, intonation information, etc. Craig Martell |
INTERSPEECH | 1 |
| 2002 | FORM: An Extensible, Kinematically-based Gesture Annotation Scheme
Craig Martell |
LREC | 1 |