Jan Neerbek

dblp:51/569 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 first-authorTheory of computation · 2Artificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Privacy and data protection · 100%
Theoretical computer science
1 paper
Quantum computing and quantum information · 30% Algorithmic game theory and mechanism design · 23% Algorithms and data structures · 23%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection
sensitive information detection
0.312017
TABOO: Detecting Unstructured Sensitive Information Using Recursive Neural Networks · ICDE 2017
Natural language and speech › Information extraction and text analysis › text similarity
paraphrase identification
0.112017
TABOO: Detecting Unstructured Sensitive Information Using Recursive Neural Networks · ICDE 2017
Computational complexity › decision problems
element distinctness
0.012001
Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness · ICALP 2001
Algorithmic game theory and mechanism design › online decision making
ordered search
0.012001
Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness · ICALP 2001
Quantum computing and quantum information
quantum algorithms
0.012001
Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness · ICALP 2001
Algorithms and data structures › sequence algorithms
sorting
0.012001
Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness · ICALP 2001
Quantum computing and quantum information
quantum complexity theory
0.012001
Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness · ICALP 2001

Methods — techniques the papers use, named apart from their topics

recursive neural network · 0.6natural language processing · 0.6
YearPublicationVenuePosition
2020 A Real-World Data Resource of Complex Sensitive Sentences Based on Documents from the Monsanto Trial
abstract
In this work we present a corpus for the evaluation of sensitive information detection approaches that addresses the need for real world sensitive information for empirical studies. Our sentence corpus contains different notions of complex sensitive information that correspond to different aspects of concern in a current trial of the Monsanto company. This paper describes the annotations process, where we both employ human annotators and furthermore create automatically inferred labels regarding technical, legal and informal communication within and with employees of Monsanto, drawing on a classification of documents by lawyers involved in the Monsanto court case. We release corpus of high quality sentences and parse trees with these two types of labels on sentence level. We characterize the sensitive information via several representative sensitive information detection models, in particular both keyword-based (n-gram) approaches and recent deep learning models, namely, recurrent neural networks (LSTM) and recursive neural networks (RecNN). Data and code are made publicly available.
Jan Neerbek, Morten Eskildsen, Peter Dolog, Ira Assent
LREC1
2019 Selective Training: A Strategy for Fast Backpropagation on Sentence Embeddings
Jan Neerbek, Peter Dolog, Ira Assent
PAKDD (3)1
2018 Detecting Complex Sensitive Information via Phrase Structure in Recursive Neural Networks
Jan Neerbek, Ira Assent, Peter Dolog
PAKDD (3)1
2017 TABOO: Detecting Unstructured Sensitive Information Using Recursive Neural Networks
abstract
Leak of sensitive information from unstructured text documents is a costly problem both for government and for industrial institutions. Traditional approaches for data leak prevention are commonly based on the hypothesis that sensitive information is reflected in the presence of distinct sensitive words. However, for complex sensitive information, this hypothesis may not hold. Our TABOO system detects complex sensitive information in text documents by learning the semantic and syntactic structure of text documents. Our approach is based on natural language processing methods for paraphrase detection, and uses recursive neural networks to assign sensitivity scores to the semantic components of the sentence structure. The demonstration of TABOO focuses on interactive detection of sensitive information with the TABOO system. Users may work with real documents, alter documents or prepare free text, and subject it to information detection. TABOO allows users to work with our TABOO engine or with traditional approaches, and to compare results. Users may verify that single words can change sensitivity according to context, thereby giving hands-on experience with complex cases of sensitive information.
Jan Neerbek, Ira Assent, Peter Dolog
ICDE1
2002 Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness
Peter Høyer, Jan Neerbek, Yaoyun Shi
Algorithmica2
2001 Quantum Complexities of Ordered Searching, Sorting, and Element Distinctness
Peter Høyer, Jan Neerbek, Yaoyun Shi
ICALP2