Siawpeng Er

dblp:258/5023 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 33% Deep learning architectures and training · 29% Learning theory · 15%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
convolutional residual networks
0.612022
Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint · ICML 2022
Machine learning › Reinforcement learning
function approximation
0.612022
Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint · ICML 2022
Machine learning › Learning theory › neural network theory
neural network approximation theory
0.612022
Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint · ICML 2022
Machine learning › Deep learning architectures and training
overparameterized neural network
0.612022
Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint · ICML 2022
Natural language and speech › Information extraction and text analysis › named entity recognition
distantly supervised NER
0.412020
BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision · KDD 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision · KDD 2020
Natural language and speech › Information extraction and text analysis › named entity recognition
open-domain named entity recognition
0.412020
BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision · KDD 2020
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.212022
Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint · ICML 2022
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning
0.112020
BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision · KDD 2020

Methods — techniques the papers use, named apart from their topics

self-training · 0.4pre-trained language model · 0.4distant supervision · 0.4
YearPublicationVenuePosition
2025 Taxonomy-Based Negative Sampling in Personalized Semantic Search for E-Commerce
Uthman Jinadu, Siawpeng Er, Chen Liang 0006, Aleksandar Velkoski
IEEE Big Data2
2022 Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint
abstract
Overparameterized neural networks enjoy great representation power on complex data, and more importantly yield sufficiently smooth output, which is crucial to their generalization and robustness. Most existing function approximation theories suggest that with sufficiently many parameters, neural networks can well approximate certain classes of functions in terms of the function value. The neural network themselves, however, can be highly nonsmooth. To bridge this gap, we take convolutional residual networks (ConvResNets) as an example, and prove that large ConvResNets can not only approximate a target function in terms of function value, but also exhibit sufficient first-order smoothness. Moreover, we extend our theory to approximating functions supported on a low-dimensional manifold. Our theory partially justifies the benefits of using deep and wide networks in practice. Numerical experiments on adversarial robust image classification are provided to support our theory.
Hao Liu 0028, Minshuo Chen, Siawpeng Er, Wenjing Liao, Tong Zhang 0001, Tuo Zhao
ICML3
2020 BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision
abstract
We study the open-domain named entity recognition (NER) problem under distant supervision. The distant supervision, though does not require large amounts of manual annotations, yields highly incomplete and noisy distant labels via external knowledge bases. To address this challenge, we propose a new computational framework -- BOND, which leverages the power of pre-trained language models (e.g., BERT and RoBERTa) to improve the prediction performance of NER models. Specifically, we propose a two-stage training algorithm: In the first stage, we adapt the pre-trained language model to the NER tasks using the distant labels, which can significantly improve the recall and precision; In the second stage, we drop the distant labels, and propose a self-training approach to further improve the model performance. Thorough experiments on 5 benchmark datasets demonstrate the superiority of BOND over existing distantly supervised NER methods. The code and distantly labeled data have been released in https://github.com/cliang1453/BOND.
Chen Liang 0006, Yue Yu 0001, Haoming Jiang, Siawpeng Er, Tuo Zhao, Chao Zhang 0014
KDD4
2019 SEACOIN2.0: an interactive mining and visualization tool for information retrieval, summarization and knowledge discovery
abstract
The rapidly increasing size of biomedical databases such as Medline requires the use of intelligent data mining methods for information extraction and summarization. Existing biomedical text-mining tools have limited capabilities for incorporating citation information during document ranking and for inferring topological and network relationships between biomedical terms. Often too much is returned during summarization leading to information overload. Furthermore, literature-based discoveries could be hard to interpret if the network is too complex. SEACOIN2.0 can incorporate citation information during document ranking and uses a unique association rule mining algorithm to generate multi-level k-ary trees. The multi-level trees facilitate efficient information retrieval, visual data exploration, summarization, and hypothesis generation. The system presents graphical summarization via multiple dynamic visualization panels and an interactive word cloud. LexRank algorithm is used to identify salient sentences in top abstracts related to the query. An average F-measure of 94% was achieved for document retrieval, and an average precision of 88% was obtained for identification of top co-occurrence terms. SEACOIN2.0 was also used to replicate previously published findings using the literature-based discovery and EMR-based PheWAS approaches. We present herein SEACOIN2.0 (https://newton.isye.gatech.edu/SEACOIN2/), an interactive visual mining tool for improved information retrieval, automated multi-level summarization of Medline abstracts, and literature-based discovery. SEACOIN2.0 addresses the problem of “information overload” and allows clinicians and biomedical researchers to meet their information needs.
Eva K. Lee, Karan Uppal, Siawpeng Er
BIBM3