Ke Shen 0003

dblp:22/6778-3 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3600-0435ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 39% Knowledge representation and reasoning · 30% Trustworthy machine learning · 30%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.812024
The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning · AAAI 2024
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model
0.812024
The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning · AAAI 2024
Data integration and cleaning
entity resolution
0.812024
Multipartite Entity Resolution: Motivating a K-Tuple Perspective (Student Abstract) · AAAI 2024
Data integration and cleaning › entity resolution
entity resolution evaluation
0.812024
Multipartite Entity Resolution: Motivating a K-Tuple Perspective (Student Abstract) · AAAI 2024
Natural language and speech › Language models and text generation › natural language understanding › question answering
multiple-choice question answering
0.212024
The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning · AAAI 2024

Methods — techniques the papers use, named apart from their topics

risk-centric evaluation framework · 0.8precision and recall metrics · 0.8confidence distribution analysis · 0.8
YearPublicationVenuePosition
2025 Defining and evaluating decision and composite risk in language models applied to natural language inference
Ke Shen 0003, Mayank Kejriwal
Eng. Appl. Artif. Intell.1
2024 The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning
abstract
The advent of powerful transformer-based discriminative language models and, more recently, generative GPT-family models, has led to notable advancements in natural language processing (NLP), particularly in commonsense reasoning tasks. One such task is commonsense reasoning, where performance is usually evaluated through multiple-choice question-answering benchmarks. Till date, many such benchmarks have been proposed and `leaderboards' tracking state-of-the-art performance on those benchmarks suggest that transformer-based models are approaching human-like performance. However, due to documented problems such as hallucination and bias, the research focus is shifting from merely quantifying accuracy on the task to an in-depth, context-sensitive probing of LLMs' generalization and robustness. To gain deeper insight into diagnosing these models' performance in commonsense reasoning scenarios, this thesis addresses three main studies: the generalization ability of transformer-based language models on commonsense reasoning, the trend in confidence distribution of these language models confronted with ambiguous inference tasks, and a proposed risk-centric evaluation framework for both discriminative and generative language models.
Ke Shen 0003
AAAI1
2024 Multipartite Entity Resolution: Motivating a K-Tuple Perspective (Student Abstract)
abstract
Entity Resolution (ER) is the problem of algorithmically matching records, mentions, or entries that refer to the same underlying real-world entity. Traditionally, the problem assumes (at most) two datasets, between which records need to be matched. There is considerably less research in ER when k > 2 datasets are involved. The evaluation of such multipartite ER (M-ER) is especially complex, since the usual ER metrics assume (whether implicitly or explicitly) k < 3. This paper takes the first step towards motivating a k-tuple approach for evaluating M-ER. Using standard algorithms and k-tuple versions of metrics like precision and recall, our preliminary results suggest a significant difference compared to aggregated pairwise evaluation, which would first decompose the M-ER problem into independent bipartite problems and then aggregate their metrics. Hence, M-ER may be more challenging and warrant more novel approaches than current decomposition-based pairwise approaches would suggest.
Adin Aberbach, Mayank Kejriwal, Ke Shen 0003
AAAI3
2023 An experimental study measuring the generalization of fine-tuned language representation models across commonsense reasoning benchmarks
abstract
Abstract In the last 5 years, language representation models, such as BERT and GPT‐3, based on transformer neural networks, have led to enormous progress in natural language processing (NLP). One such NLP task is commonsense reasoning, where performance is usually evaluated through multiple‐choice question answering benchmarks. Till date, many such benchmarks have been proposed, and ‘leaderboards’ tracking state‐of‐the‐art performance on those benchmarks suggest that transformer‐based models are approaching human‐like performance. Because these are commonsense benchmarks, however, such a model should be expected to generalize, that is, at least in aggregate, should not exhibit excessive performance loss across independent commonsense benchmarks regardless of the specific benchmark on (the training set of) which it has been fine‐tuned. In this article, we evaluate this expectation by proposing a methodology and experimental study to measure the generalization ability of language representation models using a rigorous and intuitive metric. Using five established commonsense reasoning benchmarks, our experimental study shows that the models do not generalize well, and may be (potentially) susceptible to issues such as dataset bias. The results therefore suggest that current performance on benchmarks may be an over‐estimate, especially if we want to use such models on novel commonsense problems for which a ‘training’ dataset may not be available, for the language representation model, to fine‐tune on.
Ke Shen 0003, Mayank Kejriwal
Expert Syst. J. Knowl. Eng.1
2022 Transfer-based taxonomy induction over concept labels
Mayank Kejriwal, Ke Shen 0003, Chien-Chun Ni, Nicolas Torzec
Eng. Appl. Artif. Intell.2
2021 Unsupervised real-time induction and interactive visualization of taxonomies over domain-specific concepts
abstract
Given a domain-specific set of concept labels, taxonomy induction is the problem of inducing a taxonomy over the concept labels. Despite its importance in problems such as e-commerce, and some algorithmic research as a consequence, practical tools for taxonomy induction and interactive visualization do not currently exist. To be truly useful, such a tool must permit a reasonable solution in a relatively unsupervised setting, and be applicable to general subsets of concept labels. In this paper, we present an unsupervised, end-to-end taxonomy induction system for arbitrary concept-labels from the e-commerce domain. Our system only takes a simple text file as input and yields a tree-like taxonomy that can be rendered on a browser, and that a non-technical user can interact with. Important components of the system can also be customized by a technically experienced user.
Mayank Kejriwal, Ke Shen 0003
ASONAM2
2017 Relation Linking for Wikidata Using Bag of Distribution Representation
Shiya Ren, Yuan Li 0050, Ke Shen 0003, Guoyin Wang 0001
NLPCC4