Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Aaron Soh

dblp:429/6763 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 67% Language models and text generation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
1.012026
Native Speech Processing with LLMs · AAAI 2026
Natural language and speech › Language models and text generation
large language model
1.012026
Native Speech Processing with LLMs · AAAI 2026
Natural language and speech › Speech recognition and synthesis
speaker diarization
1.012026
Native Speech Processing with LLMs · AAAI 2026

Methods — techniques the papers use, named apart from their topics

large language model · 1.0context-based prompting · 1.0
YearPublicationVenuePosition
2026 Native Speech Processing with LLMs
abstract
Recent advances in Large Language Models (LLMs) have achieved state-of-the-art performance in Automatic Speech Recognition (ASR), surpassing ASR-only systems such as Whisper. However, their application to other speech processing tasks, particularly speaker diarisation (SD), remains underexplored. This work proposes extending existing speech-aware LLM architectures with diarisation-specific training and context-based prompting to enable joint transcription and segmentation of multi-speaker audio. By exploiting the semantic reasoning and multilingual capabilities of pretrained LLMs, the proposed approach aims to improve diarisation accuracy, enhancing accessibility for assistive technologies and real-time captioning applications that rely on accurate speaker-aware transcriptions.
Aaron Soh
AAAI1