VLDB 2026 Research / reviewers in the wild / expert
Calvin Huang
dblp:180/2943
· DBLP profile ↗
2ranked-venue papers
0as first author
1since 2021 · last 2026
0000-0002-5054-4714ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 100% | |
| Artificial intelligence
1 paper |
Speech recognition and synthesis · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Wearable and physiological sensing
brain-computer interface |
1.0 | 1 | 2026 | CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding · AAAI 2026 |
Wearable and physiological sensing
electromyography |
1.0 | 1 | 2026 | CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding · AAAI 2026 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
tone recognition |
0.3 | 1 | 2026 | CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
multimodal fusion · 2.0domain adversarial training · 2.0cross-attention fusion · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone DecodingabstractBrain-computer interface (BCI) speech decoding has emerged as a promising tool for assisting individuals with speech impairments. In this context, the integration of electroencephalography (EEG) and electromyography (EMG) signals offers strong potential for enhancing decoding performance. Mandarin tone classification presents particular challenges, as tonal variations convey distinct meanings even when phonemes remain identical. In this study, we propose a novel cross-subject multimodal BCI decoding framework that fuses EEG and EMG signals to classify four Mandarin tones under both audible and silent speech conditions. Inspired by the cooperative mechanisms of neural and muscular systems in speech production, our neural decoding architecture combines spatial-temporal feature extraction branches with a cross-attention fusion mechanism, enabling informative interaction between modalities. We further incorporate domain-adversarial training to improve cross-subject generalization. We collected 4,800 EEG trials and 4,800 EMG trials from 10 participants using only twenty EEG and five EMG channels, demonstrating the feasibility of minimal-channel decoding. Despite employing lightweight modules, our model outperforms state-of-the-art baselines across all conditions, achieving average classification accuracies of 87.83\% for audible speech and 88.08\% for silent speech. In cross-subject evaluations, it still maintains strong performance with accuracies of 83.27\% and 85.10\% for audible and silent speech, respectively. We further conduct ablation studies to validate the effectiveness of each component. Our findings suggest that tone-level decoding with minimal EEG-EMG channels is feasible and potentially generalizable across subjects, contributing to the development of practical BCI applications. Yifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou 0001, Jiawei Ju |
AAAI | 2 |
| 2016 | Distributional semantics for understanding spoken meal descriptionsabstractThis paper presents ongoing language understanding experiments conducted as part of a larger effort to create a nutrition dialogue system that automatically extracts food concepts from a user's spoken meal description. We first discuss the technical approaches to understanding, including three methods for incorporating word vector features into conditional random field (CRF) models for semantic tagging, as well as classifiers for directly associating foods with properties. We report experiments on both text and spoken data from an in-domain speech recognizer. On text data, we show that the addition of word vector features significantly improves performance, achieving an F1 test score of 90.8 for semantic tagging and 86.3 for food-property association. On speech, the best model achieves an F1 test score of 87.5 for semantic tagging and 86.0 for association. Finally, we conduct an end-to-end system evaluation through a user study with human ratings of 83% semantic tagging accuracy. Mandy Korpusik, Calvin Huang, Michael Price 0001, James R. Glass |
ICASSP | 2 |