Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Han Shu

dblp:30/4042 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Efficient and distributed learning · 43% Language models and text generation · 12% Segmentation and scene understanding · 12%
Computer networks
1 paper
Network optimization and economics · 100%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.842025
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking · ICML 2024
Frequency Domain Compact 3D Convolutional Neural Networks · CVPR 2020
Co-Evolutionary Compression for Unpaired Image Translation · ICCV 2019
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.222026
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding · NeurIPS 2025
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference · INFOCOM 2026
Network optimization and economics › resource allocation
fair resource allocation
1.012026
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference · INFOCOM 2026
Network optimization and economics › throughput maximization
goodput maximization
1.012026
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference · INFOCOM 2026
Computer vision › Segmentation and scene understanding › image segmentation
efficient segmentation
0.912025
TinySAM: Pushing the Envelope for Efficient Segment Anything Model · AAAI 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding · NeurIPS 2025
Computer vision › Segmentation and scene understanding › prompt-based segmentation
segment anything model
0.912025
TinySAM: Pushing the Envelope for Efficient Segment Anything Model · AAAI 2025
Computer vision › Vision and language › vision-language model
vision-language model inference
0.912025
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.832025
Distilling Portable Generative Adversarial Networks for Image Translation · AAAI 2020
TinySAM: Pushing the Envelope for Efficient Segment Anything Model · AAAI 2025
Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer · ECCV (6) 2020
Natural language and speech › Language models and text generation
large language model
0.812024
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking · ICML 2024
Natural language and speech › Language models and text generation
large language model training
0.812024
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking · ICML 2024
Machine learning › Efficient and distributed learning › efficient neural network design
adder neural network
0.512021
Adder Attention for Vision Transformer · NeurIPS 2021
Machine learning › Deep learning architectures and training
attention mechanism
0.512021
Adder Attention for Vision Transformer · NeurIPS 2021
Machine learning › Efficient and distributed learning › inference efficiency
energy-efficient inference
0.512021
Adder Attention for Vision Transformer · NeurIPS 2021
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.512021
Adder Attention for Vision Transformer · NeurIPS 2021
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
3d convolutional neural network
0.412020
Frequency Domain Compact 3D Convolutional Neural Networks · CVPR 2020
Machine learning › Generative modeling
generative adversarial network
0.412020
Distilling Portable Generative Adversarial Networks for Image Translation · AAAI 2020
Computer vision › 3D vision › motion estimation
optical flow
0.412020
Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer · ECCV (6) 2020
Visual content generation and editing › style transfer
video style transfer
0.412020
Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer · ECCV (6) 2020
Machine learning › Efficient and distributed learning › model compression › generative model compression
GAN compression
0.412019
Co-Evolutionary Compression for Unpaired Image Translation · ICCV 2019
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.412019
Co-Evolutionary Compression for Unpaired Image Translation · ICCV 2019
Computer vision › Face, body and person analysis
pedestrian attribute recognition
0.412019
Attribute Aware Pooling for Pedestrian Attribute Recognition · IJCAI 2019
Machine learning › Generative modeling › generative adversarial network › image-to-image translation
unsupervised image-to-image translation
0.412019
Co-Evolutionary Compression for Unpaired Image Translation · ICCV 2019
Natural language and speech › Language models and text generation
large language model inference
0.312026
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference · INFOCOM 2026
Computer vision › Vision and language
vision-language model
0.312025
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.212024
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking · ICML 2024
Visual content generation and editing
image-to-image translation
0.112020
Distilling Portable Generative Adversarial Networks for Image Translation · AAAI 2020

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 2.6utility maximization · 2.0gradient scheduling · 2.0fluid sample path analysis · 2.0vision adaptor · 0.9speculative decoding · 0.9post-training quantization · 0.9weight-momentum joint shrinking · 0.8residual computation · 0.8non-uniform quantization · 0.8identity mapping · 0.5adder operation · 0.5optical flow · 0.4adversarial learning · 0.4
YearPublicationVenuePosition
2026 GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
abstract
Large language models (LLMs) have revolutionized natural language processing, yet their high computational demands pose significant challenges for real-time inference, especially in multi-user server speculative decoding and resource-constrained environments. Speculative decoding has emerged as a promising technique to accelerate LLM inference by using lightweight draft models to generate candidate tokens, which are subsequently verified by a larger, more accurate model. However, ensuring both high goodput (the effective rate of accepted tokens) and fairness across multiple draft servers cooperating with a central verification server remains an open challenge. This paper introduces GOODSPEED, a novel distributed inference framework that optimizes goodput through adaptive speculative decoding. GOODSPEED employs a central verification server that coordinates a set of heterogeneous draft servers, each running a small language model to generate speculative tokens. To manage resource allocation effectively, GOODSPEED incorporates a gradient scheduling algorithm that dynamically assigns token verification tasks, maximizing a logarithmic utility function to ensure proportional fairness across servers. By processing speculative outputs from all draft servers in parallel, the framework enables efficient collaboration between the verification server and distributed draft generators, streamlining both latency and throughput. Through rigorous fluid sample path analysis, we show that GOODSPEED converges to the optimal goodput allocation in steady-state conditions and maintains near-optimal performance with provably bounded error under dynamic workloads. These results demonstrate that GOODSPEED provides a scalable, fair and efficient solution for multi-server speculative decoding in distributed LLM inference systems.
Phuong Tran, Tzu-Hao Liu, Tung-Anh Nguyen, Van Quan La, Eason Yu, Han Shu, Choong Seon Hong, Nguyen H. Tran
INFOCOM7
2026 Federated Koopman-Reservoir Learning for Multivariate Time-Series Anomaly Detection in IoT
abstract
The rapid expansion of the Internet of Things (IoT) has led to unprecedented growth in multivariate time-series (MVTS) data, which are vital for real-world applications such as industrial monitoring, cyber-physical security, and smart city operations. These data streams are susceptible to anomalies that may indicate system malfunctions, security breaches, or environmental hazards. However, existing MVTS anomaly detection (MTAD) approaches, typically trained in centralized settings, struggle in IoT deployments due to data heterogeneity, resource constraints, and privacy concerns. We propose FEDKO, a novel federated learning (FL) framework that couples Reservoir Computing with Koopman operator theory for efficient, privacy-preserving MTAD in distributed IoT networks. At its core, ReKO, a lightweight spatio-temporal Reservoir-Koopman model, lifts nonlinear MVTS dynamics into a linear space for stable prediction and reconstruction. We formulate the FL training as a bi-level optimization procedure where the inner level learns locally stable Koopman dynamics, and the outer level refines lifted feature representations and reconstruction mappings. We further provide theoretical convergence guarantees, anomaly discriminability analysis, and a structural privacy characterization of the framework. Experiments on four IoT MVTS datasets and deployment on an NVIDIA Jetson edge device show that FEDKO achieves a balanced precision–recall profile with competitive F1-scores under heterogeneous federated settings, while substantially reducing communication and memory footprints compared with MTAD baselines.
Nhat Huy Le, Han Shu, Zilong Jin, Nguyen Binh Truong, Nguyen Hoang Tran, Zhu Han 0001, Choong Seon Hong
IEEE Internet Things J.4
2025 TinySAM: Pushing the Envelope for Efficient Segment Anything Model
abstract
Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed various applications based on the pre-trained SAM and achieved impressive performance on downstream vision tasks. However, SAM consists of heavy architectures and requires massive computational capacity, which hinders the further application of SAM on computation constrained edge devices. To this end, in this paper we propose a framework to obtain a tiny segment anything model (TinySAM) while maintaining the strong zero-shot performance. We first propose a full-stage knowledge distillation method with hard prompt sampling and hard mask weighting strategy to distill a lightweight student model. We also adapt the post-training quantization to the prompt-based segmentation task and further reduce the computational cost. Moreover, a hierarchical segmenting everything strategy is proposed to accelerate the everything inference by 2× with almost no performance degradation. With all these proposed methods, our TinySAM leads to orders of magnitude computational reduction and pushes the envelope for efficient segment anything task. Extensive experiments on various zero-shot transfer tasks demonstrate the significantly advantageous performance of our TinySAM against counterpart methods.
Han Shu, Yehui Tang 0001, Houqiang Li, Yunhe Wang 0001, Xinghao Chen 0001
AAAI1
2025 ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
abstract
Speculative decoding is a widely adopted technique for accelerating inference in large language models (LLMs), yet its application to vision-language models (VLMs) remains underexplored, with existing methods achieving only modest speedups ($<1.5\times$). This gap is increasingly significant as multimodal capabilities become central to large-scale models. We hypothesize that large VLMs can effectively filter redundant image information layer by layer without compromising textual comprehension, whereas smaller draft models struggle to do so. To address this, we introduce Vision-Aware Speculative Decoding (ViSpec), a novel framework tailored for VLMs. ViSpec employs a lightweight vision adaptor module to compress image tokens into a compact representation, which is seamlessly integrated into the draft model's attention mechanism while preserving original image positional information. Additionally, we extract a global feature vector for each input image and augment all subsequent text tokens with this feature to enhance multimodal coherence. To overcome the scarcity of multimodal datasets with long assistant responses, we curate a specialized training dataset by repurposing existing datasets and generating extended outputs using the target VLM with modified prompts. Our training strategy mitigates the risk of the draft model exploiting direct access to the target model's hidden states, which could otherwise lead to shortcut learning when training solely on target model outputs. Extensive experiments validate ViSpec, achieving, to our knowledge, the first substantial speedup in VLM speculative decoding.
Jialiang Kang, Han Shu, Yingjie Zhai, Xinghao Chen 0001
NeurIPS2
2025 Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
abstract
The proliferation of edge devices has dramatically increased the generation of multivariate time-series (MVTS) data, essential for applications from healthcare to smart cities. Such data streams, however, are vulnerable to anomalies that signal crucial problems like system failures or security incidents. Traditional MVTS anomaly detection methods, encompassing statistical and centralized machine learning approaches, struggle with the heterogeneity, variability, and privacy concerns of large-scale, distributed environments. In response, we introduce FedKO, a novel unsupervised Federated Learning framework that leverages the linear predictive capabilities of Koopman operator theory along with the dynamic adaptability of Reservoir Computing. This enables effective spatiotemporal processing and privacy-preserving for MVTS data. FedKO is formulated as a bi-level optimization problem, utilizing a specific federated algorithm to explore a shared Reservoir-Koopman model across diverse datasets. Such a model is then deployable on edge devices for efficient detection of anomalies in local MVTS streams. Experimental results across various datasets showcase FedKO’s superior performance against state-of-the-art methods in MVTS anomaly detection. Moreover, FedKO reduces up to 8x communication size and 2x memory usage, making it highly suitable for large-scale systems.
Tung-Anh Nguyen, Han Shu, Suranga Seneviratne, Choong Seon Hong, Nguyen H. Tran
SDM3
2025 MAEST: accurately spatial domain detection in spatial transcriptomics with graph masked autoencoder
abstract
Spatial transcriptomics (ST) technology provides gene expression profiles with spatial context, offering critical insights into cellular interactions and tissue architecture. A core task in ST is spatial domain identification, which involves detecting coherent regions with similar spatial expression patterns. However, existing methods often fail to fully exploit spatial information, leading to limited representational capacity and suboptimal clustering accuracy. Here, we introduce MAEST, a novel graph neural network model designed to address these limitations in ST data. MAEST leverages graph masked autoencoders to denoise and refine representations while incorporating graph contrastive learning to prevent feature collapse and enhance model robustness. By integrating one-hop and multi-hop representations, MAEST effectively captures both local and global spatial relationships, improving clustering precision. Extensive experiments across diverse datasets, including the human brain, mouse hippocampus, olfactory bulb, brain, and embryo, demonstrate that MAEST outperforms seven state-of-the-art methods in spatial domain identification. Furthermore, MAEST showcases its ability to integrate multi-slice data, identifying joint domains across horizontal tissue sections with high accuracy. These results highlight MAEST's versatility and effectiveness in unraveling the spatial organization of complex tissues. The source code of MAEST can be obtained at https://github.com/clearlove2333/MAEST.
Han Shu, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Zhen Tian 0004, Tao Wang 0082
Briefings Bioinform.2
2024 ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
abstract
Large language models (LLM) have recently attracted significant attention in the field of artificial intelligence. However, the training process of these models poses significant challenges in terms of computational and storage capacities, thus compressing checkpoints has become an urgent problem. In this paper, we propose a novel Extreme Checkpoint Compression (ExCP) framework, which significantly reduces the required storage of training checkpoints while achieving nearly lossless performance. We first calculate the residuals of adjacent checkpoints to obtain the essential but sparse information for higher compression ratio. To further excavate the redundancy parameters in checkpoints, we then propose a weight-momentum joint shrinking method to utilize another important information during the model optimization, i.e., momentum. In particular, we exploit the information of both model and optimizer to discard as many parameters as possible while preserving critical information to ensure optimal performance. Furthermore, we utilize non-uniform quantization to further compress the storage of checkpoints. We extensively evaluate our proposed ExCP framework on several models ranging from 410M to 7B parameters and demonstrate significant storage reduction while maintaining strong performance. For instance, we achieve approximately $70\times$ compression for the Pythia-410M model, with the final performance being as accurate as the original model on various downstream tasks. Codes will be available at https://github.com/Gaffey/ExCP.
Xinghao Chen 0001, Han Shu, Yehui Tang 0001, Yunhe Wang 0001
ICML3
2024 Accurately deciphering spatial domains for spatially resolved transcriptomics with stCluster
abstract
Spatial transcriptomics provides valuable insights into gene expression within the native tissue context, effectively merging molecular data with spatial information to uncover intricate cellular relationships and tissue organizations. In this context, deciphering cellular spatial domains becomes essential for revealing complex cellular dynamics and tissue structures. However, current methods encounter challenges in seamlessly integrating gene expression data with spatial information, resulting in less informative representations of spots and suboptimal accuracy in spatial domain identification. We introduce stCluster, a novel method that integrates graph contrastive learning with multi-task learning to refine informative representations for spatial transcriptomic data, consequently improving spatial domain identification. stCluster first leverages graph contrastive learning technology to obtain discriminative representations capable of recognizing spatially coherent patterns. Through jointly optimizing multiple tasks, stCluster further fine-tunes the representations to be able to capture complex relationships between gene expression and spatial organization. Benchmarked against six state-of-the-art methods, the experimental results reveal its proficiency in accurately identifying complex spatial domains across various datasets and platforms, spanning tissue, organ, and embryo levels. Moreover, stCluster can effectively denoise the spatial gene expression patterns and enhance the spatial trajectory inference. The source code of stCluster is freely available at https://github.com/hannshu/stCluster.
Tao Wang 0082, Han Shu, Jialu Hu, Yongtian Wang, Jin Chen 0004, Jiajie Peng, Xuequn Shang 0001
Briefings Bioinform.2
2021 Adder Attention for Vision Transformer
abstract
Transformer is a new kind of calculation paradigm for deep learning which has shown strong performance on a large variety of computer vision tasks. However, compared with conventional deep models (e.g., convolutional neural networks), vision transformers require more computational resources which cannot be easily deployed on mobile devices. To this end, we present to reduce the energy consumptions using adder neural network (AdderNet). We first theoretically analyze the mechanism of self-attention and the difficulty for applying adder operation into this module. Specifically, the feature diversity, i.e., the rank of attention map using only additions cannot be well preserved. Thus, we develop an adder attention layer that includes an additional identity mapping. With the new operation, vision transformers constructed using additions can also provide powerful feature representations. Experimental results on several benchmarks demonstrate that the proposed approach can achieve highly competitive performance to that of the baselines while achieving an about 2~3× reduction on the energy consumption.
Han Shu, Jiahao Wang 0005, Hanting Chen, Yujiu Yang 0001, Yunhe Wang 0001
NeurIPS1
2020 Distilling Portable Generative Adversarial Networks for Image Translation
abstract
Despite Generative Adversarial Networks (GANs) have been widely used in various image-to-image translation tasks, they can be hardly applied on mobile devices due to their heavy computation and storage cost. Traditional network compression methods focus on visually recognition tasks, but never deal with generation tasks. Inspired by knowledge distillation, a student generator of fewer parameters is trained by inheriting the low-level and high-level information from the original heavy teacher generator. To promote the capability of student generator, we include a student discriminator to measure the distances between real images, and images generated by student and teacher generators. An adversarial learning process is therefore established to optimize student generator and student discriminator. Qualitative and quantitative analysis by conducting experiments on benchmark datasets demonstrate that the proposed method can learn portable generative models with strong performance.
Hanting Chen, Yunhe Wang 0001, Han Shu, Changyuan Wen, Chunjing Xu, Boxin Shi, Chao Xu 0006, Chang Xu 0002
AAAI3
2020 Frequency Domain Compact 3D Convolutional Neural Networks
abstract
This paper studies the compression and acceleration of 3-dimensional convolutional neural networks (3D CNNs). To reduce the memory cost and computational complexity of deep neural networks, a number of algorithms have been explored by discovering redundant parameters in pre-trained networks. However, most of existing methods are designed for processing neural networks consisting of 2-dimensional convolution filters (i.e. image classification and detection) and cannot be straightforwardly applied for 3-dimensional filters (i.e. time series data). In this paper, we develop a novel approach for eliminating redundancy in the time dimensionality of 3D convolution filters by converting them into the frequency domain through a series of learned optimal transforms with extremely fewer parameters. Moreover, these transforms are forced to be orthogonal, and the calculation of feature maps can be accomplished in the frequency domain to achieve considerable speed-up rates. Experimental results on benchmark 3D CNN models and datasets demonstrate that the proposed Frequency Domain Compact 3D CNNs (FDC3D) can achieve the state-of-the-art performance, \eg a 2x speed-up ratio on the 3D-ResNet-18 without obviously affecting its accuracy.
Hanting Chen, Yunhe Wang 0001, Han Shu, Yehui Tang 0001, Chunjing Xu, Boxin Shi, Chao Xu 0006, Qi Tian 0001, Chang Xu 0002
CVPR3
2020 Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer
Xinghao Chen 0001, Yunhe Wang 0001, Han Shu, Chunjing Xu, Chang Xu 0002
ECCV (6)4
2019 Co-Evolutionary Compression for Unpaired Image Translation
abstract
Generative adversarial networks (GANs) have been successfully used for considerable computer vision tasks, especially the image-to-image translation. However, generators in these networks are of complicated architectures with large number of parameters and huge computational complexities. Existing methods are mainly designed for compressing and speeding-up deep neural networks in the classification task, and cannot be directly applied on GANs for image translation, due to their different objectives and training procedures. To this end, we develop a novel co-evolutionary approach for reducing their memory usage and FLOPs simultaneously. In practice, generators for two image domains are encoded as two populations and synergistically optimized for investigating the most important convolution filters iteratively. Fitness of each individual is calculated using the number of parameters, a discriminator-aware regularization, and the cycle consistency. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed method for obtaining compact and effective generators.
Han Shu, Yunhe Wang 0001, Xu Jia 0012, Kai Han 0002, Hanting Chen, Chunjing Xu, Qi Tian 0001, Chang Xu 0002
ICCV1
2019 Attribute Aware Pooling for Pedestrian Attribute Recognition
abstract
This paper expands the strength of deep convolutional neural networks (CNNs) to the pedestrian attribute recognition problem by devising a novel attribute aware pooling algorithm. Existing vanilla CNNs cannot be straightforwardly applied to handle multi-attribute data because of the larger label space as well as the attribute entanglement and correlations. We tackle these challenges that hampers the development of CNNs for multi-attribute classification by fully exploiting the correlation between different attributes. The multi-branch architecture is adopted for fucusing on attributes at different regions. Besides the prediction based on each branch itself, context information of each branch are employed for decision as well. The attribute aware pooling is developed to integrate both kinds of information. Therefore, attributes which are indistinct or tangled with others can be accurately recognized by exploiting the context information. Experiments on benchmark datasets demonstrate that the proposed pooling method appropriately explores and exploits the correlations between attributes for the pedestrian attribute recognition.
Kai Han 0002, Yunhe Wang 0001, Han Shu, Chuanjian Liu, Chunjing Xu, Chang Xu 0002
IJCAI3
2017 A highly accurate facial region network for unconstrained face detection
abstract
In this paper, a new face detection method with very high accuracy is proposed. We introduce a novel facial region network to detect faces in unconstrained conditions. Firstly, a face proposal net is raised to generate possible face regions in the input image. Then, a novel weighted grid feature is applied to calculate features of face regions. Owing to that, faces with large pose variation and severe occlusion can be detected correctly. Furthermore, we use millions of general object data to pre-train the network to enhance the robustness of the extracted feature. Our method is evaluated on several public face detection datasets and achieves state-of-the-art performance on all of them. Specially, our method demonstrates a very high recall rate of 96.4% when false positives are 300 on the challenging FDDB benchmark, ranking first not only in the academic list but also in the commercial list which is much more competitive than the previous one.
Han Shu, Dangdang Chen, Yali Li 0001, Shengjin Wang
ICIP1
2006 Flexible Multi-Stream Framework for Speech Recognition using Multi-Tape Finite-State Transducers
abstract
We present an approach to general multi-stream recognition utilizing multi-tape finite-state transducers (FSTs). The approach is novel in that each of the multiple "streams'" of features can represent either a sequence (e.g., fixed- or variable-rate frames) or a directed acyclic graph (e.g., containing hypothesized phonetic segmentations). Each transition of the multi-tape FST specifies the models to be applied to each stream and the degree of feature stream asynchrony to be allowed. We show how this framework can easily represent the 2-stream variable-rate landmark and segment modeling utilised by our baseline SUMMIT speech recognizer. We present experiments merging standard hidden Markov models (HMMs) with landmark models on the Wall Street Journal speech recognition task, and find that some degree of asynchrony can be critical when combining different types of models. We also present experiments performing audio-visual speech recognition on the AV-TIMIT task
I. Lee Hetherington, Han Shu, James R. Glass
ICASSP (1)2
2005 A probabilistic approach to unit selection for corpus-based speech synthesis
abstract
In this paper, we present a novel statistical approach to corpus-based speech synthesis. Unit selection is directed by probabilistic models for F0 contour, duration, and spectral characteristics of the synthesis units. The F0 targets for units are modeled by statistical additive models, and duration targets are modeled by regression trees. Spectral targets for a unit is modeled by Gaussian mixtures on MFCC-based features. Goodness of concatenation of two units is modeled by conditional Gaussian models on MFCC-based features. Although the system is in its early stage of development, we implemented an English speech synthesizer with CMU Arctic corpora and confirmed the effectiveness of this new framework. 1.
Shinsuke Sakai, Han Shu
INTERSPEECH2
2005 Pronunciation modeling using a finite-state transducer representation
Timothy J. Hazen, I. Lee Hetherington, Han Shu, Karen Livescu
Speech Commun.3
2002 EM training of finite-state transducers and its application to pronunciation modeling
abstract
Recently, nite-state transducers (FSTs) have been shown to be useful for a number of applications in speech and language processing. FST operations such as composition, determinization, and minimization make manipulating FSTs very simple. In this paper, we present a method to learn weights for arbitrary FSTs using the EM algorithm. We show that this FST EM algorithm is able to learn pronunciation weights that improve the word error rate for a spontaneous speech recognition task.
Han Shu, I. Lee Hetherington
INTERSPEECH1
2000 The 2000 BBN Byblos LVCSR system
abstract
This paper describes the 2000 BBN Byblos Large Vocabulary Continuous Speech Recognition (LVCSR) system. We briefly outline the training and decoding procedures used in the system, and explain in detail the new features we have added to the system in the past year. These new features include multiple adaptation stages, parallel path rescoring, and a new word confidence system. Word error rate results for all of these additions are presented for Hub-5 English test sets containing both Switchboard II and CallHome speakers. 1. Introduction The 2000 BBN Byblos LVCSR system aims for state of the art performance on decoding spontaneous, conversational telephone speech. In particular, it is explicitly designed to achieve low word error rate (WER) on the NIST sponsored Hub-5 evaluations, the test sets of which consist of equal parts of Switchboard II and CallHome conversations [1]. In the March 2000 Hub-5 evaluation, Byblos achieved a WER of 29.1%. This paper has two parts. First, we describ...
Thomas Colthurst, Owen Kimball, Fred Richardson, Han Shu, Chuck Wooters, Rukmini Iyer, Herbert Gish
INTERSPEECH4
2000 The BBN Byblos 2000 conversational Mandarin LVCSR system
abstract
This paper describes the year 2000 BBN Byblos Mandarin large vocabulary conversational speech recognition (LVCSR) system, the winning (and only) Mandarin system from the Spring 2 000 Hub-5 evaluation sponsored by NIST. We first outline the training and d ecoding procedures used in the system, and describe the performance of the system used in the e valuation. We then d escribe the e ffect of several features that were not in the e valuation system but have been added since, including Jacobian compensated Vocal Tract Length Normalization (VTLN), system combination, a higher number of system parameters, and additional training data. Together these give a n additional 5.4% relative improvement on character error r ate (CER) from the evaluation system.
Han Shu, Chuck Wooters, Owen Kimball, Thomas Colthurst, Fred Richardson, Spyridon Matsoukas, Herbert Gish
INTERSPEECH1
1995 Duration modeling in large vocabulary speech recognition
abstract
This paper presents a study of different methods for phoneme duration modeling in large vocabulary speech recognition. We investigate the employment of phoneme duration and the effect of context, speaking rate and lexical stress in the duration of phoneme segments in a large vocabulary speech recognition system. The duration models are used in a postprocessing phase of BYBLOS, our baseline HMM-based recognition system, to rescore the N-Best hypotheses. We describe experiments with the 5 K word ARPA Wall Street Journal (WSJ) corpus. The results show that integration of duration models that take into account context and speaking rate can improve the word accuracy of the baseline recognition system.
Tasos Anastasakos, Richard M. Schwartz, Han Shu
ICASSP3