Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Like Hui

dblp:173/3760 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 51% Trustworthy machine learning · 26% Learning theory · 23%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
feature separation
0.912025
Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.912025
Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025
Machine learning › Deep learning architectures and training › activation function
ReLU
0.912025
Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks
0.912025
Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025
Machine learning › Deep learning architectures and training › loss function design
classification loss function
0.712023
Cut your Losses with Squentropy · ICML 2023
Machine learning › Deep learning architectures and training
loss function design
0.712023
Cut your Losses with Squentropy · ICML 2023
Machine learning › Trustworthy machine learning › calibration
model calibration
0.712023
Cut your Losses with Squentropy · ICML 2023
Machine learning › Learning theory
loss function
0.512021
Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification Tasks · ICLR 2021

Methods — techniques the papers use, named apart from their topics

square loss · 1.2cross-entropy loss · 1.2gradient descent convergence analysis · 0.9
YearPublicationVenuePosition
2025 Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks
abstract
Nonlinear activation functions are widely recognized for enhancing the expressivity of neural networks, which is the primary reason for their widespread implementation. In this work, we focus on ReLU activation and reveal a novel and intriguing property of nonlinear activations. By comparing enabling and disabling the nonlinear activations in the neural network, we demonstrate their specific effects on wide neural networks: (a) *better feature separation*, i.e., a larger angle separation for similar data in the feature space of model gradient, and (b) *better NTK conditioning*, i.e., a smaller condition number of neural tangent kernel (NTK). Furthermore, we show that the network depth (i.e., with more nonlinear activation operations) further amplifies these effects; in addition, in the infinite-width-then-depth limit, all data are equally separated with a fixed angle in the model gradient feature space, regardless of how similar they are originally in the input space. Note that, without the nonlinear activation, i.e., in a linear neural network, the data separation remains the same as for the original inputs and NTK condition number is equivalent to the Gram matrix, regardless of the network depth. Due to the close connection between NTK condition number and convergence theories, our results imply that nonlinear activation helps to improve the worst-case convergence rates of gradient based methods.
Han Bi, Like Hui
NeurIPS3
2023 Cut your Losses with Squentropy
abstract
Nearly all practical neural models for classification are trained using the cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Recent work (Hui & Belkin, 2020) suggests that training using the (rescaled) square loss is often superior in terms of the classification accuracy. In this paper we propose the "squentropy" loss, which is the sum of two terms: the cross-entropy loss and the average square loss over the incorrect classes. We provide an extensive set of experiment on multi-class classification problems showing that the squentropy loss outperforms both the pure cross-entropy and rescaled square losses in terms of the classification accuracy. We also demonstrate that it provides significantly better model calibration than either of these alternative losses and, furthermore, has less variance with respect to the random initialization. Additionally, in contrast to the square loss, squentropy loss can frequently be trained using exactly the same optimization parameters, including the learning rate, as the standard cross-entropy loss, making it a true ''plug-and-play'' replacement. Finally, unlike the rescaled square loss, multiclass squentropy contains no parameters that need to be adjusted.
Like Hui, Mikhail Belkin, Stephen Wright
ICML1
2021 Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification Tasks
Like Hui, Mikhail Belkin
ICLR1
2019 Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr
abstract
In this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covariance matrix for the beamformer. Minimum Variance Distortionless Response (MVDR) beamformer and Generalized Eigenvalue (GEV) beamformer are both investigated under the CRM-based joint training architecture. We also propose a robust mask pooling strategy among multiple channels. A long short-term memory (LSTM) based language model is utilized to re-score hypotheses which further improves the overall performance. We evaluate the proposed methods on CHiME-4 challenge dataset. The CRM based system achieves a relative 10% reduction on word error rate (WER) compared with the IRM based system. Without sequence discriminative training, our best single system already achieves an average WER 2.72% on the test set which is comparable to the state-of-the-art.
Yong Xu 0004, Chao Weng, Like Hui, Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP3
2019 Kernel Machines Beat Deep Neural Networks on Mask-Based Single-Channel Speech Enhancement
abstract
We apply a fast kernel method for mask-based single-channel speech enhancement.Specifically, our method solves a kernel regression problem associated to a non-smooth kernel function (exponential power kernel) with a highly efficient iterative method (EigenPro).Due to the simplicity of this method, its hyper-parameters such as kernel bandwidth can be automatically and efficiently selected using line search with subsamples of training data.We observe an empirical correlation between the regression loss (mean square error) and regular metrics for speech enhancement.This observation justifies our training target and motivates us to achieve lower regression loss by training separate kernel model per frequency subband.We compare our method with the state-of-the-art deep neural networks on mask-based HINT and TIMIT.Experimental results show that our kernel method consistently outperforms deep neural networks while requiring less training time.
Like Hui, Mikhail Belkin
INTERSPEECH1
2015 High-performance Swahili keyword search with very limited language pack: The THUEE system for the OpenKWS15 evaluation
abstract
This paper presents the Swahili keyword search system developed by the THUEE team for the OpenKWS15 evaluation, which is conducted by NIST under the IARPA Babel program. There are several highlights in the development of the system, including automatic generation of the pronunciation lexicon, aggressive data augmentation, the multilingual bottleneck feature extractor trained from 6 languages, text selection from web data for language model training, semi-supervised training for acoustic models and language models, out-of-vocabulary keyword detection using morphemes and a rich diversity of the systems for combination. A wide variety of acoustic modeling techniques are explored and compared. Up to 12 different individual systems are used for combination. The system achieves the state-of-the-art performance in the required condition of the evaluation.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Jia Liu 0001
ASRU5
2015 Improved system fusion for keyword search
abstract
It has been demonstrated that system fusion can significantly improve the performance of keyword search. In this paper, we compare the performance of several widely-used arithmetic-based fusion methods using different normalization pipeline and try to find the best pipeline. A novel arithmetic-based fusion method is proposed in this work. The method supplies a more effective way to incorporate the number of systems which have non-zero scores for a detection. When tested on the development test dataset of the OpenKWS15 Evaluation, the proposed method achieves the highest maximum term-weighted value (MTWV) and actual term-weighted value (ATWV) among all other arithmetic-based fusion methods. Usually, discriminative fusion methods employing classifiers can outperform arithmetic-based fusion methods. A DNN-based fusion method is explored in this work. After word-burst information is added, the DNN-based fusion method outperforms all other methods. In addition, it is notable that our arithmetic-based method achieves the same MTWV as the DNN-based method.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Weiqiang Zhang 0001, Jia Liu 0001
ASRU5