VLDB 2026 Research / reviewers in the wild / expert
Like Hui
dblp:173/3760
· DBLP profile ↗
7ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 51% Trustworthy machine learning · 26% Learning theory · 23% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
feature separation |
0.9 | 1 | 2025 | Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.9 | 1 | 2025 | Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › activation function
ReLU |
0.9 | 1 | 2025 | Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks |
0.9 | 1 | 2025 | Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › loss function design
classification loss function |
0.7 | 1 | 2023 | Cut your Losses with Squentropy · ICML 2023 |
Machine learning › Deep learning architectures and training
loss function design |
0.7 | 1 | 2023 | Cut your Losses with Squentropy · ICML 2023 |
Machine learning › Trustworthy machine learning › calibration
model calibration |
0.7 | 1 | 2023 | Cut your Losses with Squentropy · ICML 2023 |
Machine learning › Learning theory
loss function |
0.5 | 1 | 2021 | Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification Tasks · ICLR 2021 |
Methods — techniques the papers use, named apart from their topics
square loss · 1.2cross-entropy loss · 1.2gradient descent convergence analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural NetworksabstractNonlinear activation functions are widely recognized for enhancing the expressivity of neural networks, which is the primary reason for their widespread implementation. In this work, we focus on ReLU activation and reveal a novel and intriguing property of nonlinear activations. By comparing enabling and disabling the nonlinear activations in the neural network, we demonstrate their specific effects on wide neural networks: (a) *better feature separation*, i.e., a larger angle separation for similar data in the feature space of model gradient, and (b) *better NTK conditioning*, i.e., a smaller condition number of neural tangent kernel (NTK). Furthermore, we show that the network depth (i.e., with more nonlinear activation operations) further amplifies these effects; in addition, in the infinite-width-then-depth limit, all data are equally separated with a fixed angle in the model gradient feature space, regardless of how similar they are originally in the input space.
Note that, without the nonlinear activation, i.e., in a linear neural network, the data separation remains the same as for the original inputs
and NTK condition number is equivalent to the Gram matrix, regardless of the network depth. Due to the close connection between NTK condition number and convergence theories, our results imply that nonlinear activation
helps to improve the worst-case convergence rates of gradient based methods. Han Bi, Like Hui |
NeurIPS | 3 |
| 2023 | Cut your Losses with SquentropyabstractNearly all practical neural models for classification are trained using the cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Recent work (Hui & Belkin, 2020) suggests that training using the (rescaled) square loss is often superior in terms of the classification accuracy. In this paper we propose the "squentropy" loss, which is the sum of two terms: the cross-entropy loss and the average square loss over the incorrect classes. We provide an extensive set of experiment on multi-class classification problems showing that the squentropy loss outperforms both the pure cross-entropy and rescaled square losses in terms of the classification accuracy. We also demonstrate that it provides significantly better model calibration than either of these alternative losses and, furthermore, has less variance with respect to the random initialization. Additionally, in contrast to the square loss, squentropy loss can frequently be trained using exactly the same optimization parameters, including the learning rate, as the standard cross-entropy loss, making it a true ''plug-and-play'' replacement. Finally, unlike the rescaled square loss, multiclass squentropy contains no parameters that need to be adjusted. Like Hui, Mikhail Belkin, Stephen Wright |
ICML | 1 |
| 2021 | Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification Tasks
Like Hui, Mikhail Belkin |
ICLR | 1 |
| 2019 | Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust AsrabstractIn this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covariance matrix for the beamformer. Minimum Variance Distortionless Response (MVDR) beamformer and Generalized Eigenvalue (GEV) beamformer are both investigated under the CRM-based joint training architecture. We also propose a robust mask pooling strategy among multiple channels. A long short-term memory (LSTM) based language model is utilized to re-score hypotheses which further improves the overall performance. We evaluate the proposed methods on CHiME-4 challenge dataset. The CRM based system achieves a relative 10% reduction on word error rate (WER) compared with the IRM based system. Without sequence discriminative training, our best single system already achieves an average WER 2.72% on the test set which is comparable to the state-of-the-art. Yong Xu 0004, Chao Weng, Like Hui, Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 3 |
| 2019 | Kernel Machines Beat Deep Neural Networks on Mask-Based Single-Channel Speech EnhancementabstractWe apply a fast kernel method for mask-based single-channel speech enhancement.Specifically, our method solves a kernel regression problem associated to a non-smooth kernel function (exponential power kernel) with a highly efficient iterative method (EigenPro).Due to the simplicity of this method, its hyper-parameters such as kernel bandwidth can be automatically and efficiently selected using line search with subsamples of training data.We observe an empirical correlation between the regression loss (mean square error) and regular metrics for speech enhancement.This observation justifies our training target and motivates us to achieve lower regression loss by training separate kernel model per frequency subband.We compare our method with the state-of-the-art deep neural networks on mask-based HINT and TIMIT.Experimental results show that our kernel method consistently outperforms deep neural networks while requiring less training time. Like Hui, Mikhail Belkin |
INTERSPEECH | 1 |
| 2015 | High-performance Swahili keyword search with very limited language pack: The THUEE system for the OpenKWS15 evaluationabstractThis paper presents the Swahili keyword search system developed by the THUEE team for the OpenKWS15 evaluation, which is conducted by NIST under the IARPA Babel program. There are several highlights in the development of the system, including automatic generation of the pronunciation lexicon, aggressive data augmentation, the multilingual bottleneck feature extractor trained from 6 languages, text selection from web data for language model training, semi-supervised training for acoustic models and language models, out-of-vocabulary keyword detection using morphemes and a rich diversity of the systems for combination. A wide variety of acoustic modeling techniques are explored and compared. Up to 12 different individual systems are used for combination. The system achieves the state-of-the-art performance in the required condition of the evaluation. Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Jia Liu 0001 |
ASRU | 5 |
| 2015 | Improved system fusion for keyword searchabstractIt has been demonstrated that system fusion can significantly improve the performance of keyword search. In this paper, we compare the performance of several widely-used arithmetic-based fusion methods using different normalization pipeline and try to find the best pipeline. A novel arithmetic-based fusion method is proposed in this work. The method supplies a more effective way to incorporate the number of systems which have non-zero scores for a detection. When tested on the development test dataset of the OpenKWS15 Evaluation, the proposed method achieves the highest maximum term-weighted value (MTWV) and actual term-weighted value (ATWV) among all other arithmetic-based fusion methods. Usually, discriminative fusion methods employing classifiers can outperform arithmetic-based fusion methods. A DNN-based fusion method is explored in this work. After word-burst information is added, the DNN-based fusion method outperforms all other methods. In addition, it is notable that our arithmetic-based method achieves the same MTWV as the DNN-based method. Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Weiqiang Zhang 0001, Jia Liu 0001 |
ASRU | 5 |