EDBT 2026 Demo / reviewers in the wild / expert
Zejiang Hou
dblp:206/8398
· DBLP profile ↗
20ranked-venue papers
14as first author
13since 2021 · last 2025
0000-0002-6836-8288ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 7 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hyper-adapter for Parameter-Efficient Multilingual ASR AdaptationabstractThis work proposes a new parameter-efficient adaptation approach for multilingual ASR based on the hyper-network. Existing multilingual ASR adaptation methods apply either one residual adapter for all the languages, or language dependent adapters for each individual language. The residual adapter cannot compete the full finetuning in terms of WER, because it is agnostic to language information. Whereas the language dependent adapters introduce high parameter overhead without a parameter sharing strategy. In contrast, we leverage a hyper-network to generate the weights for the adapters across different languages. To achieve the best parameter sharing strategy that scales with a large number of languages, we propose multi-level conditioning vector fusion, orthogonal regularization to improve the hyper-network output diversity, and language loss weighting during the model training. The proposed approach demonstrates comparable or better WER and better parameter efficiency compared to previous multilingual ASR adaptation approaches on commonly used multilingual ASR benchmarks. Zejiang Hou, Daniel Garcia-Romero, Kyu J. Han |
ICASSP | 1 |
| 2025 | Masked Image Pretraining on Language Assisted RepresentationabstractSelf-attention based transformer models have been dominating many computer vision tasks in the past few years. Their superb model qualities heavily depend on the excessively large labeled image datasets. In order to reduce the reliance on large labeled datasets, reconstruction based masked autoencoders are gaining popularity, which learn high quality transferable representations from unlabeled images. For the same purpose, recent weakly supervised image pretraining methods explore language supervision from text captions accompanying the images. In this work, we propose Masked Image pretraining on Language Assisted representatioN, dubbed as MILAN. Instead of predicting raw pixels or low level features, our pretraining objective is to reconstruct the image features with substantial semantic signals that are obtained using caption supervision. Moreover, to accommodate our reconstruction target, we propose a more efficient prompting decoder architecture and a semantic aware mask sampling mechanism, which further advance the transfer performance of the pretrained model. Experimental results demonstrate that MILAN delivers higher accuracy than the previous works. When the masked autoencoder is pretrained and finetuned on ImageNet-1K dataset with an input resolution of 224×224, MILAN achieves a top-1 accuracy of 85.4% on ViT-Base, surpassing previous state-of-the-arts by 1%. In the downstream semantic segmentation task, MILAN achieves 52.7 mIoU using ViT-Base on ADE20K dataset, outperforming previous masked pre-training results by 4 points. Zejiang Hou, Sun-Yuan Kung |
ICASSP | 1 |
| 2025 | Zero-resource Speech Translation and Recognition with LLMsabstractDespite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language. Karel Mundnich, Xing Niu 0001, Prashant Mathur, Srikanth Ronanki, Brady Houston, Veera Raghavendra Elluru, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff |
ICASSP | 8 |
| 2025 | Contextual ASR with Retrieval Augmented Large Language ModelabstractAutomatic speech recognition (ASR) systems can benefit from incorporating contextual information to improve recognition accuracy, especially for uncommon words or phrases. Current approaches like custom vocabularies or prompting with previous transcript segments provide limited contextual control. Compared to existing context biasing methods, RAG promises more flexible and scalable contextual control by leveraging LLMs’ broad knowledge. To this end, we propose leveraging large language models (LLMs) and retrieval-augmented generation (RAG) to enhance the contextual capabilities of ASR systems. Specifically, we propose systems based on text and audio LLMs to perform contextual error correction with context retrieved by querying a text-based retriever using the ASR module’s firstpass ASR hypotheses and a frequency-based custom vocabulary (CV) list. Our experiments reveal that the fine-tuned system has effectively learned to extract the relevant context to perform error correction while maintaining robustness against noise. Cihan Xiao, Zejiang Hou, Daniel Garcia-Romero, Kyu J. Han |
ICASSP | 2 |
| 2024 | Revisiting Convolution-free Transformer for Speech Recognition
Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff |
INTERSPEECH | 1 |
| 2024 | Improving Multilingual ASR Robustness to Errors in Language Input
Brady Houston, Omid Sadjadi, Zejiang Hou, Srikanth Vishnubhotla, Kyu J. Han |
INTERSPEECH | 3 |
| 2022 | CHEX: CHannel EXploration for CNN Model CompressionabstractChannel pruning has been broadly recognized as an effective technique to reduce the computation and memory cost of deep convolutional neural networks. However, conventional pruning methods have limitations in that: they are restricted to pruning process only, and they require a fully pre-trained large model. Such limitations may lead to sub-optimal model quality as well as excessive memory and training cost. In this paper, we propose a novel Channel Exploration methodology, dubbed as CHEX, to rectify these problems. As opposed to pruning-only strategy, we propose to repeatedly prune and regrow the channels throughout the training process, which reduces the risk of pruning important channels prematurely. More exactly: From intra-Layer's aspect, we tackle the channel pruning problem via a well-known column subset selection (CSS) formulation. From inter-Layer's aspect, our regrowing stages open a path for dynamically re-allocating the number of channels across all the layers under a global channel sparsity constraint. In addition, all the exploration process is done in a single training from scratch without the need of a pre-trained large model. Experimental results demonstrate that CHEX can effectively reduce the FLOPs of diverse CNN architectures on a variety of computer vision tasks, including image classification, object detection, instance segmentation, and 3D vision. For example, our compressed ResNet-50 model on ImageNet dataset achieves 76% top-l accuracy with only 25% FLOPs of the original ResNet-50 model, outperforming previous state-of-the-art channel pruning methods. The checkpoints and code are available at here. Zejiang Hou, Minghai Qin, Fei Sun 0002, Kun Yuan 0001, Yi Xu 0008, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001, Sun-Yuan Kung |
CVPR | 1 |
| 2022 | Effective Model Sparsification by Scheduled Grow-and-Prune Methods
Minghai Qin, Fei Sun 0002, Zejiang Hou, Kun Yuan 0001, Yi Xu 0008, Yanzhi Wang 0001, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001 |
ICLR | 4 |
| 2022 | Multi-Dimensional Model Compression of Vision TransformerabstractVision transformers (ViT) have recently attracted considerable attentions, but the huge computational cost remains an issue for practical deployment. Previous ViT pruning methods tend to prune the model along one dimension solely, which may suffer from excessive reduction and lead to sub-optimal model quality. In contrast, we advocate a multi-dimensional ViT compression paradigm, and propose to harness the redundancy reduction from attention head, neuron and sequence dimensions jointly. We firstly propose a statistical dependence based pruning criterion that is generalizable to different dimensions for identifying deleterious components. Moreover, we cast the multi-dimensional compression as an optimization, learning the optimal pruning policy across the three dimensions that maximizes the compressed model's accuracy under a computational budget. The problem is solved by our adapted Gaussian process search with expected improvement. Experimental results show that our method effectively reduces the computational cost of various ViT models. For example, our method reduces 40% FLOPs without top-1 accuracy loss for DeiT and T2T-ViT models, outperforming previous state-of-the-arts. Zejiang Hou, Sun-Yuan Kung |
ICME | 1 |
| 2022 | Multi-Dimensional Dynamic Model Compression for Efficient Image Super-ResolutionabstractModern single image super-resolution (SR) system based on convolutional neural networks achieves substantial progress. However, most SR deep networks are computationally expensive and require excessively large activation memory footprints, impeding their effective deployment to resource-limited devices. Based on the observation that the activation patterns in SR networks exhibit high input-dependency, we propose Multi-Dimensional Dynamic Model Compression method that can reduce both spatial and channel wise redundancy in an SR deep network for different input images. To reduce the spatial-wise redundancy, we propose to perform convolution on scaled-down feature-maps where the down-scaling factor is made adaptive to different input images. To reduce the channel-wise redundancy, we introduce a low-cost channel saliency predictor for each convolution to dynamically skip the computation of unimportant channels based on the Gumbel-Softmax. To better capture the feature-maps information and facilitate input-adaptive decision, we employ classic image processing metrics, e.g., Spatial Information, to guide the saliency predictors. The proposed method can be readily applied to a variety of SR deep networks and trained end-to-end with standard super-resolution loss, in combination with a sparsity criterion. Experiments on several benchmarks demonstrate that our method can effectively reduce the FLOPs of both lightweight and non-compact SR models with negligible PSNR loss. Moreover, our compressed models achieve competitive PSNR-FLOPs Pareto frontier compared with SOTA NAS-based SR methods. Zejiang Hou, Sun-Yuan Kung |
WACV | 1 |
| 2022 | Meta-Learning the Difference: Preparing Large Language Models for Efficient AdaptationabstractAbstract Large pretrained language models (PLMs) are often domain- or task-adapted via finetuning or prompting. Finetuning requires modifying all of the parameters and having enough data to avoid overfitting while prompting requires no training and few examples but limits performance. Instead, we prepare PLMs for data- and parameter-efficient adaptation by learning to learn the difference between general and adapted PLMs. This difference is expressed in terms of model weights and sublayer structure through our proposed dynamic low-rank reparameterization and learned architecture controller. Experiments on few-shot dialogue completion, low-resource abstractive summarization, and multi-domain language modeling show improvements in adaptation time and performance over direct finetuning or preparation via domain-adaptive pretraining. Ablations show our task-adaptive reparameterization (TARP) and model search (TAMS) components individually improve on other parameter-efficient transfer like adapters and structure-learning methods like learned sparsification. Zejiang Hou, Julian Salazar, George Polovets |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Parameter Efficient Dynamic Convolution via Tensor Decomposition
Zejiang Hou, Sun-Yuan Kung |
BMVC | 1 |
| 2021 | Meta-Learning with Attention for Improved Few-Shot LearningabstractWe consider few-shot learning (FSL), where a model learns from very few labeled examples such that it can generalize to unseen examples. Model-agnostic meta-learning (MAML) has been proposed to solve FSL. However, the low performance of MAML suggests its difficulty in tackle diverse tasks, due to the restriction of sharing a single model initialization for fast adaptation. In this paper, we propose meta-learning with attention mechanisms. Our method meta-learns attention modules to instantiate task-specific model initialization for fast adaptation, which can obtain high-quality solution to a new task using few gradient descent steps. To further improve generalization during inference, we propose to incorporate an entropy regularizer into the adaptation objective to penalize the Shannon entropy of prediction probability. Extensive experiments under various FSL scenarios show that our method achieves state-of-the-art performance on the mini-ImageNet and tiered-ImageNet. Zejiang Hou, Anwar Elwalid, Sun-Yuan Kung |
ICASSP | 1 |
| 2020 | Scalable Kernel Learning Via the Discriminant InformationabstractKernel approximation methods create explicit, low-dimensional kernel feature maps to deal with the high computational and memory complexity of standard techniques. This work studies a supervised kernel learning methodology to optimize such mappings. We utilize the Discriminant Information criterion, a measure of class separability with a strong connection to Discriminant Analysis. By generalizing this measure to cover a wider range of kernel maps and learning settings, we develop scalable methods to learn kernel features with high discriminant power. Experimental results on several datasets showcase that our techniques can improve optimization and generalization performances over state of the art kernel learning methods. Mert Al, Zejiang Hou, Sun-Yuan Kung |
ICASSP | 2 |
| 2020 | Efficient Image Super Resolution Via Channel Discriminative Deep Neural Network PruningabstractDeep convolutional neural networks (CNN) have demonstrated superior performance in image super-resolution (SR) problem. However, CNNs are known to be heavily over-parameterized, and suffer from abundant redundancy. The growing size of CNNs may be incompatible with their deployment on mobile or embedded devices. Network pruning has benefited classification tasks by removing redundant parameters and associated computation. However, it has rarely been studied for SR, because existing methods assume the channel-wise features are of equal importance to the final reconstruction. On the contrary, we show the existence of uninformative feature-maps with no contribution to the task. In order to identify and remove such uninformative channels, we propose a new pruning criterion, Discriminant Information, by characterizing the dependency of the output w.r.t to the hidden-layer feature-maps. Empirically, our DI-based channel pruning algorithm is able to trim the state-of-the-art SR networks significantly (e.g. 8.7x model size compression and 3.6x CPU acceleration on SRResNet), with no quantitative or visual performance loss. Zejiang Hou, Sun-Yuan Kung |
ICASSP | 1 |
| 2020 | Hierarchically Aggregated Residual Transformation for Single Image Super ResolutionabstractVisual patterns usually appear at different scales/sizes in natural images. Multi-scale feature representation is of great importance for the single-image super-resolution (SISR) task to reconstruct image objects at different scales. However, such characteristic has been rarely considered by CNN-based SISR methods. In this work, we propose a novel building block, i.e. hierarchically aggregated residual transformation (HART), to achieve multi-scale feature representation in each layer of the network. Within each HART block, we connect multiple convolutions in a hierarchical residual-like manner, which efficiently provides a wide range of effective receptive fields at a more granular level to detect both local and global image features. To theoretically understand the proposed HART block, we recast SISR as an optimal control problem and show that HART effectively approximates the classical 4th-order Runge-Kutta method, which has the merit of small local truncation error for solving numerical ordinary differential equation. By cascading the proposed HART blocks, we establish our high-performing HARTnet. Through extensive experiments on various benchmark datasets under different degradation models, we demonstrate that HARTnet compares favourably against existing state-of-the-art methods (including those in the NTIRE 2019 SR Challenge leaderboard) in terms of both quantitative metrics and visual quality. Moreover, the same HARTnet architecture achieves promising performance on such other image restoration tasks as image denoising and low-light image enhancement. Zejiang Hou, Sun-Yuan Kung |
ICPR | 1 |
| 2020 | A Discriminant Information Approach to Deep Neural Network PruningabstractNetwork pruning has become the de facto tool to accelerate deep neural networks for mobile and edge applications. Recently, feature-map discriminant based channel pruning has shown promising results, as it aligns well with the CNN's objective of differentiating multiple classes and offers better interpretability of the pruning decision. However, existing discriminant-based methods are challenged by computation inefficiency, as there is a lack of theoretical guidance on quantifying the feature-map discriminant power. In this paper, we develop a mathematical formulation to accurately and efficiently quantify the feature-map discriminativeness, which gives rise to a novel criterion, Discriminant Information (DI). We analyze the theoretical property of DI, specifically the non-decreasing property, that makes DI a valid channel selection criterion. By measuring the differential discriminant, we can identify and remove those channels with minimum influence to the discriminant power. The versatility of DI criterion also enables an intra-layer mixed precision quantization to further compress the network. Moreover, we propose a DI-based greedy pruning algorithm and structure distillation technique to automatically decide the pruned structure that satisfies certain resource budget, which is a common requirement in reality. Extensive experiments demonstrate the effectiveness of our method: our pruned ResNet50 on ImageNet achieves 44% FLOPs reduction without any Top-1 accuracy loss compared to unpruned model. Zejiang Hou, Sun-Yuan Kung |
ICPR | 1 |
| 2019 | Methodical Design and Trimming of Deep Learning Networks: Enhancing External BP Learning with Internal Omnipresent-supervision Training ParadigmabstractBack-propagation (BP) is now a classic learning paradigm whose source of supervision is exclusively from the external (input/output) nodes. Consequently, BP is easily vulnerable to curse-of-depth in (very) Deep Learning Networks (DLNs). This prompts us to advocate Internal Neuron’s Learnablility (INL) with (1)internal teacher labels (ITL); and (2)internal optimization metrics (IOM) for evaluating hidden layers/nodes. Conceptually, INL is a step beyond the notion of Internal Neuron’s Explainablility (INE), championed by DARPA’s XAI (or AI3.0). Practically, INL facilitates a structure/parameter NP-iterative learning for (supervised) deep compression/quantization: simultaneously trimming hidden nodes and raising accuracy. Pursuant to our simulations, the NP-iteration appears to outperform several prominent pruning methods in the literature. Sun-Yuan Kung, Zejiang Hou, Yuchen Liu 0002 |
ICASSP | 2 |
| 2019 | A Kernel Discriminant Information Approach to Non-linear Feature SelectionabstractFeature selection has become a de facto tool for analyzing high-dimensional data, especially in bioinformatics. It is effective in improving learning algorithms' scalability and facilitating feature generalization or interpretability by removing noise and redundancy. Our focus is placed on the paradigm of supervised feature selection, which aims to find an optimal feature subset to best predict the target. We propose a nonlinear approach for finding a feature subset that achieves the highest inter-class separability in terms of the kernel Discriminant Information (KDI) measure. Theoretically, we prove the existence of good prediction hypotheses for feature subsets with high KDI value. We also establish the equivalency between maximizing the KDI statistic and minimizing a functional dependency measure of label variable on data. Moreover, we asymptotically prove the concentration property of the optimal feature subset found by maximizing the KDI measure. Practically, we provide an efficient gradient optimization algorithm for solving the KDI feature selection problem. We evaluate the proposed method based on 19 benchmark datasets in various domains, and demonstrates a noticeable improvement against state-of-the-art baselines on the majority of classification and regression tasks. Notably, our method is robust to the choice of hyper-parameters, works well with various downstream classifiers, has competitive computational complexity among the kernel based methods considered, and scales well the large-scale object recognition dataset, with generalization enhancement on CIFAR. Zejiang Hou, Sun-Yuan Kung |
IJCNN | 1 |
| 2017 | Distributed optimal power flow: An Augmented Lagrangian-Sequential Quadratic Programming approachabstractThis paper presents a distributed optimal power flow approach based on Augmented Lagrangian (AL) and Sequential Quadratic Programming (SQP). It is able to separate the OPF into smaller sub-problems, which could be iteratively solved individually using the SQP. This utilizes the SQP for largescale problems with non-linear objective functions and constraints. Simulation and comparison using the IEEE 30 and 118 buses examples show that the proposed distributed approach is able to achieve comparable performance with other benchmark centralized solvers provided by the FMINCON in MATPOWER. This suggests the proposed approach may serve an attractive alternative to other OPF algorithms. Zejiang Hou, Ho-Chun Wu 0001, S. C. Chan 0001 |
ISCAS | 1 |