EDBT 2026 Demo / reviewers in the wild / expert
Yufei Cui
dblp:188/0049
· DBLP profile ↗
36ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-9663-0079ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 16 since 2021Systems, architecture and hardware · 12 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMsabstractPost-training hybridization of large language models (LLMs) often replaces quadratic selfattention with sliding-window attention (SWA) to reduce KV cache usage and improve latency.Existing hybridization schemes are typically defined either at the layer level (e.g., interleaving) or at the head level via static rankings from local to global.Layer-level schemes ignore that local and global dependencies are routed through heads within the same layer, while static head-level rankings suffer from entanglement: a head's local/global behavior can change after hybridization.We propose BOSCH, Black-box Binary Optimization for Short-context Head Selection, a training-free method that formulates the problem as a Large Neighborhood Search and decomposes it into three subproblems: (i) layer-importance detection via small-budget black-box probes, (ii) adaptive per-layer SWA-ratio assignment based on these sensitivities, and (iii) grouped headlevel optimization within ratio buckets.Extensive experiments on 4 LLMs ranging from 1.7B to 30B parameters, across 4 SWA ratios, show that BOSCH consistently outperforms layerlevel heuristics and 6 strong static head-level methods, with larger gains at higher SWA ratios.Under continual pretraining, BOSCH recover original long-context performance faster and to a higher level.Analysis of the selected heads reveals substantial turnover for BOSCH across different SWA ratios, underscoring the importance of performing head-level selection for each target ratio rather than relying on fixed locality rankings. Abbas Ghaddar, Ivan Kobyzev, Boxing Chen, Yufei Cui |
ACL (1) | 4 |
| 2026 | MATCH: Modulating Attention via In-Context Retrieval for Long-Context TransformersabstractLinrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu, Xihao Yuan, Hanting Chen, Kai Han, Xinghao Chen, Chengjun Zhan, Hanlin xu, Yichun Yin, Lifeng Shang, Feng Wen, Boxing Chen, Yufei Cui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Linrui Ma, Chun Hei Lo, Xinyu Wang 0061, Peng Lu 0006, Xihao Yuan, Hanting Chen, Kai Han 0002, Xinghao Chen 0001, Chengjun Zhan, Hanlin Xu, Yichun Yin, Lifeng Shang, Boxing Chen, Yufei Cui |
ACL (1) | 15 |
| 2026 | ColdCode: Cold Data Encoding for Enhanced Reliability and Lifetime in 3D NAND FlashabstractCold storage, which stores rarely accessed data, dominates modern data centers but is poorly served by NAND flash's data randomization. As a common practice today by flash vendors, data randomization is applied in NAND flash chips to avoid extreme data patterns that generate the worst-case raw bit error rate (RBER). However, this paper demonstrates that data randomization rules out the opportunity to explore data patterns with very low RBERs, through a comprehensive analysis on data randomization in 3D NAND flash chips (across 8 models). Motivated by this, we propose ColdCode, a novel data coding framework to replace the conventional randomizer in 3D high-density flash for cold data storage. Using a tag that indicates coldness information passed from the file system to solid-state drive (SSD) controllers, the controller encodes cold data to enhance reliability and extend the lifetime. ColdCode employs two coding techniques: skewed coding and reversed Huffman coding, applied based on the data entropy. These techniques effectively reduce the RBER of encoded data compared to conventional randomization. Experimental results on real high-density 3D flash chips show that, under the same error conditions, the skewed coding and reversed Huffman coding reduce the average RBER by 42% and 30%, respectively. Consequently, the lifetime of the flash chips is prolonged by factors of 2.97× and 1.7×, respectively, compared to data randomization. Qiao Li 0001, Shangyu Wu, Yufei Cui, Jie Zhang 0048, Chun Jason Xue |
EuroSys | 4 |
| 2026 | R2 R: A Post-training Framework for Multi-domain Decoder-Only Rerankers
Hanwei Wu, Qingchen Hu, Zhenghan Tai, Jingrui Tian, Lei Ding 0013, Jijun Chi, Hailin He, Tung Sum Thomas Kwok, Yufei Cui, Sicheng Lyu, Muzhi Li 0001, Peng Lu 0006, Xinyu Wang 0061 |
PAKDD (2) | 9 |
| 2026 | VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question AnsweringabstractRetrieval-Augmented Generation (RAG) is becoming increasingly essential for Question Answering (QA) in the financial sector, where accurate and contextually grounded insights from complex public disclosures are crucial. However, existing financial RAG systems face two significant challenges: (1) they struggle to process heterogeneous data formats, such as text, tables, and figures; and (2) they encounter difficulties in balancing general-domain applicability with company-specific adaptation. To overcome these challenges, we present VeritasFi, an innovative hybrid RAG framework that incorporates a multi-modal preprocessing pipeline alongside a cutting-edge two-stage training strategy for its re-ranking component. VeritasFi enhances financial QA through three key innovations: (1) A multi-modal preprocessing pipeline that seamlessly transforms heterogeneous data into a coherent, machine-readable format. (2) A tripartite hybrid retrieval engine that operates in parallel, combining deep multi-path retrieval over a semantically indexed document corpus, real-time data acquisition through tool utilization, and an expert-curated memory bank for high-frequency questions, ensuring comprehensive scope, accuracy, and efficiency. (3) A two-stage training strategy for the document re-ranker, which initially constructs a general, domain-specific model using anonymized data, followed by rapid fine-tuning on company-specific data for targeted applications. By integrating our proposed designs, VeritasFi presents a novel framework that greatly enhances the adaptability and robustness of financial RAG systems, providing a scalable solution for both general-domain and company-specific QA tasks. Code accompanying this work is available at https://github.com/simplew4y/VeritasFi.git. Zhenghan Tai, Hanwei Wu, Qingchen Hu, Jijun Chi, Hailin He, Lei Ding 0013, Tung Sum Thomas Kwok, Bohuai Xiao, Yuchen Hua, Suyuchen Wang, Peng Lu 0006, Muzhi Li 0001, Yihong Wu 0006, Liheng Ma, Jerry Huang, Jiayi Zhang 0017, Gonghao Zhang, Chaolong Jiang, Jingrui Tian, Sicheng Lyu, Fengran Mo, Yufei Cui, Xinyu Wang 0061 |
WWW | 25 |
| 2025 | Transtreaming: Adaptive Delay-aware Transformer for Real-time Streaming PerceptionabstractReal-time object detection is critical for the decision-making process for many real-world applications, such as collision avoidance and path planning in autonomous driving. This work presents an innovative real-time streaming perception method, Transtreaming, which addresses the challenge of real-time object detection with dynamic computational delays. The core innovation of Transtreaming lies in its adaptive delay-aware transformer, which can concurrently predict multiple future frames and select the output that best matches the real-world present time, compensating for any system-induced computational delays. The proposed model outperforms existing state-of-the-art methods, even in single-frame detection scenarios, by leveraging a transformer-based methodology. It demonstrates robust performance across a range of devices, from powerful V100 to modest 2080Ti, achieving the highest level of perceptual accuracy on all platforms. Unlike most state-of-the-art methods that struggle to complete computation within a single frame on less powerful devices, Transtreaming meets the stringent real-time processing requirements on all kinds of devices. The experimental results emphasize the system's adaptability and its potential to significantly improve the safety and reliability of many real-world systems, such as autonomous driving. Yufei Cui, Chenchen Fu, Weiwei Wu 0001 |
AAAI | 2 |
| 2025 | GeneQuery: Generalized Gene Expression Prediction from Histology Images via Image-Gene QAabstractGene expression profiling provides profound insights into molecular mechanisms, but its time-consuming and costly nature often presents significant challenges. Recent advancements have utilized histological images to predict spatially resolved gene expression profiles. The gene prediction problem has two main characteristics, i.e., spatial heterogeneity and gene interdependency. Existing works only focus on addressing spatial heterogeneity, ignoring the importance of relationships between genes. To address the above limitation, this paper presents GeneQuery, which aims to solve this gene expression prediction task in a question-answering (QA) manner for better generality and flexibility. Specifically, GeneQuery takes gene meta-information as queries and whole-slide images as contexts and then predicts the queried gene expression values. GeneQuery learns to dynamically fuse spatial image features with gene semantics, and uses an attention mechanism to explicitly capture the spatial information and a shared regressor to implicitly capture the gene relationship. A variant of GeneQuery can use the attention mechanism to capture gene relationships for more complex tissue cases. This QA-based reformulation also grants the model the ability to generalize, enabling the prediction of unseen gene expression without requiring model retraining. Comprehensive experiments on spatial transcriptomics datasets show that the proposed GeneQuery outperforms existing state-of-the-art methods on known and unseen genes. More results also demonstrate that GeneQuery can help analyze the tissue structure. Linjing Liu, Yufei Cui, Shangyu Wu, Xue (Steve) Liu, Antoni B. Chan, Chun Jason Xue |
BIBM | 3 |
| 2025 | FinSage: A Multi-aspect RAG System for Financial Filings Question AnsweringabstractLeveraging large language models in real-world settings often entails a need to utilize domain-specific data and tools in order to follow the complex regulations that need to be followed for acceptable use. Within financial sectors, modern enterprises increasingly rely on Retrieval-Augmented Generation (RAG) systems to address complex information retrieval in financial document workflows. However, existing solutions struggle to account for the inherent heterogeneity of data (e.g., text, tables, diagrams) and evolving complexity in financial filings, leading to compromised accuracy in critical information extraction. We propose the FinSage framework as a solution, utilizing a multi-aspect RAG framework tailored for data retrieval and summarization in multi-modal financial documents. øurmodel introduces three innovative components: (1) a multi-modal pre-processing pipeline that unifies diverse data formats and generates chunk-level metadata summaries, (2) a multi-path sparse-dense retrieval system augmented with query expansion (HyDE) and metadata-aware semantic search, and (3) a domain-specialized re-ranking module fine-tuned via Direct Preference Optimization to prioritize ground-truth-related content. Extensive experiments demonstrate that FinSage achieves an impressive recall of 92.51% on 75 expert-curated questions derived from surpasses the best baseline method on the FinanceBench question answering datasets by 24.06% in accuracy. Moreover, FinSage has been successfully deployed as financial question-answering system in online meetings, where it has already served more than 1,200 people. The implementation is publicly available at https://github.com/simplew4y/finsage. Xinyu Wang 0061, Jijun Chi, Zhenghan Tai, Tung Sum Thomas Kwok, Hailin He, Zhuhong Li, Yuchen Hua, Muzhi Li 0001, Peng Lu 0006, Suyuchen Wang, Yihong Wu 0006, Jerry Huang, Jingrui Tian, Fengran Mo, Yufei Cui |
CIKM | 15 |
| 2025 | Advancing Multiple Instance Learning with Continual Learning for Whole Slide ImagingabstractAdvances in medical imaging and deep learning have propelled progress in whole slide image (WSI) analysis, with multiple instance learning (MIL) showing promise for efficient and accurate diagnostics. However, conventional MIL models often lack adaptability to evolving datasets, as they rely on static training that cannot incorporate new information without extensive retraining. Applying continual learning (CL) to MIL models is a possible solution, but often sees limited improvements. In this paper, we analyze CL in the context of attention MIL models and find that the model forgetting is mainly concentrated in the attention layers of the MIL model. Using the results of this analysis we propose two components for improving CL on MIL: Attention Knowledge Distillation (AKD) and the Pseudo-Bag Memory Pool (PMP). AKD mitigates catastrophic forgetting by focusing on retaining attention layer knowledge between learning sessions, while PMP reduces the memory footprint by selectively storing only the most informative patches, or "pseudo-bags" from WSIs. Experimental evaluations demonstrate that our method significantly improves both accuracy and memory efficiency on diverse WSI datasets, outperforming current state-of-the-art CL methods. This work provides a foundation for CL in large-scale, weakly annotated clinical datasets, paving the way for more adaptable and resilient diagnostic models. Xianrui Li, Yufei Cui, Antoni B. Chan |
CVPR | 2 |
| 2025 | PoT-PTQ: Two-Step Power-of-Two Post-Training for LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable performance across various natural language processing (NLP) tasks. However, their deployment is challenging due to the substantial computational resources required. Power-of-two (PoT) quantization is a general tool to counteract this difficulty. Albeit previous works on PoT quantization can be efficiently dequantized on CPUs using fixed-point addition, it showed less effectiveness on GPUs. The reason is entanglement of the sign bit and sequential bit manipulations needed for dequantization. We propose a novel POT quantization framework for LLM weights that (i) outperforms state-of-the-art accuracy in extremely low-precision number formats, and (ii) enables faster inference through more efficient dequantization. To maintain the accuracy of the quantized model, we introduce a two-step post-training algorithm: (i) initialize the quantization scales with a robust starting point, and (ii) refine these scales using a minimal calibration set. The performance of our PoT post-training algorithm surpasses the current state-of-the-art in integer quantization, particularly at low precisions such as 2- and 3-bit formats. Our PoT quantization accelerates the dequantization step required for the floating point inference and leads to 3.67× speed up on a NVIDIA V100, and 1.63× on a NVIDIA RTX 4090, compared to uniform integer dequantization. Xinyu Wang 0061, Vahid Partovi Nia, Peng Lu 0006, Jerry Huang, Xiao-Wen Chang, Boxing Chen, Yufei Cui |
ECAI | 7 |
| 2025 | RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented LearningabstractAs end-to-end autonomous driving advances toward real-world deployment, ensuring the safety of autonomous vehicles (AVs) has become a critical requirement for their commercial viability. While rule-based AVs have traditionally undergone rigorous testing in both real-world and simulated environments before deployment, data-driven autonomous models are typically trained on real-world datasets, limiting their generalization to simulation environments. This poses a significant challenge for the development and testing of end-to-end autonomous driving. To address this issue, we propose Retrieval-Augmented Learning for Autonomous Driving (RALAD), a novel framework designed to bridge the real-to-sim gap in a cost-effective manner. RALAD consists of three key components: (1) domain adaptation via an enhanced Optimal Transport (OT) method, which retrieves the most similar scenarios between real and simulated environments; (2) feature fusion across similar scenarios, enabling the construction of a feature mapping between real-world and simulated domains; and (3) feature extraction freezing with fine-tuning on the fused features, allowing the model to learn simulation-specific characteristics through feature mapping. We evaluate RALAD on three monocular 3D object detection models, and the results demonstrate that our approach significantly improves model accuracy in simulation. Additionally, we use real autonomous vehicle for testing in real-world scenarios, and have established simulated scenes similar to reality for further testing, which illustrate the effectiveness of our method. Jiacheng Zuo, Zikang Zhou, Yufei Cui, Ziquan Liu, Jianping Wang 0001, Nan Guan, Jin Wang 0009, Chun Jason Xue |
IROS | 4 |
| 2025 | Mamba Modulation: On the Length Generalization of Mamba ModelsabstractThe quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged as a leading architecture, achieving state-of-the-art results across a range of language modeling tasks. However, Mamba’s performance significantly deteriorates when applied to contexts longer than those seen during pre-training, revealing a sharp sensitivity to context length extension. Through detailed analysis, we attribute this limitation to the out-of-distribution behavior of its state-space dynamics, particularly within the parameterization of the state transition matrix $A$. Unlike recent works which attribute this sensitivity to the vanished accumulation of discretization time steps, $\exp(-\sum_{t=1}^N{\Delta}_t)$, we establish a connection between state convergence behavior as the input length approaches infinity and the spectrum of the transition matrix $A$, offering a well-founded explanation of its role in length extension. Next, to overcome this challenge, we propose an approach that applies spectrum scaling to pre-trained Mamba models to enable robust long-context generalization by selectively modulating the spectrum of $A$ matrices in each layer. We show that this can significantly improve performance in settings where simply modulating ${\Delta}_t$ fails, validating our insights and providing avenues for better length generalization of state-space models with structured transition matrices. Peng Lu 0006, Jerry Huang, Qiuhao Zeng, Xinyu Wang 0061, Boxing Chen, Philippe Langlais, Yufei Cui |
NeurIPS | 7 |
| 2024 | ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation FusionabstractRetrieval-based augmentations (RA) incorporating knowledge from an external database into language models have greatly succeeded in various knowledge-intensive (KI) tasks. However, integrating retrievals in non-knowledge-intensive (NKI) tasks is still challenging.
Existing works focus on concatenating retrievals with inputs to improve model performance. Unfortunately, the use of retrieval concatenation-based augmentations causes an increase in the input length, substantially raising the computational demands of attention mechanisms.
This paper proposes a new paradigm of RA named \textbf{ReFusion}, a computation-efficient \textbf{Re}trieval representation \textbf{Fusion} with bi-level optimization. Unlike previous works, ReFusion directly fuses the retrieval representations into the hidden states of models.
Specifically, ReFusion leverages an adaptive retrieval integrator to seek the optimal combination of the proposed ranking schemes across different model layers. Experimental results demonstrate that the proposed ReFusion can achieve superior and robust performance in various NKI tasks. Shangyu Wu, Yufei Cui, Xue (Steve) Liu, Buzhou Tang, Tei-Wei Kuo, Chun Jason Xue |
ICLR | 3 |
| 2024 | The Pitfalls and Promise of Conformal Inference Under Adversarial AttacksabstractIn safety-critical applications such as medical imaging and autonomous driving, where decisions have profound implications for patient health and road safety, it is imperative to maintain both high adversarial robustness to protect against potential adversarial attacks and reliable uncertainty quantification in decision-making. With extensive research focused on enhancing adversarial robustness through various forms of adversarial training (AT), a notable knowledge gap remains concerning the uncertainty inherent in adversarially trained models. To address this gap, this study investigates the uncertainty of deep learning models by examining the performance of conformal prediction (CP) in the context of standard adversarial attacks within the adversarial defense community. It is first unveiled that existing CP methods do not produce informative prediction sets under the commonly used $l_{\infty}$-norm bounded attack if the model is not adversarially trained, which underpins the importance of adversarial training for CP. Our paper next demonstrates that the prediction set size (PSS) of CP using adversarially trained models with AT variants is often worse than using standard AT, inspiring us to research into CP-efficient AT for improved PSS. We propose to optimize a Beta-weighting loss with an entropy minimization regularizer during AT to improve CP-efficiency, where the Beta-weighting loss is shown to be an upper bound of PSS at the population level by our theoretical analysis. Moreover, our empirical study on four image classification datasets across three popular AT baselines validates the effectiveness of the proposed Uncertainty-Reducing AT (AT-UR). Ziquan Liu, Yufei Cui, Yan Yan 0006, Yi Xu 0008, Xiangyang Ji, Xue (Steve) Liu, Antoni B. Chan |
ICML | 2 |
| 2024 | ProtoTree-MIL: Interpretable Multiple Instance Learning for Whole Slide Image ClassificationabstractWhole slide image (WSI) classification is one of the important fields of digital pathology, and is generally solved as a weakly supervised learning problem by adopting multiple instance learning (MIL). However, a common but crucial challenge faced by existing MIL models is their inability to provide convincing explanations that can win the trust of pathologists and be applied to clinical diagnosis. In addition, most attention-based MIL models use attention scores to represent the importance of each patch in the WSI rather than inferring patch probabilities directly, which does not accurately detect the critical patches. To address these two challenges, we propose a ProtoTree based MIL model for WSI classification, called ProtoTree-MIL, where ProtoTree is an interpretable model that combines the advantages of prototype-learning and decision tree. ProtoTree-MIL not only explains why some patches are important for the final prediction through prototype-learning, but also provides global and local explanation through decision tree. We also propose a method to infer patch probabilities and measure their importance under the framework of ProtoTree-MIL. By conducting various experiments on three public WSI datasets, Camelyon16, TCGA-NSCLC, and TCGA-RCC, we demonstrate that our proposed ProtoTree-MIL can achieve a competitive performance to the state-of-the-art MIL models but provide more persuasive explanations than them. Explicitly generating patch probabilities also makes ProtoTree-MIL more accurate to detect the key patches than other attention-based MIL models. Specially, by evaluating our model on a real clinical gastritis and gastric cancer dataset, we show the explanations provided by ProtoTree-MIL are significant and faithful. Zhifeng Wu, Luning Wang, Shendi Wang, Yufei Cui, Jiahai Wang |
IJCNN | 5 |
| 2023 | Faster and Stronger Lossless Compression with Optimized Autoregressive FrameworkabstractNeural AutoRegressive (AR) framework has been applied in general-purpose lossless compression recently to improve compression performance. However, this paper found that directly applying the original AR framework causes the duplicated processing problem and the in-batch distribution variation problem, which leads to deteriorated compression performance. The key to address the duplicated processing problem is to disentangle the processing of the history symbol set at the input side. Two new types of neural blocks are first proposed. An individual-block performs separate feature extraction on each history symbol while a mix-block models the correlation between extracted features and estimates the probability. A progressive AR-based compression framework (PAC) is then proposed, which only requires one history symbol from the host at a time rather than the whole history symbol set. In addition, we introduced a trainable matrix multiplication to model the ordered importance, replacing previous hardware-unfriendly Gumble-Softmax sampling. The in-batch distribution variation problem is caused by AR-based compression’s structured batch construction. Based on this observation, a batch-location-aware individual block is proposed to capture the heterogeneous in-batch distributions precisely, improving the performance without efficiency losses. Experimental results show the proposed framework can achieve an average of 130% speed improvement with an average of 3% compression ratio gain across data domains compared to the state-of-the-art. Yu Mao 0001, Jingzong Li, Yufei Cui, Chun Jason Xue |
DAC | 3 |
| 2023 | Bayes-MIL: A New Probabilistic Perspective on Attention-based Multiple Instance Learning for Whole Slide Images
Yufei Cui, Ziquan Liu, Xue (Steve) Liu, Cong Wang 0001, Tei-Wei Kuo, Chun Jason Xue, Antoni B. Chan |
ICLR | 1 |
| 2023 | Retrieval-Augmented Multiple Instance LearningabstractMultiple Instance Learning (MIL) is a crucial weakly supervised learning method applied across various domains, e.g., medical diagnosis based on whole slide images (WSIs). Recent advancements in MIL algorithms have yielded exceptional performance when the training and test data originate from the same domain, such as WSIs obtained from the same hospital. However, this paper reveals a performance deterioration of MIL models when tested on an out-of-domain test set, exemplified by WSIs sourced from a novel hospital. To address this challenge, this paper introduces the Retrieval-AugMented MIL (RAM-MIL) framework, which integrates Optimal Transport (OT) as the distance metric for nearest neighbor retrieval. The development of RAM-MIL is driven by two key insights. First, a theoretical discovery indicates that reducing the input's intrinsic dimension can minimize the approximation error in attention-based MIL. Second, previous studies highlight a link between input intrinsic dimension and the feature merging process with the retrieved data. Empirical evaluations conducted on WSI classification demonstrate that the proposed RAM-MIL framework achieves state-of-the-art performance in both in-domain scenarios, where the training and retrieval data are in the same domain, and more crucially, in out-of-domain scenarios, where the (unlabeled) retrieval data originates from a different domain. Furthermore, the use of the transportation matrix derived from OT renders the retrieval results interpretable at the instance level, in contrast to the vanilla $l_2$ distance, and allows for visualization for human experts. *Code can be found at \url{https://github.com/ralphc1212/ram-mil*. Yufei Cui, Ziquan Liu, Xue (Steve) Liu, Tei-Wei Kuo, Miguel R. D. Rodrigues, Chun Jason Xue, Antoni B. Chan |
NeurIPS | 1 |
| 2023 | Variational Nested DropoutabstractNested dropout is a variant of dropout operation that is able to order network parameters or features based on the pre-defined importance during training. It has been explored for: I. Constructing nested nets Cui et al. 2020, Cui et al. 2021: the nested nets are neural networks whose architectures can be adjusted instantly during testing time, e.g., based on computational constraints. The nested dropout implicitly ranks the network parameters, generating a set of sub-networks such that any smaller sub-network forms the basis of a larger one. II. Learning ordered representation Rippel et al. 2014: the nested dropout applied to the latent representation of a generative model (e.g., auto-encoder) ranks the features, enforcing explicit order of the dense representation over dimensions. However, the dropout rate is fixed as a hyper-parameter during the whole training process. For nested nets, when network parameters are removed, the performance decays in a human-specified trajectory rather than in a trajectory learned from data. For generative models, the importance of features is specified as a constant vector, restraining the flexibility of representation learning. To address the problem, we focus on the probabilistic counterpart of the nested dropout. We propose a variational nested dropout (VND) operation that draws samples of multi-dimensional ordered masks at a low cost, providing useful gradients to the parameters of nested dropout. Based on this approach, we design a Bayesian nested neural network that learns the order knowledge of the parameter distributions. We further exploit the VND under different generative models for learning ordered latent distributions. In experiments, we show that the proposed approach outperforms the nested network in terms of accuracy, calibration, and out-of-domain detection in classification tasks. It also outperforms the related generative models on data generation tasks. Yufei Cui, Yu Mao 0001, Ziquan Liu, Qiao Li 0001, Antoni B. Chan, Xue (Steve) Liu, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | CacheSifter: Sifting Cache Files for Boosted Mobile Performance and Lifetime
Yu Liang 0004, Riwei Pan, Yufei Cui, Rachata Ausavarungnirun, Xianzhang Chen, Changlong Li 0006, Tei-Wei Kuo, Chun Jason Xue |
FAST | 4 |
| 2022 | Accelerating General-purpose Lossless Compression via Simple and Scalable ParameterizationabstractThe storage of multi-media data can benefit from the advancements in general-purpose lossless compression. The explosive growth of multi-media data volume in data centers demands a higher compression ratio and better compressors' run-time speed. However, recent deep-learning-based compressors with a high compression ratio usually build complicated dependencies on history symbols, leading to a long compression time. This paper investigates the behavior of historical symbols and finds an approximate order of importance. Namely, recent symbols have a substantially larger influence on the probability estimation of the next unknown symbol. This observation guides the designing of an interpretable structure for data compression, rather than learning implicitly from data like Recurrent Neural Network (RNN) and attention. Based on this observation, we disentangle the compression model into order learning and feature learning, which were fused in a large module in previous works. A parameterized ordered mask unit is established to learn the ordered importance of history symbols. A fast Multi-Layer Perceptron (MLP) network is designed for efficient feature learning. The proposed compressor can improve both compression performance and computational efficiency compared with transformer-based or RNN-based compressors. To further enhance computational efficiency, we propose a branch-MLP block to replace the original MLP layer. This block reduces the parameters and the FLOPs of the original MLP to a half, without sacrificing compression performance. Experiments on multi-media data demonstrate that our model improves the compression ratio by 10% on average across data domains while accelerating compression speed by 100% compared with the state-of-the-art. The source code and appendix are released at https://github.com/mynotwo/compressor_via_simple_and_scalable_parameterization.git. Yu Mao 0001, Yufei Cui, Tei-Wei Kuo, Chun Jason Xue |
ACM Multimedia | 2 |
| 2022 | TRACE: A Fast Transformer-based General-Purpose Lossless CompressorabstractDeep-learning-based compressor has received interests recently due to much improved compression ratio. However, modern approaches suffer from long execution time. To ease this problem, this paper targets on cutting down the execution time of deep-learning-based compressors. Building history-dependencies sequentially (e.g., recurrent neural networks) is responsible for long inference latency. Instead, we introduce transformer into deep learning compressors to build history-dependencies in parallel. However, existing transformer is too heavy in computation and incompatible to compression tasks. Yu Mao 0001, Yufei Cui, Tei-Wei Kuo, Chun Jason Xue |
WWW | 2 |
| 2022 | NFL: Robust Learned Index via Distribution TransformationabstractRecent works on learned index open a new direction for the indexing field. The key insight of the learned index is to approximate the mapping between keys and positions with piece-wise linear functions. Such methods require partitioning key space for a better approximation. Although lots of heuristics are proposed to improve the approximation quality, the bottleneck is that the segmentation overheads could hinder the overall performance. This paper tackles the approximation problem by applying a distribution transformation to the keys before constructing the learned index. A two-stage Normalizing-Flow-based Learned index framework (NFL) is proposed, which first transforms the original complex key distribution into a near-uniform distribution, then builds a learned index leveraging the transformed keys. For effective distribution transformation, we propose a Numerical Normalizing Flow (Numerical NF). Based on the characteristics of the transformed keys, we propose a robust After-Flow Learned Index (AFLI). To validate the performance, comprehensive evaluations are conducted on both synthetic and real-world workloads, which shows that the proposed NFL produces the highest throughput and the lowest tail latency compared to the state-of-the-art learned indexes. Shangyu Wu, Yufei Cui, Jinghuan Yu, Xuan Sun 0003, Tei-Wei Kuo, Chun Jason Xue |
Proc. VLDB Endow. | 2 |
| 2022 | Online Rare Category Identification and Data Diversification for Edge ComputingabstractIdentifying rare categories is an important data management problem in many application fields, including video surveillance, ecological environment monitoring, and precision medicine. Previously, solutions in the literature require all data instances to be first delivered to the server. Then, the rare category identification algorithms are executed on the pool of data to find informative instances for human annotators to label. This incurs large bandwidth consumption and high latency. To deal with the problems, we propose a lightweight rare category identification framework. At the sensor side, the designed online algorithm filters less informative data instances from the data stream and only sends the informative ones to the servers for annotating. After labeling, the server only sends labels of the corresponding data instances in response. The sensor-side algorithm is extended to enable cooperation between embedded devices for the cases that data are collected in a distributed manner. For enhancing diversity of selected data, a representative selection algorithm is proposed to run during the idle time of the system or after the execution of a rare category identification algorithm. Experiments are conducted to show that our framework dramatically outperforms the baseline. The network traffic is reduced by 75% on average. Yufei Cui, Qiao Li 0001, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Bits-Ensemble: Toward Light-Weight Robust Deep Ensemble by Bits-SharingabstractRobustness and uncertainty estimation is crucial to the safety of deep neural networks (DNNs) deployed on the edge. The deep ensemble model, composed of a set of individual DNNs (namely members), has strong performance in accuracy, uncertainty estimation, and robustness to out-of-distribution data and adversarial attacks. However, the storage and memory consumption increases linearly with the number of members within an ensemble. Previous works focus on selecting better members, layer-wise low-rank approximation of ensemble parameters, and designing partial ensemble model for reducing the ensemble size, thus lowering storage and memory consumption. In this work, we pay attention to the quantization of the ensemble, which serves as the last mile of network deployment. We propose a differentiable and parallelizable bit sharing scheme that allows the members to share the less significant bits of parameters, without hurting the performance, leaving alone the more significant bits. The intuition is that, numerically, more significant bits (e.g., the bit for the sign) are more useful in distinguishing a member from other members. For real deployment of the bit-sharing scheme, we further propose an efficient encoding-decoding scheme with minimal storage overhead. The experimental results show that, BitsEnsemble reduces the storage size of ensemble for over$22\times $, with only$0.36\times $increase in training latency, and no sacrifice of inference latency. The code is available inhttps://github.com/ralphc1212/bitsensemble. Yufei Cui, Shangyu Wu, Qiao Li 0001, Antoni B. Chan, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Resolving the Reliability Issues of Open Blocks for 3-D NAND Flash: Observations and StrategiesabstractWhile the block size of 3-D NAND flash memory increases with the density and capacity, the raw bit-error rates (RBER) of open blocks could be significantly increased. This article conducts a systematic study over reliability issues caused by open blocks, and reports several new observations. We found that the reliability degradation, due to long open time in writing a block, could happen over all layers in a 3-D NAND block, even after the block is closed. To address the reliability issues of open blocks, this article first proposes to adaptively allocate active blocks to serve write requests based on the workload characteristics for open time reduction. We then propose a partial-block refreshing strategy to alleviate the amplified RBER variations in open blocks and, thus, avoid unnecessary refreshing operations in low-RBER layers. Experimental results show that the proposed method can reduce the RBER by 43% through the reduction of the open time by 28% on average, and reduce the extra write operations for refreshing by 23% on average. Qiao Li 0001, Yufei Cui, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Accelerating Monte Carlo Bayesian Prediction via Approximating Predictive Uncertainty Over the SimplexabstractEstimating the predictive uncertainty of a Bayesian learning model is critical in various decision-making problems, e.g., reinforcement learning, detecting the adversarial attack, self-driving car. As the model posterior is almost always intractable, most efforts were made on finding an accurate approximation to the true posterior. Even though a decent estimation of the model posterior is obtained, another approximation is required to compute the predictive distribution over the desired output. A common accurate solution is to use Monte Carlo (MC) integration. However, it needs to maintain a large number of samples, and evaluate the model repeatedly, and average multiple model outputs. In many real-world cases, this is computationally prohibitive. In this work, assuming that the exact posterior or a decent approximation is obtained, we propose a generic framework to approximate the output probability distribution induced by the model posterior with a parameterized model and in an amortized fashion. The aim is to approximate the predictive uncertainty of a specific Bayesian model, meanwhile alleviating the heavy workload of MC integration at testing time. The proposed method is universally applicable to Bayesian classification models that allow for posterior sampling. Theoretically, we show that the idea of amortization incurs no additional costs on approximation performance. Empirical results validate the strong practical performance of our approach. Yufei Cui, Wuguannan Yao, Qiao Li 0001, Antoni B. Chan, Chun Jason Xue |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Bayesian Nested Neural Networks for Uncertainty Calibration and Adaptive CompressionabstractNested networks or slimmable networks are neural networks whose architectures can be adjusted instantly during testing time, e.g., based on computational constraints. Recent studies have focused on a "nested dropout" layer, which is able to order the nodes of a layer by importance during training, thus generating a nested set of subnetworks that are optimal for different configurations of resources. However, the dropout rate is fixed as a hyperparameter over different layers during the whole training process. Therefore, when nodes are removed, the performance decays in a human-specified trajectory rather than in a trajectory learned from data. Another drawback is the generated sub-networks are deterministic networks without well-calibrated uncertainty. To address these two problems, we develop a Bayesian approach to nested neural networks. We propose a variational ordering unit that draws samples for nested dropout at a low cost, from a proposed Downhill distribution, which provides useful gradients to the parameters of nested dropout. Based on this approach, we design a Bayesian nested neural network that learns the order knowledge of the node distributions. In experiments, we show that the proposed approach outperforms the nested network in terms of accuracy, calibration, and out-of-domain detection in classification tasks. It also outperforms the related approach on uncertainty-critical tasks in computer vision. Yufei Cui, Ziquan Liu, Qiao Li 0001, Antoni B. Chan, Chun Jason Xue |
CVPR | 1 |
| 2020 | Fully Nested Neural Network for Adaptive Compression and QuantizationabstractNeural network compression and quantization are important tasks for fitting state-of-the-art models into the computational, memory and power constraints of mobile devices and embedded hardware. Recent approaches to model compression/quantization are based on reinforcement learning or search methods to quantize the neural network for a specific hardware platform. However, these methods require multiple runs to compress/quantize the same base neural network to different hardware setups. In this work, we propose a fully nested neural network (FN3) that runs only once to build a nested set of compressed/quantized models, which is optimal for different resource constraints. Specifically, we exploit the additive characteristic in different levels of building blocks in neural network and propose an ordered dropout (ODO) operation that ranks the building blocks. Given a trained FN3, a fast heuristic search algorithm is run offline to find the optimal removal of components to maximize the accuracy under different constraints. Compared with the related works on adaptive neural network designed only for channels or bits, the proposed approach is applicable to different levels of building blocks (bits, neurons, channels, residual paths and layers). Empirical results validate strong practical performance of proposed approach. Yufei Cui, Ziquan Liu, Wuguannan Yao, Qiao Li 0001, Antoni B. Chan, Tei-Wei Kuo, Chun Jason Xue |
IJCAI | 1 |
| 2020 | Shaving Retries with Sentinels for Fast Read over High-Density 3D FlashabstractHigh-density flash-memory chips are under tremendous demands with the exponential growth of data. At the same time, the slow read performance of these high-density flash-memory chips becomes a new challenge. In this work, we analyze the high raw bit error rates (RBER) issue by characterizing the error behaviours of 3D QLC flash-memory chips. A preferred read voltage to a QLC cell could vary among layers and might even change in a short period of time due to the temperature. A sentinel-cell approach is thus proposed to utilize the error characteristics among cells. We propose to infer the optimal read voltages of a wordline based on errors introduced on sentinel cells. An on-line calibration procedure is further presented to resolve the problem of possible non-uniform error distribution on some wordlines. With optimal voltages being inferred, the number of read retries will be significantly reduced. Experiments show that optimal read voltages can be instantly obtained in 94% cases on average over the evaluated QLC flash memory with at most 2 read retries, and with merely 0.2% space overheads for adopting sentinel cells. The number of read retries could be reduced by 82% on average, and the read performance can be improved by 74% on average through a series of extensive experiments over 3D TLC and QLC flash-memory chips. Qiao Li 0001, Yufei Cui, Liang Shi 0001, Tei-Wei Kuo, Chun Jason Xue |
MICRO | 3 |
| 2020 | Exploiting Asymmetric Errors for LDPC Decoding Optimization on 3D NAND Flash MemoryabstractBy stacking layers vertically, the adoption of 3D NAND has significantly increased the capacity for storage systems. The complex structure of 3D NAND introduces more errors than planer flash. To address the reliability issue, low-density parity-check (LDPC) code with a strong error correction capability is now widely applied on 3D NAND flash memory. However, LDPC has long decoding latency when the raw bit error rates (RBER) are high. This is because it needs fine-grained soft sensing between voltage states to iteratively decode the raw data. Multiple sensing voltages are applied on flash cell array to gain necessary information for decoding. In this article, a new sensing level placement scheme with reduced number of sensing levels is proposed. The basic idea for the placement scheme is motivated by three asymmetric error characteristics of flash memory: the asymmetric errors between different states, the asymmetric errors caused by voltage left-shifts and right-shifts and asymmetric errors among layers in a 3D NAND flash block. With awareness of these three types of error characteristics, reduced number of sensing levels are placed to achieve reduced read latency for LDPC decoding while maintaining the error correction capability of LDPC. Experiment analysis shows that the proposed scheme achieves significant performance improvement. Qiao Li 0001, Liang Shi 0001, Yufei Cui, Chun Jason Xue |
IEEE Trans. Computers | 3 |
| 2020 | Pruning Deep Reinforcement Learning for Dual User Experience and Storage Lifetime Improvement on Mobile DevicesabstractBackground segment cleaning in log-structured file system has a significant impact on mobile devices. A low triggering frequency of the cleaning activity cannot reclaim enough free space for subsequent I/O, thus incurring foreground segment cleaning and impacting the user experience. In contrast, a high triggering frequency could generate excessive block migrations (BMs) and impair the storage lifetime. Prior works address this issue either by performance-biased solutions or incurring excessive memory overhead. In this article, a pruned reinforcement learning-based approach, MOBC, is proposed. Through learning the behaviors of I/O workloads and the statuses of logical address space, MOBC adaptively reduces the number of BMs and the number of triggered foreground segment cleanings. In order to integrate MOBC to resource-constraint mobile devices, a structured pruning method is proposed to reduce the time and space cost. The experimental results show that the pruned MOBC can reduce the worst case latency by 32.5%-68.6% at the 99.9th percentile, and improve the storage endurance by 24.3% over existing approaches, with significantly reduced overheads. Chao Wu 0006, Yufei Cui, Cheng Ji 0002, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Online Rare Category Detection for Edge ComputingabstractIdentifying rare categories is an important data management problem in many application fields including video surveillance, ecological environment monitoring and precision medicine. Previous solutions in literature require all data instances to be first delivered to the server. Then, the rare categories identification algorithms are executed on the pool of data to find informative instances for human annotators to label. This incurs large bandwidth consumption and high latency. To deal with the problems, we propose a light-weight rare categories identification framework. At the sensor side, the designed online algorithm filters less informative data instances from the data stream and only sends the informative ones to human annotators. After labeling, the server only sends labels of the corresponding data instances in response. The sensor-side algorithm is extended to enable cooperation between embedded devices for the cases that data is collected in a distributed manner. Experiments are conducted to show our framework dramatically outperforms the baseline. The network traffic is reduced by 75% on average. Yufei Cui, Qiao Li 0001, Sarana Nutanong, Chun Jason Xue |
DATE | 1 |
| 2019 | Sentinel Cells Enabled Fast Read for NAND Flash
Qiao Li 0001, Yufei Cui, Liang Shi 0001, Chun Jason Xue |
HotStorage | 3 |
| 2019 | A Hardware-Accelerated Solution for Hierarchical Index-Based Merge-Join(Extended Abstract)abstractHardware acceleration through field programmable gate arrays (FPGAs) has recently become a technique of growing interest for many data-intensive applications. Join query is one of the most fundamental database query types useful in relational database management systems. However, the available solutions so far have been beset by higher costs in comparison with other query types. In this paper, we develop a novel solution to accelerate the processing of sort-merge join queries with low match rates. Specifically, our solution makes use of hierarchical indexes to identify result-yielding regions in the solution space in order to take advantage of result sparseness. Further, in addition to one-dimensional equi-join query processing, our solution supports processing of multidimensional similarity join queries. Experimental results show that our solution is superior to the best existing method in a low match rate setting; the method achieves a speedup factor of 4.8 for join queries with a match rate of 5%. Zimeng Zhou, Chenyun Yu, Sarana Nutanong, Yufei Cui, Chenchen Fu, Chun Jason Xue |
ICDE | 4 |
| 2019 | A Hardware-Accelerated Solution for Hierarchical Index-Based Merge-JoinabstractHardware acceleration through field programmable gate arrays (FPGAs) has recently become a technique of growing interest for many data-intensive applications. Join query is one of the most fundamental database query types useful in relational database management systems. However, the available solutions so far have been beset by higher costs in comparison to other query types. In this paper, we develop a novel solution to accelerate the processing of sort-merge join queries with low match rates. Specifically, our solution makes use of hierarchical indexes to identify result-yielding regions in the solution space in order to take advantage of result sparseness. Further, in addition to one-dimensional equi-join query processing, our solution supports processing of multidimensional similarity join queries. Experimental results show that our solution is superior to the best existing method in a low match rate setting; the method achieves a speedup factor of 4.8 for join queries with a match rate of 5 percent. Zimeng Zhou, Chenyun Yu, Sarana Nutanong, Yufei Cui, Chenchen Fu, Chun Jason Xue |
IEEE Trans. Knowl. Data Eng. | 4 |