Yijin Liu

dblp:242/7766 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamical analysis and secure communication application of parameter-controlled multiscroll attractors in memristive chaotic system
Yijin Liu, Qiang Lai, Huangtao Wang, Yongxian Zhang
Integr.1
2025 LT-OAQ: Learnable Threshold Based Outlier-Aware Quantization and its Energy-Efficient Accelerator for Low-Precision On-Chip Training
abstract
Low-precision training has emerged as a powerful technique for reducing computational and storage costs in Deep Neural Network (DNN) model training, enabling on-chip training or fine-tuning on edge devices. However, existing low-precision training methods often require higher bit-widths to maintain accuracy as model sizes increase. In this paper, we introduce an outlier-aware quantization strategy for low-precision training. While traditional value-aware quantization methods require costly online distribution statistics operations on computational data, impeding the efficiency gains of low-precision training, our approach addresses this challenge through a novel Learnable Threshold based Outlier-Aware Quantization (LT-OAQ) training framework. This method concurrently updates outlier thresholds and model weights through gradient descent, eliminating the need for costly data-statistics operations. To efficiently support the LT-OAQ training framework, we designed a hardware accelerator based on the systolic array architecture. This accelerator introduces a processing element (PE) fusion mechanism that dynamically fuses adjacent PEs into clusters to support outlier computations, optimizing the mapping of outlier computation tasks, enabling mixed-precision training, and implementing online quantization. Our approach maintains model accuracy while significantly reducing computational complexity and storage resource requirements. Experimental results demonstrate that our design achieves a 2.9 ×speedup in performance and a 2.17 ×reduction in energy consumption compared to state-of-the-art low-precision accelerators.
Qinkai Xu, Yijin Liu, Yunlong Mao, Li Erran Li
DATE2
2025 VETA-DiT: Variance-Equalized and Temporally Adaptive Quantization for Efficient 4-bit Diffusion Transformers
abstract
Diffusion Transformers (DiTs) have recently demonstrated remarkable performance in visual generation tasks, surpassing traditional U-Net-based diffusion models by significantly improving image and video generation quality and scalability. However, the large model size and iterative denoising process introduce substantial computational and memory overhead, limiting their deployment in real-world applications. Post-training quantization (PTQ) is a promising solution that compresses models and accelerates inference by converting weights and activations to low-bit representations. Despite its potential, PTQ faces significant challenges when applied to DiTs, often resulting in severe degradation of generative quality. To address these issues, we propose VETA-DiT (**V**ariance-**E**qualized and **T**emporal **A**daptation for **Di**ffusion **T**ransformers), a dedicated quantization framework for DiTs. Our method first analyzes the sources of quantization error from the perspective of inter-channel variance and introduces a Karhunen–Loève Transform enhanced alignment to equalize variance across channels, facilitating effective quantization under low bit-widths. Furthermore, to handle the temporal variation of activation distributions inherent in the iterative denoising steps of DiTs, we design an incoherence-aware adaptive method that identifies and properly calibrates timesteps with high quantization difficulty. We validate VETA-DiT on extensive image and video generation tasks, preserving acceptable visual quality under the more aggressive W4A4 configuration. Specifically, VETA-DiT reduces FID by 33.65 on the DiT-XL/2 model and by 45.76 on the PixArt-$\Sigma$ model compared to the baseline under W4A4, demonstrating its strong quantization capability and generative performance. Code is available at: https://github.com/xululi0223/VETA-DiT.
Qinkai Xu, Yijin Liu, Lin Yang 0011, Li Erran Li
NeurIPS2
2025 A long-short-term feature extraction network based on soft-parameter-sharing for high-speed train bogies multi-object fault diagnosis under long-tailed distribution
Yijin Liu, Jinglong Chen, Tongyang Pan, Jingsong Xie
Expert Syst. Appl.1
2025 Generating Grid Multiscroll Memristive Chua's Circuit and Its Predefined-Time Synchronization for Secure Communication
abstract
With consideration of the inherent nonlinearity and distinctive memory characteristics, memristors are excellent candidates for constructing multiscroll attractors. This paper seeks to introduce a memristor into the Chua’s circuit to operate in conjunction with a novel piecewise nonlinear resistor which replaces the Chua’s diode to design the circuit that can generate grid multiscroll attractors, which is designated as multiscroll memristive Chua’s circuit (MMCC). The designed MMCC is capable of generating any number of scrolls, with the number of scrolls expanding in accordance with the internal variables of the memristor and the nonlinear resistor. By utilizing phase portraits, bifurcation diagrams, Lyapunov exponents (LEs), coexisting attractors and amplitude modulation thoroughly examined its property. The feasibility of the MMCC is demonstrated through circuit implementation. Furthermore, we design the predefined-time synchronization (PTS) controller for the MMCC, serving as the foundation for a multi-channel segmented secure communication scheme, whose effectiveness is rigorously validated through experimental testing.
Qiang Lai, Yijin Liu, Feng Liu 0011, Xiao-Wen Zhao
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Multi-Scale Fusion of Gated Neighborhood Attention Transformers for Single Image Deraining
abstract
Since the diverse geometric appearances and densities of rain streaks, local-global information is equally essential for single image deraining. Balancing local-global information becomes a challenge. Thus, a Multi-Scale Fusion of Gated Neighborhood Attention Transformers (MSF-GNAT) for single image deraining is proposed in this paper. Firstly, a Gated Neighborhood Attention Transformer (GNAT) block is designed to achieve complex condition deraining by complementarily fusing global and local information. Secondly, a Multi-Scale Fusion (MSF) block is developed to fuse multi-scale information for richer representation. Moreover, a Simplified Gated Feed-forward Network (SGFN) is proposed to ensure that the convolutional results primarily attend to valid pixels. Evaluation experiments on synthetic and real-world datasets demonstrate the superiority of our MSF-GNAT over state-of-the-art methods. The source code is available at https://github.com/SWU-CS-MediaLab/MSF-GNAT.
Yijin Liu, Guoqiang Xiao 0001, Michael S. Lew, Song Wu 0003
ICASSP1
2024 Language Generation with Strictly Proper Scoring Rules
abstract
Language generation based on maximum likelihood estimation (MLE) has become the fundamental approach for text generation. Maximum likelihood estimation is typically performed by minimizing the log-likelihood loss, also known as the logarithmic score in statistical decision theory. The logarithmic score is strictly proper in the sense that it encourages honest forecasts, where the expected score is maximized only when the model reports true probabilities. Although many strictly proper scoring rules exist, the logarithmic score is the only local scoring rule among them that depends exclusively on the probability of the observed sample, making it capable of handling the exponentially large sample space of natural text. In this work, we propose a straightforward strategy for adapting scoring rules to language generation, allowing for language modeling with any non-local scoring rules. Leveraging this strategy, we train language generation models using two classic strictly proper scoring rules, the Brier score and the Spherical score, as alternatives to the logarithmic score. Experimental results indicate that simply substituting the loss function, without adjusting other hyperparameters, can yield substantial improvements in model’s generation capabilities. Moreover, these improvements can scale up to large language models (LLMs) such as LLaMA-7B and LLaMA-13B. Source code: https://github.com/shaochenze/ScoringRulesLM.
Chenze Shao, Fandong Meng, Yijin Liu, Jie Zhou 0016
ICML3
2024 Dynamical Analysis and Fixed-Time Synchronization for Secure Communication of Hidden Multiscroll Memristive Chaotic System
abstract
In view of the superiority of memristors in strengthening dynamical complexity and the significant application potiential of multiscroll chaos, this paper attempts to introduce two memristors with scalable memductances into simple seed chaotic system for designing multiscroll memristive chaotic system (MMCS). The designed MMCS yields hidden grid multiscroll chaotic attractors with any number of scrolls expanding along with the internal variables of memristors. By varying the parameters, the multiscroll attractors can be broken into coexisting attractors with different numbers and scrolls dependent on parameters, and their oscillation amplitudes can be increased (or decreased) without changing the chaotic features. Dynamical analysis and circuit implementation are given to reveal the complexity and feasibility of the MMCS. The fixed-time synchronization (FxTS) is studied by using adaptive controller and the sufficient condition for FxTS is established via Lyapunov stability theory (LST). A multilevel secure communication scheme based on the FxTS of MMCS is designed and the experimental tests on the image, audio and data secure communication verify its effectiveness, which to some extent shows the application availability of MMCS.
Qiang Lai, Yijin Liu, Luigi Fortuna
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Subspace Modeling Enabled High-Sensitivity X-Ray Chemical Imaging
abstract
Resolving morphological chemical phase transformations at the nanoscale is of vital importance to many scientific and industrial applications across various disciplines. The TXM-XANES imaging technique, by combining full-field transmission X-ray microscopy (TXM) and X-ray absorption near edge structure (XANES), has been an emerging tool that operates by acquiring a series of microscopy images with multi-energy X-rays and fitting to obtain the chemical map. Its capability, however, is limited by the poor signal-to-noise ratios due to system errors and low exposure illuminations for fast acquisition. In this work, by exploiting the intrinsic properties and subspace modeling of the TXM-XANES imaging data, we introduce a simple and robust denoising approach to improve the image quality, which enables fast and high-sensitivity chemical characterization. Extensive experiments on both synthetic and real datasets demonstrate the superior performance of the proposed method.
Jizhou Li, Bin Chen 0019, Guibin Zan, Guannan Qian, Piero Pianetta, Yijin Liu
ICASSP6
2022 Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine Translation
abstract
Songming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, Jian Liu, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Songming Zhang 0001, Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032, Jie Zhou 0016
ACL (1)2
2021 Faster Depth-Adaptive Transformers
abstract
Depth-adaptive neural networks can dynamically adjust depths according to the hardness of input words, and thus improve efficiency. The main challenge is how to measure such hardness and decide the required depths (i.e., layers) to conduct. Previous works generally build a halting unit to decide whether the computation should continue or stop at each layer. As there is no specific supervision of depth selection, the halting unit may be under-optimized and inaccurate, which results in suboptimal and unstable performance when modeling sentences. In this paper, we get rid of the halting unit and estimate the required depths in advance, which yields a faster depth-adaptive model. Specifically, two approaches are proposed to explicitly measure the hardness of input words and estimate corresponding adaptive depth, namely 1) mutual information (MI) based estimation and 2) reconstruction loss based estimation. We conduct experiments on the text classification task with 24 datasets in various sizes and domains. Results confirm that our approaches can speed up the vanilla Transformer (up to 7x) while preserving high accuracy. Moreover, efficiency and robustness are significantly improved when compared with other depth-adaptive approaches.
Yijin Liu, Fandong Meng, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
AAAI1
2021 Prevent the Language Model from being Overconfident in Neural Machine Translation
abstract
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou 0016
ACL/IJCNLP (1)3
2021 Scheduled Sampling Based on Decoding Steps for Neural Machine Translation
abstract
Scheduled sampling is widely used to mitigate the exposure bias problem for neural machine translation.Its core motivation is to simulate the inference scene during training by replacing ground-truth tokens with predicted tokens, thus bridging the gap between training and inference.However, vanilla scheduled sampling is merely based on training steps and equally treats all decoding steps.Namely, it simulates an inference scene with uniform error rates, which disobeys the real inference scene, where larger decoding steps usually have higher error rates due to error accumulations.To alleviate the above discrepancy, we propose scheduled sampling methods based on decoding steps, increasing the selection chance of predicted tokens with the growth of decoding steps.Consequently, we can more realistically simulate the inference scene during training, thus better bridging the gap between training and inference.Moreover, we investigate scheduled sampling based on both training steps and decoding steps for further improvements.Experimentally, our approaches significantly outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks.Additionally, our approaches also generalize well to the text summarization task on two popular benchmarks.
Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
EMNLP (1)1
2019 GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling
abstract
Current state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs).However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information restrict the potential performance of those models.In this paper, we try to address these issues, and thus propose a Global Context enhanced Deep Transition architecture for sequence labeling named GCDT.We deepen the state transition path at each position in a sentence, and further assign every token with a global representation learned from the entire sentence.Experiments on two standard sequence labeling tasks show that, given only training data and the ubiquitous word embeddings (Glove), our GCDT achieves 91.96 F 1 on the CoNLL03 NER task and 95.43 F 1 on the CoNLL2000 Chunking task, which outperforms the best reported results under the same settings.Furthermore, by leveraging BERT as an additional resource, we establish new stateof-the-art results with 93.47 F 1 on NER and 97.30F 1 on Chunking 1 .
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)1
2019 CM-Net: A Novel Collaborative Memory Network for Spoken Language Understanding
abstract
Yijin Liu, Fandong Meng, Jinchao Zhang, Jie Zhou, Yufeng Chen, Jinan Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
EMNLP/IJCNLP (1)1