Jie Ou

dblp:37/3622 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-1159-2043ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GateRA: Token-aware Modulation for Parameter-Efficient Fine-tuning
abstract
Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, DoRA, and HiRA, enable lightweight adaptation of large pre-trained models via low-rank updates. However, existing PEFT approaches apply static, input-agnostic updates to all tokens, disregarding the varying importance and difficulty of different inputs. This uniform treatment can lead to overfitting on trivial content or under-adaptation on more informative regions, especially in autoregressive settings with distinct prefill and decoding dynamics. In this paper, we propose GateRA, a unified framework that introduces token-aware modulation to dynamically adjust the strength of PEFT updates. By incorporating adaptive gating into standard PEFT branches, GateRA enables selective, token-level adaptation—preserving pre-trained knowledge for well-modeled inputs while focusing capacity on challenging cases. Empirical visualizations reveal phase-sensitive behaviors, where GateRA automatically suppresses updates for redundant prefill tokens while emphasizing adaptation during decoding. To promote confident and efficient modulation, we further introduce an entropy-based regularization that encourages near-binary gating decisions. This regularization prevents diffuse update patterns and leads to interpretable, sparse adaptation without hard thresholding. Finally, we present a theoretical analysis showing that GateRA induces a soft gradient-masking effect over the PEFT path, enabling continuous and differentiable control over adaptation. Experiments on multiple commonsense reasoning benchmarks demonstrate that GateRA consistently outperforms or matches prior PEFT methods.
Jie Ou, Shuaihong Jiang, Yingjun Du, Cees Snoek
AAAI1
2026 AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache Reuse
abstract
Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Ruiqi Wu, Zhaokun Wang, Wenyi Li, Wenhong Tian. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Zhaokun Wang, Wenhong Tian
ACL (1)1
2026 CAP: Controllable Alignment Prompting for Unlearning in LLMs
abstract
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Xunlei Chen, Jie Ou, Guangchun Luo, Wenhong Tian
ACL (1)7
2026 Accelerating long-context inference of large language models via dynamic attention load balancing
Jie Ou, Jinyu Guo, Shuaihong Jiang, Ruini Xue, Wenhong Tian, Rajkumar Buyya
Knowl. Based Syst.1
2025 Efficient and Automatic 3D Parallelism Strategies Search via Contrastive Reinforcement Learning Pretrained Neural Networks
Jie Ou, Xiaowang Li, Yueming Chen, Jiahong Qian, Teng Su, Wenhong Tian
DSAA1
2025 Enhancing Large Language Model Inference Efficiency via Lookahead Cache Filtering
abstract
The large Key-Value (KV) cache is a significant challenge in deploying Large Language Models (LLMs). Current research addressing these issues employs cache compression techniques, which we find suffer from information loss and the "lost-in-the-middle" problem. We propose the Lookahead Cache Filtering (LCF), which retains the full cache in host memory to keep complete details, using the Important Key-Value Lookahead Prediction by Approximate Sorting and the Gather-based Matrix Multiplication on CPU, to filter important KV cache to reduce the overhead of loading cache to GPU and improve the throughput of inference. Finally, we use the Multi-Scale Pyramid Information Fusion mechanism to enhance information fusion to further improve the effectiveness of inference results. Experiments demonstrate that LCF effectively maintains accuracy without introducing additional latency while reducing memory requirements. Our code will be available on GitHub1.
Jie Ou, Yueming Chen, Shuaihong Jiang, Wenhong Tian
ICASSP1
2025 Can LLM Be a Good Path Planner Based on Prompt Engineering? Mitigating the Hallucination for Path Planning
Hourui Deng, Jie Ou, Chaosheng Feng
ICIC (23)3
2025 A Neural Network-Based Pipeline Parallel Strategy Solver for Heterogeneous Environments
abstract
The widespread application of large language models(LLMs) has made distributed training increasingly important, especially pipeline parallelism, which is a fundamental technique for ultra-large-scale LLMs. Current research in this field mainly employs combinatorial optimization algorithms such as dynamic programming. However, as the problem size increases, these methods become difficult to solve quickly in large-scale scenarios due to their high search time. Online optimization algorithms that combine neural networks with reinforcement learning require real-time interaction with the cluster environment to obtain feedback, resulting in high resource overhead and low search efficiency. Moreover, current research lacks studies on heterogeneous computing environments, which are frequently used by small research teams. To address these issues, we designed a novel Neural Network-based Pipeline Parallel strategy solver (NN-Piper) for heterogeneous environments. NN-Piper can perceive computational and communication costs, the number of stages to be divided, and the number of micro-batches. In addition, it can directly provide the strategy for allocating specific devices to each pipeline stage. To avoid an online training process that requires interaction with the cluster environment, we propose the Virtual Contrastive Training Algorithm (VCTA) to enable efficient training of NN-Piper without collecting large amounts of real data. After training, NN-Piper can be transferred to many different scenarios without further training or fine-tuning, and it can search for strategies within a few dozen milliseconds. Compared with the state-of-the-art method, NN-Piper can improve the training speed on average by 16-25% in different environments for the transformer-based models.
Jie Ou, Jinyu Guo, Yueming Chen, Shuaihong Jiang, Ruini Xue, Wenhong Tian
IJCNN1
2025 Low-Rank Decomposition Assisted Quantization and Inference Compensation for Quality Large Language Model Inference
abstract
Large Language Models (LLMs) have demonstrated exceptional performance on natural language processing tasks. However, these models are computationally intensive and require substantial hardware resources for deployment. Quantization has emerged as a popular technique for LLM deployment, reducing memory requirements, but it results in accuracy degradation, particularly when using low-bit quantization. To mitigate this accuracy loss, we introduce Low-Rank Compensation (LoRC), a novel compensation mechanism that aims to recover the performance drop caused by quantization. Additionally, we propose Low-Rank Quantization (LoRQ), which further reduces the quantization-induced loss by adaptively adjusting weights at the element-wise level to help LLMs accommodate quantized computations. LoRC focuses on compensating for accuracy loss during inference, LoRQ integrates low-rank compensation directly into the quantization process, and they do not need end-to-end fine-tuning with LLM. Furthermore, we propose the Rank-α Addition Strategy (RαAS) to combine LoRC into the inference framework, which improves inference accuracy without increasing inference latency. Experimental results show that our method outperforms the state-of-the-art OmniQuant by 1.89% on several common zero-shot datasets under the W4A4 setting of the widely-used LLaMA. Through the joint design of algorithms and systems, our techniques can be easily integrated into the FlexGen inference framework without introducing additional inference latency, thereby maintaining high throughput while improving accuracy.
Jie Ou, Jinyu Guo, Shuaihong Jiang, Zhaokun Wang, Yueming Chen, Ruini Xue, Wenhong Tian
IJCNN1
2025 Noise-Robustness Through Noise: A Framework combining Asymmetric LoRA with Poisoning MoE
abstract
Current parameter-efficient fine-tuning methods for adapting pre-trained language models to downstream tasks are susceptible to interference from noisy data. Conventional noise-handling approaches either rely on laborious data pre-processing or employ model architecture modifications prone to error accumulation. In contrast to existing noise-process paradigms, we propose a noise-robust adaptation method via asymmetric LoRA poisoning experts (LoPE), a novel framework that enhances model robustness to noise only with generated noisy data. Drawing inspiration from the mixture-of-experts architecture, LoPE strategically integrates a dedicated poisoning expert in an asymmetric LoRA configuration. Through a two-stage paradigm, LoPE performs noise injection on the poisoning expert during fine-tuning to enhance its noise discrimination and processing ability. During inference, we selectively mask the dedicated poisoning expert to leverage purified knowledge acquired by normal experts for noise-robust output. Extensive experiments demonstrate that LoPE achieves strong performance and robustness purely through the low-cost noise injection, which completely eliminates the requirement of data cleaning.
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Lingfeng Chen, Hongli Pu, Jie Ou, Libo Qin 0001, Wenhong Tian
NeurIPS6
2024 Improving Chinese Emotion Classification Based on Bilingual Feature Fusion
Haocheng Lan, Jie Ou, Zhaokun Wang, Wenhong Tian
ICPR (31)2
2024 A Supervised Domain Adaptation Method with Alignment Regularization for Low-Light Facial Expression Recognition
Zhaokun Wang, Yuanlun Xie, Jie Ou, Jiahui Zhong, Wenhong Tian
PRCV (3)3
2024 Compound facial expressions recognition approach using DCGAN and CNN
Jie Ou, Yuanlun Xie, Wenhong Tian
Multim. Tools Appl.2
2021 Use Machine Learning Based Smart Sampling to Improve System Level Testing Efficiency
abstract
System level tests (SLTs) are important and expensive procedures to ensure high IC quality. In volume production stage with stable high yields, efforts such as random sampling have been used to improve testing efficiency. However random sampling doesn’t fully utilize information gathered before SLT and is not optimal. In this paper we propose both supervised (SVM) and unsupervised (AutoEncoder) machine learning algorithms to predict or estimate SLT failures based on earlier stage Final Test (FT) test data and further use the estimated pseudo probabilities to guide the selecting of dies for system level testing. Experiments on a real product dataset, consisting of 158 wafers, each with 3118 FT testing variables reveal robustness of the models. Through the gains chart of the models, we provide a flexible smart sampling strategy and demonstrate its potential of reducing SLT testing cost by 40% with minor impact on Defective Parts Per Million (DPPM). Our cases also show that such smart sampling approach is very well suited for engaging adaptive test flow optimization achieving balanced goals of improving test efficiency, reducing cost and ensuring high product quality at the same time
Chenwei Liu, Jie Ou
ITC-Asia2
2021 Smart Sampling for Efficient System Level Test: A Robust Machine Learning Approach
abstract
System level tests (SLTs) are important and expensive procedures to ensure high quality of IC products. In the volume production stage with stable yield, efforts such as random sampling have been made to improve testing efficiency. However random sampling doesn’t fully utilize information gathered before SLT and is not optimal. In this paper we propose both supervised (SVM) and unsupervised (AutoEncoder ) machine learning algorithms to predict or estimate the SLT failure based on earlier stage Final Test (FT) data and use the estimated pseudo probabilities to guide the selection of some chips for system level test. Experiments on a real product dataset, consisting of 158 wafers from 8 lots, each with 3118 FT testing variables reveal robustness of the models to data shift such as lot variations and missing test items. Through the gains chart of the models, we provide a flexible smart sampling strategy and demonstrate its potential of reducing SLT testing cost by 40% with minor impact on Defective Parts Per Million (DPPM). Our cases also show that such robust machine learning based sampling approach is very well suited for engaging adaptive test flow optimization to achieve balanced goals of improving test efficiency, reducing cost and ensuring high product quality at the same time.
Chenwei Liu, Jie Ou
ITC2
2020 Full-resolution encoder-decoder networks with multi-scale feature fusion for human pose estimation
abstract
To achieve more accurate 2D human pose estimation, we extend the successful encoder-decoder network, simple baseline network (SBN), in three ways. To reduce the quantization errors caused by the large output stride size, two more decoder modules are appended to the end of the simple baseline network to get full output resolution. Then, the global context blocks (GCBs) are added to the encoder and decoder modules to enhance them with global context features. Furthermore, we propose a novel spatial-attention-based multi-scale feature collection and distribution module (SA-MFCD) to fuse and distribute multi-scale features to boost the pose estimation. Experimental results on the MS COCO dataset indicate that our network can remarkably improve the accuracy of human pose estimation over SBN, our network using ResNet34 as the backbone network can even achieve the same accuracy as SBN with ResNet152, and our networks can achieve superior results with big backbone networks.
Jie Ou
MMAsia1
2020 Efficient Human Pose Estimation with Depthwise Separable Convolution and Person Centroid Guided Joint Grouping
Jie Ou
PRCV (2)1