EDBT 2026 Demo / reviewers in the wild / expert
Yijiang Liu
dblp:242/7663
· DBLP profile ↗
17ranked-venue papers
8as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PAT: Pruning-Aware Tuning for Large Language ModelsabstractLarge language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet, traditional post-hoc pruning often leads to significant performance loss, with limited recovery from further fine-tuning due to reduced capacity. Since the model fine-tuning refines the general and chaotic knowledge in pre-trained models, we aim to incorporate structural pruning with the fine-tuning, and propose the Pruning-Aware Tuning (PAT) paradigm to eliminate model redundancy while preserving the model performance to the maximum extend. Specifically, we insert the innovative Hybrid Sparsification Modules (HSMs) between the Attention and FFN components to accordingly sparsify the upstream and downstream linear modules. The HSM comprises a lightweight operator and a globally shared trainable mask. The lightweight operator maintains a training overhead comparable to that of LoRA, while the trainable mask unifies the channels to be sparsified, ensuring structural pruning. Additionally, we propose the Identity Loss which decouples the transformation and scaling properties of the HSMs to enhance training robustness. Extensive experiments demonstrate that PAT excels in both performance and efficiency. For example, our Llama2-7b model with a 25% pruning ratio achieves 1.33x speedup while outperforming the LoRA-finetuned model by up to 1.26% in accuracy with a similar training cost. Yijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang, Yuan Du |
AAAI | 1 |
| 2025 | FBQuant: FeedBack Quantization for Large Language ModelsabstractDeploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user privacy. However, on-device deployment is challenging due to the limited computational resources of edge devices. In particular, the key bottleneck stems from memory bandwidth constraints related to weight loading. Weight-only quantization effectively reduces memory access, yet often induces significant accuracy degradation. Recent efforts to incorporate sub-branches have shown promise for mitigating quantization errors, but these methods either lack robust optimization strategies or rely on suboptimal objectives. To address these gaps, we propose FeedBack Quantization (FBQuant), a novel approach inspired by negative feedback mechanisms in automatic control. FBQuant inherently ensures that the reconstructed weights remain bounded by the quantization process, thereby reducing the risk of overfitting. To further offset the additional latency introduced by sub-branches, we develop an efficient CUDA kernel that decreases 60% of extra inference time. Comprehensive experiments demonstrate the efficiency and effectiveness of FBQuant across various LLMs. Notably, for 3-bit Llama2-7B, FBQuant improves zero-shot accuracy by 1.2%. Yijiang Liu, Hengyu Fang, Liulu He, Rongyu Zhang, Yichuan Bai, Yuan Du |
IJCAI | 1 |
| 2024 | Cloud-Device Collaborative Learning for Multimodal Large Language ModelsabstractThe burgeoning field of Multimodal Large Language Models (MLLMs) has exhibited remarkable performance in diverse tasks such as captioning, commonsense reasoning, and visual scene understanding. However, the deployment of these large-scale MLLMs on client devices is hindered by their extensive model parameters, leading to a notable de-cline in generalization capabilities when these models are compressed for device deployment. Addressing this chal-lenge, we introduce a Cloud-Device Collaborative Contin-ual Adaptation framework, designed to enhance the performance of compressed, device-deployed MLLMs by lever-aging the robust capabilities of cloud-based, larger-scale MLLMs. Our framework is structured into three key components: a device-to-cloud uplink for efficient data transmission, cloud-based knowledge adaptation, and an optimized cloud-to-device downlink for model deployment. In the up-link phase, we employ an Uncertainty-guided Token Sam-pling (UTS) strategy to effectively filter out-of-distribution tokens, thereby reducing transmission costs and improving training efficiency. On the cloud side, we propose Adapter-based Knowledge Distillation (AKD) method to transfer refined knowledge from large-scale to compressed, pocket-size MLLMs. Furthermore, we propose a Dynamic Weight update Compression (DWC) strategy for the down-link, which adaptively selects and quantizes updated weight parameters, enhancing transmission efficiency and reducing the representational disparity between cloud and de-vice models. Extensive experiments on several multimodal benchmarks demonstrate the superiority of our proposed framework over prior Knowledge Distillation and device-cloud collaboration methods. Notably, we also validate the feasibility of our approach to real-world experiments. Guanqun Wang, Jiaming Liu 0003, Chenxuan Li 0003, Yuan Zhang 0020, Junpeng Ma, Maurice Chong, Renrui Zhang, Yijiang Liu, Shanghang Zhang |
CVPR | 10 |
| 2024 | PromptCoT: Align Prompt Distribution via Adapted Chain-of-ThoughtabstractDiffusion-based generative models have exhibited remarkable capability in the production of high-fidelity visual content such as images and videos. However, their performance is significantly contingent upon the quality of textual inputs, commonly referred to as ‘'prompts'. The process of traditional prompt engineering necessitates empirical exper-tise and poses challenges for inexperienced users. In this paper, we introduce PromptCoT, an innovative enhancer that autonomously refines prompts for users. PromptCoT is designed based on the observation that prompts, which re-semble the textual information of high-quality images during training, lead to superior generation performance. Therefore, we fine-tune the Large Language Models (LLM) using a curated text dataset that comprises descriptions of high-quality visual content. Consequently, the LLM can capture the distribution of high-quality texts, enabling it to boost the original texts. Nonetheless, one drawback of LLMs is their tendency to generate irrelevant information. We employ a tailored Chain-of-Thought (CoT) mechanism to address the problem. Our CoT can extract and amalgamate crucial information from the prompt candidates, enabling a reasonable process based on the contextual cues to produce a more comprehensive and nuanced output. Considering computational efficiency, instead of allocating a dedicated LLM to each individual model or dataset, we integrate adapters that facil-itate task-specific adaptation, leveraging a shared LLM as the foundation for this process. With independent fine-tuning of adapters, we can adapt PromptCoT to new datasets while minimally increasing training costs and memory usage. We evaluate the effectiveness of PromptCoT by assessing on widely-used latent diffusion models for visual generation. The results demonstrate significant improvements in key performance metrics. Junyi Yao, Yijiang Liu, Zhen Dong 0003, Mingfei Guo, Helan Hu, Kurt Keutzer, Daquan Zhou, Shanghang Zhang |
CVPR | 2 |
| 2024 | Improving Cross-lingual Aspect-based Sentiment Analysis with Sememe BridgeabstractAspect-based Sentiment Analysis (ABSA) comprises numerous subtasks including aspect term extraction (AE), opinion term extraction (OE), opinion pair extraction (PE), and triplet extraction (TE). Current research in Chinese ABSA primarily concentrates on aspect terms and sentiment polarity, with insufficient emphasis on opinion terms. This article aims to provide a viable solution for the unannotated Chinese ABSA subtasks such as OE, PE, and TE. First, we develop an English-Chinese parallel dataset for ABSA using a semi-automatic process involving machine translation and a word aligner. Second, we examine the efficacy of cross-lingual transfer methods. Third, we propose a plug-and-play transfer method based on sememe knowledge. Sememes are the language-independent smallest semantic units that encapsulate the components, commonalities, and attributes of things extracted from the real world, which can bridge the gap between English and Chinese. Experimental results show that our proposed method brings significant improvements for Chinese ABSA, and achieves a maximum increase of 8% on the F1 metric for model transfer and label transfer on OE, PE, and TE. Yijiang Liu, Fei Li 0021, Donghong Ji |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2023 | NoisyQuant: Noisy Bias-Enhanced Post-Training Activation Quantization for Vision TransformersabstractThe complicated architecture and high training cost of vision transformers urge the exploration of post-training quantization. However, the heavy-tailed distribution of vision transformer activations hinders the effectiveness of previous post-training quantization methods, even with advanced quantizer designs. Instead of tuning the quantizer to better fit the complicated activation distribution, this paper proposes NoisyQuant, a quantizer-agnostic enhancement for the post-training activation quantization performance of vision transformers. We make a surprising theoretical discovery that for a given quantizer, adding a fixed Uniform noisy bias to the values being quantized can significantly reduce the quantization error under provable conditions. Building on the theoretical insight, NoisyQuant achieves the first success on actively altering the heavy-tailed activation distribution with additive noisy bias to fit a given quantizer. Extensive experiments show NoisyQuant largely improves the post-training quantization performance of vision transformer with minimal computation overhead. For instance, on linear uniform 6-bit activation quantization, NoisyQuant improves SOTA top-1 accuracy on ImageNet by up to 1.7%, 1.1% and 0.5% for ViT, DeiT, and Swin Transformer respectively, achieving on-par or even higher performance than previous nonlinear, mixed-precision quantization. Yijiang Liu, Huanrui Yang, Zhen Dong 0003, Kurt Keutzer, Shanghang Zhang |
CVPR | 1 |
| 2023 | Q-Diffusion: Quantizing Diffusion ModelsabstractDiffusion models have achieved great success in image synthesis through iterative noise estimation using deep neural networks. However, the slow inference, high memory consumption, and computation intensity of the noise estimation model hinder the efficient adoption of diffusion models. Although post-training quantization (PTQ) is considered a go-to compression method for other tasks, it does not work out-of-the-box on diffusion models. We propose a novel PTQ method specifically tailored towards the unique multi-timestep pipeline and model architecture of the diffusion models, which compresses the noise estimation network to accelerate the generation process. We identify the key difficulty of diffusion model quantization as the changing output distributions of noise estimation networks over multiple time steps and the bimodal activation distribution of the shortcut layers within the noise estimation network. We tackle these challenges with timestep-aware calibration and split shortcut quantization in this work. Experimental results show that our proposed method is able to quantize full-precision unconditional diffusion models into 4-bit while maintaining comparable performance (small FID change of at most 2.34 compared to >100 for traditional PTQ) in a training-free manner. Our approach can also be applied to text-guided image generation, where we can run stable diffusion in 4-bit weights with high generation quality for the first time. Xiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang, Zhen Dong 0003, Daniel Kang 0001, Shanghang Zhang, Kurt Keutzer |
ICCV | 2 |
| 2023 | MOIT: A Novel task for mining opinions towards implicit targets
Fei Li 0021, Chong Teng, Yijiang Liu, Chunli Xiang, Donghong Ji |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Mastering the Explicit Opinion-Role Interaction: Syntax-Aided Neural Transition System for Unified Opinion Role LabelingabstractUnified opinion role labeling (ORL) aims to detect all possible opinion structures of 'opinion-holder-target' in one shot, given a text. The existing transition-based unified method, unfortunately, is subject to longer opinion terms and fails to solve the term overlap issue. Current top performance has been achieved by employing the span-based graph model, which however still suffers from both high model complexity and insufficient interaction among opinions and roles. In this work, we investigate a novel solution by revisiting the transition architecture, and augmenting it with a pointer network (PointNet). The framework parses out all opinion structures in linear-time complexity, meanwhile breaks through the limitation of any length of terms with PointNet. To achieve the explicit opinion-role interactions, we further propose a unified dependency-opinion graph (UDOG), co-modeling the syntactic dependency structure and the partial opinion-role structure. We then devise a relation-centered graph aggregator (RCGA) to encode the multi-relational UDOG, where the resulting high-order representations are used to promote the predictions in the vanilla transition system. Our model achieves new state-of-the-art results on the MPQA benchmark. Analyses further demonstrate the superiority of our methods on both efficacy and efficiency. Shengqiong Wu, Hao Fei 0001, Fei Li 0021, Meishan Zhang, Yijiang Liu, Chong Teng, Donghong Ji |
AAAI | 5 |
| 2022 | Pair-wise aspect and opinion terms extraction as graph parsing via a novel mutually-aware interaction mechanism
Yijiang Liu, Fei Li 0021, Hao Fei 0001, Donghong Ji |
Neurocomputing | 1 |
| 2022 | Limb Pose Aware Networks for Monocular 3D Pose EstimationabstractIn the task of monocular 3D pose estimation, the estimation errors of limb joints (i.e., wrist, ankle, etc) with a higher degree of freedom(DOF) are larger than that of others (i.e., hip, thorax, etc). Specifically, errors may accumulate along the physiological structure of human body parts, and trajectories of joints with higher DOF bring in higher complexity. To address this problem, we propose a limb pose aware framework, involving a kinematic constraint aware network as well as a trajectory aware temporal module, to improve the 3D prediction accuracy of limb joint positions. Two kinematic constraints named relative bone angles and absolute bone angles are introduced in this paper, the former being used for building the angular relation between adjacent bones and the latter for building the angular relation between bones and the camera plane. As a joint result of two constraints, our work suppresses errors accumulated along limbs. Furthermore, we propose a trajectory-aware network, named as Hierarchical Transformer, which takes temporal trajectories of joints as input and generates fused trajectory estimation as a result. The Hierarchical Transformer consists of Transformer Encoder blocks and aims at improving the performance of fusing temporal features. Under the effect of kinematic constraints and trajectory network, we alleviate the problem of errors accumulated along limbs and achieve promising results. Most of the off-the-shelf 2D pose estimators can be easily integrated into our framework. We perform extensive experiments on public datasets and validate the effectiveness of the framework. The ablation studies show the strength of each individual sub-module. Lele Wu, Zhenbo Yu, Yijiang Liu, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Rethinking Boundaries: End-To-End Recognition of Discontinuous Mentions with Pointer NetworksabstractA majority of research interests in irregular (e.g., nested or discontinuous) named entity recognition (NER) have been paid on nested entities, while discontinuous entities received limited attention. Existing work for discontinuous NER, however, either suffers from decoding ambiguity or predicting using token-level local features. In this work, we present an innovative model for discontinuous NER based on pointer networks, where the pointer simultaneously decides whether a token at each decoding frame constitutes an entity mention and where the next constituent token is. Our model has three major merits compared with previous work: (1) The pointer mechanism is memory-augmented, which enhances the mention boundary detection and interactions between the current decision and prior recognized mentions. (2) The encoder-decoder architecture can linearize the complexity of structure prediction, and thus reduce search costs. (3) The model makes every decision using global information, i.e., by consulting all the input, encoder and previous decoder output in a global view. Experimental results on the CADEC and ShARe13 datasets show that our model outperforms flat and hypergraph models as well as a state-of-the-art transition-based model for discontinuous NER. Further in-depth analysis demonstrates that our model performs well in recognizing various entities including flat, overlapping and discontinuous ones. More crucially, our model is effective on boundary detection, which is the kernel source to NER. Hao Fei 0001, Donghong Ji, Bobo Li 0001, Yijiang Liu, Yafeng Ren, Fei Li 0021 |
AAAI | 4 |
| 2021 | Aspect-Based Pair-Wise Opinion Generation in Chinese automotive reviews: Design of the task, dataset and model
Yijiang Liu, Fei Li 0021, Donghong Ji |
Inf. Process. Manag. | 1 |
| 2021 | Document-level event causality identification via graph inference mechanism
Kun Zhao 0018, Donghong Ji, Fazhi He, Yijiang Liu, Yafeng Ren |
Inf. Sci. | 4 |
| 2020 | HiTrans: A Transformer-Based Context- and Speaker-Sensitive Model for Emotion Detection in ConversationsabstractEmotion detection in conversations (EDC) is to detect the emotion for each utterance in conversations that have multiple speakers. Different from the traditional non-conversational emotion detection, the model for EDC should be context-sensitive (e.g., understanding the whole conversation rather than one utterance) and speaker-sensitive (e.g., understanding which utterance belongs to which speaker). In this paper, we propose a transformer-based context- and speaker-sensitive model for EDC, namely HiTrans, which consists of two hierarchical transformers. We utilize BERT as the low-level transformer to generate local utterance representations, and feed them into another high-level transformer so that utterance representations could be sensitive to the global context of the conversation. Moreover, we exploit an auxiliary task to make our model speaker-sensitive, called pairwise utterance speaker verification (PUSV), which aims to classify whether two utterances belong to the same speaker. We evaluate our model on three benchmark datasets, namely EmoryNLP, MELD and IEMOCAP. Results show that our model outperforms previous state-of-the-art models. Donghong Ji, Fei Li 0021, Meishan Zhang, Yijiang Liu |
COLING | 5 |
| 2020 | End to End Chinese Lexical Fusion Recognition with Sememe KnowledgeabstractIn this paper, we present Chinese lexical fusion recognition, a new task which could be regarded as one kind of coreference recognition.First, we introduce the task in detail, showing the relationship with coreference recognition and differences from the existing tasks.Second, we propose an end-to-end model for the task, handling mentions as well as coreference relationship jointly.The model exploits the state-of-the-art contextualized BERT representations as the encoder, and is further enhanced with the sememe knowledge from HowNet by graph attention networks.We manually annotate a benchmark dataset for the task and then conduct experiments on it.Results demonstrate that our final model is effective and competitive for the task.Detailed analysis is offered for comprehensively understanding the new task and our proposed model. Yijiang Liu, Meishan Zhang, Donghong Ji |
COLING | 1 |
| 2019 | Identifying individual expectations in service recovery through natural language processing and machine learning
Yijiang Liu, Yinghong Wan |
Expert Syst. Appl. | 1 |