VLDB 2026 Research / reviewers in the wild / expert
Jiawang Liu
dblp:91/9436
· DBLP profile ↗
13ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Adaptation from Visible to Thermal Domain for Learned Image Compression
Jiawang Liu, Hualong Yu, Lu Yu 0003 |
ISCAS | 1 |
| 2025 | SiQA: A Large Multi-Modal Question Answering Model for Structured Images Based on RAGabstractExisting Large Multimodal Models (LMMs) demonstrate excellent performance in handling visual tasks in everyday scenarios. However, they still face challenges in understanding structured images, such as flowcharts and organizational charts, which are characterized by text-rich and complex hierarchical components. In this paper, we propose SiQA, a knowledge construction and Retrieval-Augmented Generation(RAG)-based multimodal Question-Answering model designed for Structured Images. SiQA operates in three stages: Knowledge Graph (KG) generation, retrieval-augmented, and answer generation. First, a KG representing the semantics of the structured images is generated through component analysis. We then performed similarity retrieval between the KG and queries, using a node-first algorithm to construct the most relevant subgraph. Finally, after performing an encoding alignment on the multimodal information, it is fed into the LLM to generate the answer. Additionally, we introduce a new dataset, OCQA1, which includes 5,112 questions derived from 1,000 Organizational Charts. We evaluated SiQA’s structured image detection and question-answering capabilities on the FD-DETR (a flowchart dataset) and SCQA, and verified its effectiveness and strong generalization ability through comparisons with existing state-of-the-art (SOTA) methods. Jiawang Liu, Ye Tao 0002, Fei Wang 0077, Hui Li 0010, Xiugong Qin |
ICASSP | 1 |
| 2025 | Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
Xiaolong Lin, Jiawang Liu, Shixi Huang, Zhenpeng Zhan |
INTERSPEECH | 3 |
| 2025 | Assessing the Reusability of Cloud-Received Feature Streams on Advanced NetworksabstractThe goal of this paper is to raise awareness of challenges and opportunities in the Collaborative Intelligence (CI) field and promote research on related standards. We begin by identifying a key challenge in CI applications, i.e., is it still possible for feature streams received in the cloud to be reused in the future by more advanced multitasking networks to achieve effective task accuracy? We then propose a framework to explore the generalization ability of cloud-received feature streams on more advanced networks from a coarse-grained to a fine-grained manner. We design a series of adapters of varying complexity to further explore the potential of feature streams for task network adaptation. Experiments show that sharing feature streams across multiple task networks could achieve an average of nearly 80% bitrate saving compared to Versatile Video Coding (VVC), which demonstrates the reuse potential of cloud-received feature streams. In addition, we make theoretical inferences about the adaptation range of shared feature streams, especially for those networks with high precision. Jiawang Liu, Hualong Yu, Heming Sun, Lu Yu 0003 |
ISCAS | 1 |
| 2025 | Learning-based Image Coding for Machine Intelligence with Variable-RateabstractImage Coding for Machines (ICM) has yielded significant developments recently. Variable-rate support is necessary for image coding, while performance gap still exists, in learning-based image coding, between the single-model and multiple-fixed-models methods. This paper proposes a Machine Intelligence Variable-Rate Codec (MIVRCodec) with single-model method. We introduce a method to generate, compress, and utilize image semantic feature information, enabling the codec to adaptively process different semantic content of the image. Additionally, current studies employ fixed methods to remove redundant information between luminance and chrominance components, neglecting the dynamic characteristics of this redundancy and leading to its inappropriate utilization. We further propose a Color Dynamic Fusion Module (CDFM), which adaptively fuses image color component features based on various conditions (e.g., bitrate and image content) to utilize the redundancy among image color components as appropriately as possible. Lastly, we propose a Progressive Training Strategy (PTS) for training MIVRCodec. These proposed methods not only reduce performance loss in variable-rate ICM but also improve baseline performance. Experimental results demonstrate that our proposed MIVRCodec works well in the bitrate range corresponding to meaningful accuracy intervals in machine intelligence tasks using a single model, achieving coding efficiency on par with multiple fixed-rate models and surpassing existing state-of-the-art codecs. Hualong Yu, Jiawang Liu, Qiqi He, Lu Yu 0003 |
ISCAS | 3 |
| 2025 | CMDF-TTS: Text-to-speech method with limited target speaker corpus
Jiawang Liu, Chaofeng Lu, Xiugong Qin, Yunlong Tian, Yongjie Du |
Neural Networks | 2 |
| 2024 | KFEX-N : A table-text data question-answering model based on knowledge-fusion encoder and EX-N tree decoder
Jiawang Liu, Wenqian Cao, Xiugong Qin, Yunlong Tian, Yongjie Du |
Neurocomputing | 2 |
| 2023 | Attention Based Relation Network for Facial Action Units RecognitionabstractFacial action unit (AU) recognition is essential to facial expression analysis. Since there are highly positive or negative correlations between AUs, some existing AU recognition works have focused on modeling AU relations. However, previous relationship-based approaches typically embed predefined rules into their models and ignore the impact of various AU relations in different crowds. In this paper, we propose a novel Attention Based Relation Network (ABRNet) for AU recognition, which can automatically capture AU relations without unnecessary or even disturbing predefined rules. ABRNet uses several relation learning layers to automatically capture different AU relations. The learned AU relation features are then fed into a self-attention fusion module, which aims to refine individual AU features with attention weights to enhance the feature robustness. Furthermore, we propose an AU relation dropout strategy and AU relation loss (AUR-Loss) to better model AU relations, which can further improve AU recognition. Extensive experiments show that our approach achieves state-of-the-art performance on the DISFA and DISFA+ datasets. Haoxiang Wang 0002, Jiawang Liu |
ICASSP | 4 |
| 2023 | Evaluation on the generalization of coded features across neural networks of different tasksabstractRecent advances in deep neural networks (DNNs) for computer vision tasks have made intelligent analysis on edge devices more prevalent and practical. To better distribute computational load between edge devices and the cloud, a novel deep learning deployment strategy called Collaborative Intelligence (CI) has been proposed. In this strategy, features extracted from edge devices are first compressed and then transmitted to the cloud. However, it is unclear whether these compressed features have enough information to perform diverse downstream tasks. This paper focuses on the generalization of compressed features from one neural network among other object detection and instance segmentation task networks. We first propose a scheme to evaluate the generalization of features and further perform experiments on feature compression. Our experiments show that the extracted features contain enough information for other task networks and feature compression scheme for multi-task networks offers a 82.04% average bitrate saving compared to VVC. Jiawang Liu, Ke Jia, Hualong Yu, Lu Yu 0003 |
VCIP | 1 |
| 2022 | Graph based emotion recognition with attention pooling for variable-length utterances
Jiawang Liu, Haoxiang Wang 0002 |
Neurocomputing | 1 |
| 2021 | Graph Isomorphism Network for Speech Emotion Recognition
Jiawang Liu, Haoxiang Wang 0002 |
Interspeech | 1 |
| 2021 | A Speech Emotion Recognition Framework for Better Discrimination of Confusions
Jiawang Liu, Haoxiang Wang 0002 |
Interspeech | 1 |
| 2011 | Key Concepts Identification and Weighting in Search Engine Queries
Jiawang Liu |
APWeb | 1 |