VLDB 2026 Research / reviewers in the wild / expert
Jiaqi Wang 0006
dblp:44/740-6
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0000-1766-7774ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SS-GEN: A Social Story Generation Framework with Large Language ModelsabstractChildren with Autism Spectrum Disorder (ASD) often misunderstand social situations and struggle to participate in daily routines. Social Stories™ are traditionally crafted by psychology experts under strict constraints to address these challenges but are costly and limited in diversity. As Large Language Models (LLMs) advance, there's an opportunity to develop more automated, affordable, and accessible methods to generate Social Stories in real-time with broad coverage. However, adapting LLMs to meet the unique and strict constraints of Social Stories is a challenging issue. To this end, we propose SS-GEN, a Social Story GENeration framework with LLMs. Firstly, we develop a constraint-driven sophisticated strategy named StarSow to hierarchically prompt LLMs to generate Social Stories at scale, followed by rigorous human filtering to build a high-quality dataset. Additionally, we introduce quality assessment criteria to evaluate the effectiveness of these generated stories. Considering that powerful closed-source large models require very complex instructions and expensive API fees, we finally fine-tune smaller language models with our curated high-quality dataset, achieving comparable results at lower costs and with simpler instruction and deployment. This work marks a significant step in leveraging AI to personalize Social Stories cost-effectively for autistic children at scale, which we hope can encourage future research on special groups. Jiaqi Wang 0006, Zhuang Chen 0002, Guanqun Bi, Minlie Huang, Liping Jing, Jian Yu 0001 |
AAAI | 3 |
| 2025 | Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language ModelsabstractYi Feng, Jiaqi Wang, Wenxuan Zhang, Zhuang Chen, Shen Yutong, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiaqi Wang 0006, Zhuang Chen 0002, Yutong Shen, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu 0001 |
EMNLP | 2 |
| 2024 | Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentabstractMost works of interpretable neural networks strive for learning the semantics concepts merely from single modal information such as images. However, humans usually learn semantic concepts from multiple modalities and the semantics is encoded by the brain from fused multi-modal information. Inspired by cognitive science and vision-language learning, we propose a Prototype-Concept Alignment Network (ProCoNet) for learning visual prototypes under the guidance of textual concepts. In the ProCoNet, we have designed a visual encoder to decompose the input image into regional features of prototypes, while also developing a prompt generation strategy that incorporates in-context learning to prompt large language models to generate textual concepts. To align visual prototypes with textual concepts, we leverage the multimodal space provided by the pre-trained CLIP as a bridge. Specifically, the regional features from the vision space and the cropped regions of prototypes encoded by CLIP reside on different but semantically highly correlated manifolds, i.e. follow a multi-manifold distribution. We transform the multi-manifold distribution alignment problem into optimizing the projection matrix by Cayley transform on the Stiefel manifold. Through the learned projection matrix, visual prototypes can be projected into the multimodal space to align with semantically similar textual concept features encoded by CLIP. We conducted two case studies on the CUB-200-2011 and Oxford Flower dataset. Our experiments show that the ProCoNet provides higher accuracy and better interpretability compared to the single-modality interpretable model. Furthermore, ProCoNet offers a level of interpretability not previously available in other interpretable methods. Jiaqi Wang 0006, Pichao Wang, Huafeng Liu 0001, Chang Gao 0007, Liping Jing |
ACM Multimedia | 1 |
| 2024 | Transparent Embedding Space for Interpretable Image RecognitionabstractWhen humans explain their reasoning, such as their classification decisions, they often break down an image into parts and highlight the evidence from those parts to support the concepts they have in mind. Drawing inspiration from this cognitive process, several self-explaining models have been proposed to explain predictions by part-level concepts. However, these models can be limited by their structure and difficulty in determining the effect of individual parts on the output category. To address these challenges, we introduce a self-explaining architecture that uses a plug-in transparent embedding space (TesNet) to connect high-level input patches (e.g. feature maps or tokens) with output categories. The transparent embedding space is spanned by basis concepts and constructed on the Grassmann manifold. The basis concepts are enforced to be category-aware, and within-category concepts are orthogonal to each other, ensuring the embedding space is disentangled. To reduce concept redundancy and restore the concept space structure, we introduce two concept pruning methods and a new re-training strategy to build a slimming transparent embedding space. We verify the scalability of TesNet through experiments on deep networks such as VGG, ResNet, DenseNet, and Vision Transformer. Additionally, we design several metrics for self-explaining models to quantify interpretability and compare them with state-of-the-art self-explaining methods. Our experiments demonstrate that TesNet is much more effective for classification tasks, providing better interpretability on predictions and improving final accuracy. Jiaqi Wang 0006, Huafeng Liu 0001, Liping Jing |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Interpretable Image Recognition by Screening Class-Specific and Class-Shared Prototypes
Jiaqi Wang 0006, Liping Jing |
ICANN (2) | 2 |
| 2023 | Siamese transformer with hierarchical concept embedding for fine-grained image recognition
Yilin Lyu, Liping Jing, Jiaqi Wang 0006, Mingzhe Guo, Jian Yu 0001 |
Sci. China Inf. Sci. | 3 |
| 2023 | Deep Generative Mixture Model for Robust Imbalance ClassificationabstractDiscovering hidden pattern from imbalanced data is a critical issue in various real-world applications. Existing classification methods usually suffer from the limitation of data especially for minority classes, and result in unstable prediction and low performance. In this paper, a deep generative classifier is proposed to mitigate this issue via both model perturbation and data perturbation. Specially, the proposed generative classifier is derived from a deep latent variable model where two variables are involved. One variable is to capture the essential information of the original data, denoted as latent codes, which are represented by a probability distribution rather than a single fixed value. The learnt distribution aims to enforce the uncertainty of model and implement model perturbation, thus, lead to stable predictions. The other variable is a prior to latent codes so that the codes are restricted to lie on components in Gaussian Mixture Model. As a confounder affecting generative processes of data (feature/label), the latent variables are supposed to capture the discriminative latent distribution and implement data perturbation. Extensive experiments have been conducted on widely-used real imbalanced image datasets. Experimental results demonstrate the superiority of our proposed model by comparing with popular imbalanced classification baselines on imbalance classification task. Liping Jing, Yilin Lyu, Mingzhe Guo, Jiaqi Wang 0006, Huafeng Liu 0001, Jian Yu 0001, Tieyong Zeng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Deep Amortized Relational Model with Group-Wise Hierarchical Generative ProcessabstractIn this paper, we propose Deep amortized Relational Model (DaRM) with group-wise hierarchical generative process for community discovery and link prediction on relational data (e.g., graph, network). It provides an efficient neural relational model architecture by grouping nodes in a group-wise view rather than node-wise or edge-wise view. DaRM simultaneously learns what makes a group, how to divide nodes into groups, and how to adaptively control the number of groups. The dedicated group generative process is able to sufficiently exploit pair-wise or higher-order interactions between data points in both inter-group and intra-group, which is useful to sufficiently mine the hidden structure among data. A series of experiments have been conducted on both synthetic and real-world datasets. The experimental results demonstrated that DaRM can obtain high performance on both community detection and link prediction tasks. Huafeng Liu 0001, Jiaqi Wang 0006 |
AAAI | 3 |
| 2021 | Cluster-Wise Hierarchical Generative Model for Deep Amortized ClusteringabstractIn this paper, we propose Cluster-wise Hierarchical Generative Model for deep amortized clustering (CHiGac). It provides an efficient neural clustering architecture by grouping data points in a cluster-wise view rather than point-wise view. CHiGac simultaneously learns what makes a cluster, how to group data points into clusters, and how to adaptively control the number of clusters. The dedicated cluster generative process is able to sufficiently exploit pair-wise or higher-order interactions between data points in both inter- and intra-cluster, which is useful to sufficiently mine the hidden structure among data. To efficiently minimize the generalized lower bound of CHiGac, we design an Ergodic Amortized Inference (EAI) strategy by considering the average behavior over sequence on an inner variational parameter trajectory, which is theoretically proven to reduce the amortization gap. A series of experiments have been conducted on both synthetic and real-world data. The experimental results demonstrated that CHiGac can efficiently and accurately cluster datasets in terms of both internal and external evaluation metrics (DBI and ACC). Huafeng Liu 0001, Jiaqi Wang 0006, Liping Jing |
CVPR | 2 |
| 2021 | Interpretable Image Recognition by Constructing Transparent Embedding SpaceabstractHumans usually explain their reasoning (e.g. classification) by dissecting the image and pointing out the evidence from these parts to the concepts in their minds. Inspired by this cognitive process, several part-level interpretable neural network architectures have been proposed to explain the predictions. However, they suffer from the complex data structure and confusing the effect of the individual part to output category. In this work, an interpretable image recognition deep network is designed by introducing a plug-in transparent embedding space (TesNet) to bridge the high-level input patches (e.g. CNN feature maps) and the out- put categories. This plug-in embedding space is spanned by transparent basis concepts which are constructed on the Grassmann manifold. These basis concepts are enforced to be category-aware and within-category concepts are orthogonal to each other, which makes sure the embedding space is disentangled. Meanwhile, each basis concept can be traced back to the particular image patches, thus they are transparent and friendly to explain the reasoning process. By comparing with state-of-the-art interpretable methods, TesNet is much more beneficial to classification tasks, esp. providing better interpretability on predictions and improve the final accuracy. The code is available at https://github.com/JackeyWang96/TesNet. Jiaqi Wang 0006, Huafeng Liu 0001, Liping Jing |
ICCV | 1 |
| 2021 | Interpretable Deep Generative Recommendation ModelsabstractUser preference modeling in recommendation system aims to improve customer experience through discovering users’ intrinsic preference based on prior user behavior data. This is a challenging issue because user preferences usually have complicated structure, such as inter-user preference similarity and intra-user preference diversity. Among them, inter-user similarity indicates different users may share similar preference, while intra-user diversity indicates one user may have several preferences. In literatures, deep generative models have been successfully applied in recommendation systems due to its flexibility on statistical distributions and strong ability for non-linear representation learning. However, they suffer from the simple generative process when handling complex user preferences. Meanwhile, the latent representations learned by deep generative models are usually entangled, and may range from observed-level ones that dominate the complex correlations between users, to latent-level ones that characterize a user’s preference, which makes the deep model hard to explain and unfriendly for recommendation. Thus, in this paper, we propose an Interpretable Deep Generative Recommendation Model (InDGRM) to characterize inter-user preference similarity and intra-user preference diversity, which will simultaneously disentangle the learned representation from observed-level and latent-level. In InDGRM, the observed-level disentanglement on users is achieved by modeling the user-cluster structure (i.e., inter-user preference similarity) in a rich multimodal space, so that users with similar preferences are assigned into the same cluster. The observed-level disentanglement on items is achieved by modeling the intra-user preference diversity in a prototype learning strategy, where different user intentions are captured by item groups (one group refers to one intention). To promote disentangled latent representations, InDGRM adopts structure and sparsity-inducing penalty and integrates them into the generative procedure, which has ability to enforce each latent factor focus on a limited subset of items (e.g., one item group) and benefit latent-level disentanglement. Meanwhile, it can be efficiently inferred by minimizing its penalized upper bound with the aid of local variational optimization technique. Theoretically, we analyze the generalization error bound of InDGRM to guarantee its performance. A series of experimental results on four widely-used benchmark datasets demonstrates the superiority of InDGRM on recommendation performance and interpretability. Huafeng Liu 0001, Liping Jing, Jingxuan Wen, Pengyu Xu, Jiaqi Wang 0006, Jian Yu 0001, Michael Kwok-Po Ng |
J. Mach. Learn. Res. | 5 |
| 2020 | Deep Global and Local Generative Model for RecommendationabstractDeep generative model, especially variational auto-encoder (VAE), has been successfully employed by more and more recommendation systems. The reason is that it combines the flexibility of probabilistic generative model with the powerful non-linear feature representation ability of deep neural networks. The existing VAE-based recommendation models are usually proposed under global assumption by incorporating simple priors, e.g., a single Gaussian, to regularize the latent variables. This strategy, however, is ineffective when the user is simultaneously interested in different kinds of items, i.e., the user’s preference may be highly diverse. In this paper, thus, we propose a Deep Global and Local Generative Model for recommendation to consider both local and global structure among users (DGLGM) under the Wasserstein auto-encoder framework. Besides keeping the global structure like the existing model, DGLGM adopts a non-parametric Mixture Gaussian distribution with several components to capture the diversity of the users’ preferences. Each component is corresponding to one local structure and its optimal size can be determined via the automatic relevance determination technique. These two parts can be seamlessly integrated and enhance each other. The proposed DGLGM can be efficiently inferred by minimizing its penalized upper bound with the aid of local variational optimization technique. Meanwhile, we theoretically analyze its generalization error bounds to guarantee its performance in sparse feedback data with diversity. By comparing with the state-of-the-art methods, the experimental results demonstrate that DGLGM consistently benefits the recommendation system in top-N recommendation task. Huafeng Liu 0001, Liping Jing, Jingxuan Wen, Zhicheng Wu, Xiaoyi Sun, Jiaqi Wang 0006, Jian Yu 0001 |
WWW | 6 |