Zhixiong Zeng

dblp:23/10840 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-3822-1074ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An audio-augmented fusion model for weakly supervised video moment retrieval
Shuyi He, Qingchao Kong, Zhixiong Zeng, Wenji Mao
Neurocomputing3
2026 d-MAR: Deep Metal Artifact Reduction via Diffusion-Driven Domain Transformations
abstract
Metal implants introduce severe artifacts in CT images, compromising diagnostic reliability. Supervised metal artifact reduction (MAR) models trained on simulated data are effective but often fail due to domain gaps when applied to real clinical data. Unsupervised methods trained on real images avoid such gaps but suffer from weak artifact suppression and training instability. To address these challenges, we propose d-MAR, a novel MAR framework that performs diffusion-driven domain transformations between simulated and real image domains. Specifically, real image domain (RID) data is transformed into the simulated image domain (SID), processed by a MAR model trained on simulation-paired data, and transformed back into RID. We harness diffusion models as a transformation bridge and introduce two targeted conditional sampling techniques-conditional input and sampling enhancement-based on Fourier-extracted low-frequency image components. This enables domain alignment without random generation, ensuring consistent anatomical fidelity. The proposed d-MAR can reduce real metal artifacts originating from different scanning protocols and devices with a MAR model trained with simulated paired data. Evaluations on Clinical Head, Clinical Body, and dental CBCT datasets show that d-MAR consistently outperforms conventional MAR methods in both quantitative metrics and visual quality, demonstrating strong generalization capability.
Zhixiong Zeng, Yuyan Song, Mingjun Lu, Yaoduo Zhang, Ji He 0001, Zhibo Wen, Dong Zeng, Zhaoying Bian, Jianhua Ma 0001
IEEE J. Biomed. Health Informatics2
2025 Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
abstract
In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that facilitates higher autoregressive generation quality with faster synthesis speed. Specifically, the proposed method introduces two main innovations: 1) dynamic temperature control guided by spatial entropy of token distributions, enhancing the balance between content diversity, alignment accuracy, and structural coherence in both mask-based and scale-wise models, without extra computational overhead, and 2) entropy-aware acceptance rules in speculative decoding, achieving near-lossless generation at about 85% of the inference cost of conventional acceleration methods. Extensive experiments across multiple benchmarks using diverse AR image generation models demonstrate the effectiveness and generalizability of our approach in enhancing both generation quality and sampling speed.
Feng Zhao 0004, Pengyang Ling, Haibo Qiu, Zhixiang Wei, Hu Yu 0001, Jie Huang 0017, Zhixiong Zeng, Lin Ma 0002
NeurIPS8
2025 A multimodal embedding transfer approach for consistent and selective learning processes in cross-modal retrieval
Zhixiong Zeng, Shuyi He, Wenji Mao
Inf. Sci.1
2024 Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
abstract
Supervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks.In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in performance degradation.Consequently, we propose to mitigate this imbalance by merging SFT models fine-tuned with different data orders, thereby enhancing the overall effectiveness of SFT.Additionally, we introduce a novel technique, "parameter-selection merging," which outperforms traditional weightedaverage methods on five datasets.Further, through analysis and ablation studies, we validate the effectiveness of our method and identify the sources of performance improvements.
Yiming Ju, Ziyi Ni, Xingrun Xing, Zhixiong Zeng, Siqi Fan 0001, Zheng Zhang 0006
EMNLP4
2024 GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
abstract
Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reasoning. The potential benefits and capabilities of representing graph structures as visual images (i.e., $\textit{visual graph}$) are still unexplored. To fill the gap, we innovatively propose an end-to-end framework, called $\textbf{G}$raph to v$\textbf{I}$sual and $\textbf{T}$extual Integr$\textbf{A}$tion (GITA), which firstly incorporates visual graphs into general graph reasoning. Besides, we establish $\textbf{G}$raph-based $\textbf{V}$ision-$\textbf{L}$anguage $\textbf{Q}$uestion $\textbf{A}$nswering (GVLQA) dataset from existing graph data, which is the first vision-language dataset for general graph reasoning purposes. Extensive experiments on the GVLQA dataset and five real-world datasets show that GITA outperforms mainstream LLMs in terms of general graph reasoning capabilities. Moreover, We highlight the effectiveness of the layout augmentation on visual graphs and pretraining on the GVLQA dataset.
Yanbin Wei, Weisen Jiang, Zejian Zhang, Zhixiong Zeng, James T. Kwok, Yu Zhang 0006
NeurIPS5
2024 KLoB: a Benchmark for Assessing Knowledge Localization Methods in Language Models
Yiming Ju, Xingrun Xing, Zhixiong Zeng
PRICAI (2)4
2022 C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media Retrieval
abstract
Cross-modal retrieval is an essential area of representation learning, which aims to retrieve instances with the same semantics from different modalities. In real implementation, a key challenge for cross-modal retrieval is to narrow the heterogeneity gap between different modalities and obtain modality-invariant and discriminative features. Typically, existing approaches for this task mainly learn inter-modal invariance and focus on how to combine pair-level loss and class-level loss, which cannot effectively and adequately learn discriminative features. To address these issues, in this paper, we propose a novel Cross-Modality Cross-Instance Contrastive Learning for Cross-Media Retrieval (C3CMR) method. Specifically, to fully employ the intra-modal similarities, we introduce the intra-modal contrastive learning to enhance the discriminative power of the unimodal features. Besides, we design a supervised inter-modal contrastive learning scheme to take full advantage of the label semantic associations. In this way, cross-semantic associations and inter-modal invariance can be further learned. Moreover, pertaining to the local suboptimal semantic similarity by only mining pairwise and triplewise sample relationships, we propose the cross-instance contrastive learning to mine the similarities among multiple instances. Comprehensive experimental results on four widely-used benchmark datasets demonstrate the superiority of our proposed method over several state-of-the-art cross-modal retrieval methods.
Tiantian Gong, Zhixiong Zeng, Changchang Sun, Yan Yan 0002
ACM Multimedia3
2021 AliMe MKG: A Multi-modal Knowledge Graph for Live-streaming E-commerce
abstract
Live streaming is becoming an increasingly popular trend of sales in E-commerce. The core of live-streaming sales is to encourage customers to purchase in an online broadcasting room. To enable customers to better understand a product without jumping out, we propose AliMe MKG, a multi-modal knowledge graph that aims at providing a cognitive profile for products, through which customers are able to seek information about and understand a product. Based on the MKG, we build an online live assistant that highlights product search, product exhibition and question answering, allowing customers to skim over item list, view item details, and ask item-related questions. Our system has been launched online in the Taobao app, and currently serves hundreds of thousands of customers per day.
Guohai Xu, Hehong Chen, Feng-Lin Li, Fu Sun, Yunzhou Shi, Zhixiong Zeng, Zhongzhou Zhao, Ji Zhang 0011
CIKM6
2021 MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal Retrieval
abstract
Cross-modal retrieval is an important multimedia research area which aims to take one type of data as the query to retrieve relevant data of another type. Most of the existing methods follow the paradigm of pair-wise learning and class-level learning to generate a common embedding space, where the similarity of heterogeneous multimodal samples can be calculated. However, in contrast to large-scale cross-modal retrieval applications which often need to tackle multiple modalities, previous studies on cross-modal retrieval mainly focus on two modalities (i.e., text-image or text-video). In addition, for large-scale cross-modal retrieval with modality diversity, another important problem is that the available training data are considerably modality-imbalanced. In this paper, we focus on the challenging problem of modality-imbalanced cross-modal retrieval, and propose a Multimodal Coordinated Clustering Network (MCCN) which consists of two modules, Multimodal Coordinated Embedding (MCE) module to alleviate the imbalanced training data and Multimodal Contrastive Clustering (MCC) module to tackle the imbalanced optimization. The MCE module develops a data-driven approach to coordinate multiple modalities via multimodal semantic graph for the generation of modality-balanced training samples. The MCC module learns class prototypes as anchors to preserve the pair-wise and class-level similarities across modalities for intra-class compactness and inter-class separation, and further introduces intra-class and inter-class margins to enhance optimization flexibility. We conduct experiments on the benchmark multimodal datasets to verify the effectiveness of our proposed method.
Zhixiong Zeng, Wenji Mao
ACM Multimedia1
2021 PAN: Prototype-based Adaptive Network for Robust Cross-modal Retrieval
abstract
In practical applications of cross-modal retrieval, test queries of the retrieval system may vary greatly and come from unknown category. Meanwhile, due to the cost and difficulty of data collection as well as other issues, the available data for cross-modal retrieval are often imbalanced over different modalities. In this paper, we address two important issues to increase the robustness of cross-modal retrieval system for real-world applications: handling test queries from unknown category and modality-imbalanced training data. The first issue has not been addressed by existing methods and the second issue was not well addressed in the related research. To tackle the above issues, we take the advantage of prototype learning, and propose a prototype-based adaptive network (PAN) for robust cross-modal retrieval. Our method leverages a unified prototype to represent each semantic category across modalities, which provides discriminative information of different categories and takes unified prototypes as anchors to learn cross-modal representations adaptively. Moreover, we propose a novel prototype propagation strategy to reconstruct balanced representations which preserves the semantic consistency and modality heterogeneity. Experimental results on the benchmark datasets demonstrate the effectiveness of our method compared to the SOTA methods, and further robustness tests show the superiority of our method in solving the above issues.
Zhixiong Zeng, Nan Xu 0004, Wenji Mao
SIGIR1
2020 Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic Association
abstract
Sarcasm is a sophisticated linguistic phenomenon to express the opposite of what one really means.With the rapid growth of social media, multimodal sarcastic tweets are widely posted on various social platforms.In multimodal context, sarcasm is no longer a pure linguistic phenomenon, and due to the nature of social media short text, the opposite is more often manifested via cross-modality expressions.Thus traditional text-based methods are insufficient to detect multimodal sarcasm.To reason with multimodal sarcastic tweets, in this paper, we propose a novel method for modeling cross-modality contrast in the associated context.Our method models both cross-modality contrast and semantic association by constructing the Decomposition and Relation Network (namely D&R Net).The decomposition network represents the commonality and discrepancy between image and text, and the relation network models the semantic association in cross-modality context.Experimental results on a public dataset demonstrate the effectiveness of our model in multimodal sarcasm detection.
Nan Xu 0004, Zhixiong Zeng, Wenji Mao
ACL2
2020 Event-Driven Network for Cross-Modal Retrieval
abstract
Despite extensive research on cross-modal retrieval, existing methods focus on the matching between image objects and text words. However, for the large amount of social media, such as news reports and online posts with images, previous methods are insufficient to model the associations between long text and image. As long text contains multiple entities and relationships between them, as well as complex events sharing a common scenario of the text, it poses unique research challenge to cross-modal retrieval. To tackle the challenge, in this paper, we focus on the retrieval task on long text and image, and propose an event-driven network for cross-modal retrieval. Our approach consists of two modules, namely the contextual neural tensor network (CNTN) and cross-modal matching network (CMMN). The CNTN module captures both event-level and text-level semantics of the sequential events extracted from a long text. The CMMN module learns a common representation space to compute the similarity of image and text modalities. We construct a multimodal dataset based on the news reports in People's Daily. The experimental results demonstrate that our model outperforms the existing state-of-the-art methods and can provide semantic richer text representations to enhance the effectiveness in cross-modal retrieval.
Zhixiong Zeng, Nan Xu 0004, Wenji Mao
CIKM1