Zixun Sun

dblp:245/2700 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0003-4125-1909ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unifying Adversarial Multi-Deconfounded Learning Paradigm for Fake News Detection
abstract
In the task of fake news detection, ensuring authenticity and accuracy is of paramount importance. This task, however, is susceptible to the influence of confounders, necessitating effective confounder debiasing strategies. Conventional methods are typically designed to address specific confounders, resulting in frameworks that relatively lack generalization and overlook potential correlations among confounders. The presence of multiple confounders further escalates the complexity and challenges of debiasing learning. To tackle this issue, we introduce the Adversarial Multi-Deconfounded (AMD) Learning Paradigm, a generic training framework designed to eliminate biases from multiple confounders. Our approach leverages adversarial networks to extract confounder-invariant feature representations, guiding the model to ignore potential biases introduced by confounders and extract stable representations independent of these confounders, thereby enhancing generalization. Comprehensive experiments demonstrate that our approach outperforms state-of-the-art methods on the Weibo and GossipCop datasets, and significantly exceeds other methods in generalization evaluation on CHEF. Additionally, we validate that our AMD framework exhibits improved robustness against confounders.
Zixun Sun, Mingye Xu, Guanming Liang
KDD (1)1
2025 Anchor-Regularized GAN Priors
abstract
This study presents anchor-regularized generative adversarial network (GAN) priors to delicately explore the inherent knowledge of a pretrained generative model. Previous research leveraged the latent space of a pretrained GAN model to provide a variety of image-editing operations. However, the semantically meaningful regions within latent space are distinctly bounded; therefore, the manipulation of the latent code can easily land out of the domain. To address this problem, we introduce an anchoring mechanism that enables novel and robust image editing. The key insights driving the method are that latent space is structurally organized, and that natural coherence allows semantically correlated latent code to be located in the areas surrounding a meaningful anchor. By using different input anchors, the proposed method forms the basis for a variety of robust and flexible editing operations, including misaligned domain translation, interactive editing, and few-shot interpretable direction exploration. Extensive experiments demonstrated the superior performance of the proposed method compared with state-of-the-art editing methods.
Huiting Yang, Yang Zhou 0038, Zhansheng Li, Liangyu Chai, Panan Wu, Zixun Sun, Shengfeng He
Comput. Vis. Media7
2024 DELTA: Dynamic Embedding Learning with Truncated Conscious Attention for CTR Prediction
Liang Du 0004, Hong Chen 0011, Zixun Sun, Xin Wang 0019, Wenwu Zhu 0001
CogSci5
2024 Stable Heterogeneous Treatment Effect Estimation across Out-of-Distribution Populations
abstract
Heterogeneous treatment effect (HTE) estimation is vital for understanding the change of treatment effect across individuals or subgroups. Most existing HTE estimation methods focus on addressing selection bias induced by imbalanced distributions of confounders between treated and control units, but ignore distribution shifts across populations. Thereby, their applicability has been limited to the in-distribution (ID) population, which shares a similar distribution with the training dataset. In real-world applications, where population distributions are subject to continuous changes, there is an urgent need for stable HTE estimation across out-of-distribution (OOD) populations, which, however, remains an open problem. As pioneers in resolving this problem, we propose a novel Stable Balanced Representation Learning with Hierarchical-Attention Paradigm (SBRL-HAP) framework, which consists of 1) Balancing Regularizer for eliminating selection bias, 2) Independence Regularizer for addressing the distribution shift issue, 3) Hierarchical-Attention Paradigm for coordination between balance and independence. In this way, SBRL- HAP regresses counterfactual outcomes using ID data, while ensuring the resulting HTE estimation can be successfully generalized to out-of-distribution scenarios, thereby enhancing the model's applicability in real-world settings. Extensive experiments conducted on synthetic and real-world datasets demonstrate the effectiveness of our SBRL-HAP in achieving stable HTE estimation across OOD populations, with an average 10% reduction in the error metric PEHE and 11% decrease in the ATE bias, compared to the SOTA methods.
Anpeng Wu, Kun Kuang 0001, Liang Du 0004, Zixun Sun
ICDE5
2024 Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
abstract
Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scarcity of labeled data. In this paper, we propose an end-to-end method, PaRe, to enhance cross-modal fine-tuning, aiming to transfer a large-scale pretrained model to various target modalities. PaRe employs a gating mechanism to select key patches from both source and target data. Through a modality-agnostic Patch Replacement scheme, these patches are preserved and combined to construct data-rich intermediate modalities ranging from easy to hard. By gradually intermediate modality generation, we can not only effectively bridge the modality gap to enhance stability and transferability of cross-modal fine-tuning, but also address the challenge of limited data in the target modality by leveraging enriched intermediate modality data. Compared with hand-designed, general-purpose, task-specific, and state-of-the-art cross-modal fine-tuning approaches, PaRe demonstrates superior performance across three challenging benchmarks, encompassing more than ten modalities.
Lincan Cai, Shuang Li 0008, Wenxuan Ma 0001, Jingxuan Kang, Binhui Xie, Zixun Sun, Chengwei Zhu
ICML6
2024 Weight Diffusion for Future: Learn to Generalize in Non-Stationary Environments
abstract
Enabling deep models to generalize in non-stationary environments is vital for real-world machine learning, as data distributions are often found to continually change. Recently, evolving domain generalization (EDG) has emerged to tackle the domain generalization in a time-varying system, where the domain gradually evolves over time in an underlying continuous structure. Nevertheless, it typically assumes multiple source domains simultaneously ready. It still remains an open problem to address EDG in the domain-incremental setting, where source domains are non-static and arrive sequentially to mimic the evolution of training domains. To this end, we propose Weight Diffusion (W-Diff), a novel framework that utilizes the conditional diffusion model in the parameter space to learn the evolving pattern of classifiers during the domain-incremental training process. Specifically, the diffusion model is conditioned on the classifier weights of different historical domain (regarded as a reference point) and the prototypes of current domain, to learn the evolution from the reference point to the classifier weights of current domain (regarded as the anchor point). In addition, a domain-shared feature encoder is learned by enforcing prediction consistency among multiple classifiers, so as to mitigate the overfitting problem and restrict the evolving pattern to be reflected in the classifier as much as possible. During inference, we adopt the ensemble manner based on a great number of target domain-customized classifiers, which are cheaply obtained via the conditional diffusion model, for robust prediction. Comprehensive experiments on both synthetic and real-world datasets show the superior generalization performance of W-Diff on unseen domains in the future.
Mixue Xie, Shuang Li 0008, Binhui Xie, Chi Harold Liu, Jian Liang 0002, Zixun Sun, Chengwei Zhu
NeurIPS6
2024 RE-SORT: Removing Spurious Correlation in Multilevel Interaction for CTR Prediction
abstract
Click-through rate (CTR) prediction is a critical task in recommendation systems, serving as the ultimate filtering step to sort items for a user. Most recent cutting-edge methods primarily focus on investigating complex implicit and explicit feature interactions; however, these methods neglect the spurious correlation issue caused by confounding factors, thereby diminishing the model’s generalization ability. We propose a CTR prediction framework that REmoves Spurious cORrelations in mulTilevel feature interactions, termed RE-SORT, which has two key components. I. A multilevel stacked recurrent (MSR) structure enables the model to efficiently capture diverse nonlinear interactions from feature spaces at different levels. II. A spurious correlation elimination (SCE) module further leverages Laplacian kernel mapping and sample reweighting methods to eliminate the spurious correlations concealed within the multilevel features, allowing the model to focus on the true causal features. Extensive experiments conducted on four challenging CTR datasets, our production dataset, and an online A/B test demonstrate that the proposed method achieves state-of-the-art performance in both accuracy and speed. The utilized codes, models, and dataset will be released at https://github.com/RE-SORT.
Songli Wu, Liang Du 0004, Yuai Wang, De-Chuan Zhan, Zixun Sun
UAI7
2024 Unsupervised Modality-Transferable Video Highlight Detection With Representation Activation Sequence Learning
abstract
Identifying highlight moments of raw video materials is crucial for improving the efficiency of editing videos that are pervasive on internet platforms. However, the extensive work of manually labeling footage has created obstacles to applying supervised methods to videos of unseen categories. The absence of an audio modality that contains valuable cues for highlight detection in many videos also makes it difficult to use multimodal strategies. In this paper, we propose a novel model with cross-modal perception for unsupervised highlight detection. The proposed model learns representations with visual-audio level semantics from image-audio pair data via a self-reconstruction task. To achieve unsupervised highlight detection, we investigate the latent representations of the network and propose the representation activation sequence learning (RASL) module with k-point contrastive learning to learn significant representation activations. To connect the visual modality with the audio modality, we use the symmetric contrastive learning (SCL) module to learn the paired visual and audio representations. Furthermore, an auxiliary task of masked feature vector sequence (FVS) reconstruction is simultaneously conducted during pretraining for representation enhancement. During inference, the cross-modal pretrained model can generate representations with paired visual-audio semantics given only the visual modality. The RASL module is used to output the highlight scores. The experimental results show that the proposed framework achieves superior performance compared to other state-of-the-art approaches.
Tingtian Li, Zixun Sun, Xinyu Xiao
IEEE Trans. Image Process.2
2022 Relational Graph Reasoning Transformer for Image Captioning
abstract
The current published methods of image captioning are directly inputting the features of objects in image into model, and introduced a variety of attention mechanisms to capture the associations between the objects and specific words. But the relationships of vision and semantic between objects are not sufficiently concerned. In this paper, we propose a relational graph reasoning Transformer which explicitly incorporates the relationships of vision and semantic between objects to construct an object relational graph in Transformer. Specifically, besides the detected object features, the global spatial relationships and the semantic context between different objects is attended. Meanwhile, a graph structures feature which correlates object features, their spatial and semantic information is reasoned by a learned grafting mechanism. Finally, the contextual graph feature is integrated into the proposed Transformer decoder. Experimental results demonstrate the significance of our relationship reasoning Transformer model.
Xinyu Xiao, Zixun Sun, Tingtian Li, Yipeng Yu
ICME2
2022 CALM: Constrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis
Xiang Li 0105, Zhiyong Wu 0001, Tingtian Li, Zixun Sun, Xinyu Xiao, Chi Sun, Hui Zhan, Helen M. Meng
INTERSPEECH5
2021 Discovering Interpretable Latent Space Directions of GANs Beyond Binary Attributes
abstract
Generative adversarial networks (GANs) learn to map noise latent vectors to high-fidelity image outputs. It is found that the input latent space shows semantic correlations with the output image space. Recent works aim to interpret the latent space and discover meaningful directions that correspond to human interpretable image transformations. However, these methods either rely on explicit scores of attributes (e.g., memorability) or are restricted to binary ones (e.g., gender), which largely limits the applicability of editing tasks, especially for free-form artistic tasks like style/anime editing. In this paper, we propose an adversarial method, AdvStyle, for discovering interpretable directions in the absence of well-labeled scores or binary attributes. In particular, the proposed adversarial method simultaneously optimizes the discovered directions and the attribute assessor using the target attribute data as positive samples, while the generated ones being negative. In this way, arbitrary attributes can be edited by collecting positive data only, and the proposed method learns a controllable representation enabling manipulation of non-binary attributes like anime styles and facial characteristics. Moreover, the proposed learning strategy attenuates the entanglement between attributes, such that multi-attribute manipulation can be easily achieved without any additional constraint. Furthermore, we reveal several interesting semantics with the involuntarily learned negative directions. Extensive experiments on 9 anime attributes and 7 human attributes demonstrate the effectiveness of our adversarial approach qualitatively and quantitatively. Code is available at https://github.com/BERYLSHEEP/AdvStyle.
Huiting Yang, Liangyu Chai, Zixun Sun, Shengfeng He
CVPR5
2021 Hierarchical Multilabel Text Classification via Multitask Learning
abstract
Hierarchical multilabel classification is a variant of classification where instances might belong to multiple labels and these labels come from a hierarchy. In this paper, we solve the hierarchical multilabel text classification problem of professionally-generated content via multitask learning. More specifically, we focus on (1) how to build models that can share features well in multitask learning, (2) how to incorporate the label dependence into the training procedure of the models, and (3) how to combine the predicted labels of different levels in the hierarchy. To make the experiments simple and comparable, we bring in the state-of-art BERT model as the base model in our work. Experiment results show that the multitask models we build are competitive, the penalty loss we propose is able to improve the performance, and the union operation is the best choice to handle prediction contradiction. In other words, the time cost is reduced but performance is improved via our multitask learning approach.
Yipeng Yu, Zixun Sun, Chi Sun
ICTAI2
2021 Text2Video: Automatic Video Generation Based on Text Scripts
abstract
To make video creation simpler, in this paper we present Text2Video, a novel system to automatically produce videos using only text-editing for novice users. Given an input text script, the director-like system can generate game-related engaging videos which illustrate the given narrative, provide diverse multi-modal content, and follow video editing guidelines. The system involves five modules: (1) A material manager extracts highlights from raw live game videos, and tags each video highlight, image and audio with labels. (2) A natural language processor extracts entities and semantics from the input text scripts. (3) A refined cross-modal retrieval searches for matching candidate shots from the material manager. (4) A text to speech speaker reads the processed text scripts with synthesized human voice. (5) The selected material shots and synthesized speech are assembled artistically through appropriate video editing techniques.
Yipeng Yu, Zirui Tu, Longyu Lu, Hui Zhan, Zixun Sun
ACM Multimedia6
2021 Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs
abstract
Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on the visual modality cannot show promising performance regarding videos with fine-grained virtual contents. In this paper, we also investigate the widely added voice-overs in short videos and propose a novel framework to retrieve BGM for fine-grained short videos. In our framework, we use the self-attention (SA) and the cross-modal attention (CMA) modules to explore the intra- and the inter-relationships of different modalities respectively. For balancing the modalities, we dynamically assign different weights to the modal features via a fusion gate. For paring the query and the BGM embeddings, we introduce a triplet pseudo-label loss to constrain the semantics of the modal embeddings. As there are no existing virtual-content video-BGM retrieval datasets, we build and release two virtual-content video datasets HoK400 and CFM400. Experimental results show that our method achieves superior performance and outperforms other state-of-the-art methods with large margins.
Tingtian Li, Zixun Sun, Haoruo Zhang, Ziming Wu, Hui Zhan, Yipeng Yu, Hengcan Shi
SIGIR2
2020 Group-Skeleton-Based Human Action Recognition in Complex Events
abstract
Human action recognition as an important application of computer vision has been studied for decades. Among various approaches, skeleton-based methods recently attract increasing attention due to their robust and superior performance. However, existing skeleton-based methods ignore the potential action relationships between different persons, while the action of a person is highly likely to be impacted by another person especially in complex events. In this paper, we propose a novel group-skeleton-based human action recognition method in complex events. This method first utilizes multi-scale spatial-temporal graph convolutional networks (MS-G3Ds) to extract skeleton features from multiple persons. In addition to the traditional key point coordinates, we also input the key point speed values to the networks for better performance. Then we use multilayer perceptrons (MLPs) to embed the distance values between the reference person and other persons into the extracted features. Lastly, all the features are fed into another MS-G3D for feature fusion and classification. For avoiding class imbalance problems, the networks are trained with a focal loss. The proposed algorithm is also our solution for the Large-scale Human-centric Video Analysis in Complex Events Challenge. Results on the HiEve dataset show that our method can give superior performance compared to other state-of-the-art methods.
Tingtian Li, Zixun Sun
ACM Multimedia2
2020 Two-stage Photograph Cartoonization via Line Tracing
abstract
Abstract Cartoon is highly abstracted with clear edges, which makes it unique from the other art forms. In this paper, we focus on the essential cartoon factors of abstraction and edges, aiming to cartoonize real‐world photographs like an artist. To this end, we propose a two‐stage network, each stage explicitly targets at producing abstracted shading and crisp edges respectively. In the first abstraction stage, we propose a novel unsupervised bilateral flattening loss, which allows generating high‐quality smoothing results in a label‐free manner. Together with two other semantic‐aware losses, the abstraction stage imposes different forms of regularization for creating cartoon‐like flattened images. In the second stage we draw lines on the structural edges of the flattened cartoon with the fully supervised line drawing objective and unsupervised edge augmenting loss. We collect a cartoon‐line dataset with line tracing, and it serves as the starting point for preparing abstraction and line drawing data. We have evaluated the proposed method on a large number of photographs, by converting them to three different cartoon styles. Our method substantially outperforms state‐of‐the‐art methods in terms of visual quality quantitatively and qualitatively.
Zixun Sun, Shengfeng He
Comput. Graph. Forum4