Zhun Sun

dblp:185/6899 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0001-9336-2841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 9 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
abstract
Recent advancements in Large Language Models (LLMs) have demonstrated enhanced reasoning capabilities, evolving from Chainof-Thought (CoT) prompting to advanced, product-oriented solutions like OpenAI o1.During our re-implementation of this model, we noticed that in multimodal tasks requiring visual input (e.g., geometry problems), Multimodal LLMs (MLLMs) struggle to maintain focus on the visual information, in other words, MLLMs suffer from a gradual decline in attention to visual information as reasoning progresses, causing text-over-relied outputs.To investigate this, we ablate image inputs during long-chain reasoning.Concretely, we truncate the reasoning process midway, then re-complete the reasoning process with the input image removed.We observe only a ∼2% accuracy drop on MathVista's test-hard subset, revealing the model's textual outputs dominate the following reasoning process.Motivated by this, we propose Take-along Visual Conditioning (TVC), a strategy that shifts image input to critical reasoning stages and compresses redundant visual tokens via dynamic pruning.This methodology helps the model retain attention to the visual components throughout the reasoning.Our approach achieves state-of-the-art performance on average across five mathematical reasoning benchmarks (+3.4 points vs previous sota), demonstrating the effectiveness of TVC in enhancing multimodal reasoning systems.
Hai-Long Sun, Zhun Sun, Houwen Peng, Han-Jia Ye
ACL (1)2
2025 Fractional Tensor Recurrent Unit (fTRU): A Stable Forecasting Model With Long Memory
abstract
The tensor recurrent model is a family of nonlinear dynamical systems, of which the recurrence relation consists of a -fold (called degree- ) tensor product. Despite such models frequently appearing in advanced recurrent neural networks (RNNs), to this date, there are limited studies on their long memory properties and stability in sequence tasks. In this article, we propose a fractional tensor recurrent model, where the tensor degree is extended from the discrete domain to the continuous domain, so it is effectively learnable from various datasets. Theoretically, we prove that a large degree is essential to achieve the long memory effect in a tensor recurrent model, yet it could lead to unstable dynamical behaviors. Hence, our new model, named fractional tensor recurrent unit (fTRU), is expected to seek the saddle point between long memory property and model stability during the training. We experimentally show that the proposed model achieves competitive performance with a long memory and stable manners in several forecasting tasks compared to various advanced RNNs.
Hejia Qiu, Chao Li 0013, Ying Weng, Zhun Sun, Qibin Zhao
IEEE Trans. Neural Networks Learn. Syst.4
2024 tnGPS: Discovering Unknown Tensor Network Structure Search Algorithms via Large Language Models (LLMs)
abstract
Tensor networks are efficient for extremely high-dimensional representation, but their model selection, known as tensor network structure search (TN-SS), is a challenging problem. Although several works have targeted TN-SS, most existing algorithms are manually crafted heuristics with poor performance, suffering from the curse of dimensionality and local convergence. In this work, we jump out of the box, studying how to harness large language models (LLMs) to automatically discover new TN-SS algorithms, replacing the involvement of human experts. By observing how human experts innovate in research, we model their common workflow and propose an automatic algorithm discovery framework called tnGPS. The proposed framework is an elaborate prompting pipeline that instruct LLMs to generate new TN-SS algorithms through iterative refinement and enhancement. The experimental results demonstrate that the algorithms discovered by tnGPS exhibit superior performance in benchmarks compared to the current state-of-the-art methods. Our code is available at https://github.com/ChaoLiAtRIKEN/tngps.
Junhua Zeng, Chao Li 0013, Zhun Sun, Qibin Zhao, Guoxu Zhou
ICML3
2024 Transferring Vision-Language Models for Visual Recognition: A Classifier Perspective
abstract
Abstract Transferring knowledge from pre-trained deep models for downstream tasks, particularly with limited labeled samples, is a fundamental problem in computer vision research. Recent advances in large-scale, task-agnostic vision-language pre-trained models, which are learned with billions of samples, have shed new light on this problem. In this study, we investigate how to efficiently transfer aligned visual and textual knowledge for downstream visual recognition tasks. We first revisit the role of the linear classifier in the vanilla transfer learning framework, and then propose a new paradigm where the parameters of the classifier are initialized with semantic targets from the textual encoder and remain fixed during optimization. To provide a comparison, we also initialize the classifier with knowledge from various resources. In the empirical study, we demonstrate that our paradigm improves the performance and training speed of transfer learning tasks. With only minor modifications, our approach proves effective across 17 visual datasets that span three different data domains: image, video, and 3D point cloud.
Zhun Sun, YuXin Song 0001, Jingdong Wang 0001, Wanli Ouyang
Int. J. Comput. Vis.2
2023 Revisiting Classifier: Transferring Vision-Language Models for Video Recognition
abstract
Transferring knowledge from task-agnostic pre-trained deep models for downstream tasks is an important topic in computer vision research. Along with the growth of computational capacity, we now have open-source vision-language pre-trained models in large scales of the model architecture and amount of data. In this study, we focus on transferring knowledge for video classification tasks. Conventional methods randomly initialize the linear classifier head for vision classification, but they leave the usage of the text encoder for downstream visual recognition tasks undiscovered. In this paper, we revise the role of the linear classifier and replace the classifier with the different knowledge from pre-trained model. We utilize the well-pretrained language model to generate good semantic target for efficient transferring learning. The empirical study shows that our method improves both the performance and the training speed of video classification, with a negligible change in the model. Our simple yet effective tuning paradigm achieves state-of-the-art performance and efficient training on various video recognition scenarios, i.e., zero-shot, few-shot, general recognition. In particular, our paradigm achieves the state-of-the-art accuracy of 87.8% on Kinetics-400, and also surpasses previous methods by 20~50% absolute top-1 accuracy under zero-shot, few-shot settings on five video datasets. Code and models are available at https://github.com/whwu95/Text4Vis.
Zhun Sun, Wanli Ouyang
AAAI2
2023 What Can Simple Arithmetic Operations Do for Temporal Modeling?
abstract
Temporal modeling plays a crucial role in understanding video content. To tackle this problem, previous studies built complicated temporal relations through time sequence thanks to the development of computationally powerful devices. In this work, we explore the potential of four simple arithmetic operations for temporal modeling. Specifically, we first capture auxiliary temporal cues by computing addition, subtraction, multiplication, and division between pairs of extracted frame features. Then, we extract corresponding features from these cues to benefit the original temporal-irrespective domain. We term such a simple pipeline as an Arithmetic Temporal Module (ATM), which operates on the stem of a visual backbone with a plug-and-play style. We conduct comprehensive ablation studies on the instantiation of ATMs and demonstrate that this module provides powerful temporal modeling capability at a low computational cost. Moreover, the ATM is compatible with both CNNs- and ViTs-based architectures. Our results show that ATM achieves superior performance over several popular video benchmarks. Specifically, on Something-Something V1, V2 and Kinetics-400, we reach top-1 accuracy of 65.6%, 74.6%, and 89.4% respectively. The code is available at https://github.com/whwu95/ATM.
YuXin Song 0001, Zhun Sun, Jingdong Wang 0001, Chang Xu 0002, Wanli Ouyang
ICCV3
2023 Representation Disentanglement in Generative Models with Contrastive Learning
abstract
Contrastive learning has shown its effectiveness in image classification and generation. Recent works apply contrastive learning to the discriminator of the Generative Adversarial Networks. However, there is little work exploring if contrastive learning can be applied to the encoderdecoder structure to learn disentangled representations. In this work, we propose a simple yet effective method via incorporating contrastive learning into latent optimization, where we name it ContraLORD. Specifically, we first use a generator to learn discriminative and disentangled embeddings via latent optimization. Then an encoder and two momentum encoders are applied to dynamically learn disentangled information across a large number of samples with content-level and residual-level contrastive loss. In the meanwhile, we tune the encoder with the learned embeddings in an amortized manner. We evaluate our approach on ten benchmarks regarding representation disentanglement and linear classification. Extensive experiments demonstrate the effectiveness of our ContraLORD on learning both discriminative and generative representations.
Shentong Mo, Zhun Sun, Chao Li 0013
WACV2
2023 Multi-level Contrastive Learning for Self-Supervised Vision Transformers
abstract
Recent studies aim to establish contrastive self-supervised learning (CSL) algorithms specialized for the family of Vision Transformers (ViTs) to make them function normally as ordinary convolutional-based backbones in the training progress. Despite obtaining promising performance on related downstream tasks, one compelling property of the ViTs is ignored in those approaches. As previous studies have demonstrated, vision transformers benefit from the early stage global attention mechanics, obtaining feature representations that contain information from distant patches, even in their shallow layers. Motivated by this, we present a simple yet effective framework to facilitate the self-supervised feature learning of transformer based vision architectures, namely, Multi-level Contrastive learning for Vision Transformers (MCVT). Specifically, we equip the vision transformers with individual-based (InfoNCE) and prototypical-based (ProtoNCE) contrastive loss in different stages of the architecture to capture low-level invariance and high-level invariance between views of samples, respectively. We conduct extensive experiments to demonstrate the effectiveness of the proposed method, using two well-known vision transformer backbones, on several vision downstream tasks, including linear classification, detection, and semantic segmentation.
Shentong Mo, Zhun Sun, Chao Li 0013
WACV2
2022 Are we pruning the correct channels in image-to-image translation models?
Yiyong Li, Zhun Sun, Chao Li 0013
BMVC2
2022 Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation
Shentong Mo, Zhun Sun, Chao Li 0013
BMVC2
2021 On the Memory Mechanism of Tensor-Power Recurrent Models
abstract
Tensor-power (TP) recurrent model is a family of non-linear dynamical systems, of which the recurrence relation consists of a p-fold (a.k.a., degree-p) tensor product. Despite such the model frequently appears in the advanced recurrent neural networks (RNNs), to this date there is limited study on its memory property, a critical characteristic in sequence tasks. In this work, we conduct a thorough investigation of the memory mechanism of TP recurrent models. Theoretically, we prove that a large degree p is an essential condition to achieve the long memory effect, yet it would lead to unstable dynamical behaviors. Empirically, we tackle this issue by extending the degree p from discrete to a differentiable domain, such that it is efficiently learnable from a variety of datasets. Taken together, the new model is expected to benefit from the long memory effect in a stable manner. We experimentally show that the proposed model achieves competitive performance compared to various advanced RNNs in both the single-cell and seq2seq architectures.
Hejia Qiu, Chao Li 0013, Ying Weng, Zhun Sun, Qibin Zhao
AISTATS4
2021 Siamese Prototypical Contrastive Learning
Shentong Mo, Zhun Sun, Chao Li 0013
BMVC2
2021 Hide Chopin in the Music: Efficient Information Steganography Via Random Shuffling
abstract
Information steganography is a family of techniques that hide secret messages into a carrier; thus, the messages can only be extracted by receivers with a correct key Although many approaches have been proposed to achieve this purpose, historically, it is a difficult problem to conceal a large amount of information without occasioning human perceptible changes. In this paper, we explore the room introduced by the low-rank property of natural signals (i.e., images, audios), and propose a training-free model for efficient information steganography, which provides a capacity of hiding full-size images into carriers of the same spatial resolution. The key of our method is to randomly shuffle the secrets and carry out a simple reduction summation with the carrier. On the other hand, the secret images can be reconstructed by solving a convex optimization problem similar to the ordinary tensor decomposition. In the experimental analysis, we carry out two tasks: concealing a full-RGB-color image into a gray-scale image; concealing images into music signals. The results confirm the ability of our model to handle massive secret payloads. The code of our paper is provided in https://github.com/minogame/icassp-SIC.
Zhun Sun, Chao Li 0013, Qibin Zhao
ICASSP1
2021 Focus and retain: Complement the Broken Pose in Human Image Synthesis
abstract
Given a target pose, how to generate an image of a specific style with that target pose remains an ill-posed and thus complicated problem. Most recent works treat the human pose synthesis tasks as an image spatial transformation problem using flow warping techniques. However, we observe that, due to the inherent ill-posed nature of many complicated human poses, former methods fail to generate body parts. To tackle this problem, we propose a feature-level flow attention module and an Enhancer Network. The flow attention module produces a flow attention mask to guide the combination of the flow-warped features and the structural pose features. Then, we apply the Enhancer Network to re-fine the coarse image by injecting the pose information. We present our experimental evaluation both qualitatively and quantitatively on DeepFashion, Market-1501, and Youtube dance datasets. Quantitative results show that our method has 12.995 FID at DeepFashion, 25.459 FID at Market-1501, 14.516 FID at Youtube dance datasets, which outperforms some state-of-the-arts including Guide-Pixe2Pixe, Global-Flow-Local-Attn, and CocosNet.
Pu Ge, Qiushi Huang, Xue Jing, Yule Li, Yiyong Li, Zhun Sun
WACV7
2020 Beyond Unfolding: Exact Recovery of Latent Convex Tensor Decomposition Under Reshuffling
abstract
Exact recovery of tensor decomposition (TD) methods is a desirable property in both unsupervised learning and scientific data analysis. The numerical defects of TD methods, however, limit their practical applications on real-world data. As an alternative, convex tensor decomposition (CTD) was proposed to alleviate these problems, but its exact-recovery property is not properly addressed so far. To this end, we focus on latent convex tensor decomposition (LCTD), a practically widely-used CTD model, and rigorously prove a sufficient condition for its exact-recovery property. Furthermore, we show that such property can be also achieved by a more general model than LCTD. In the new model, we generalize the classic tensor (un-)folding into reshuffling operation, a more flexible mapping to relocate the entries of the matrix into a tensor. Armed with the reshuffling operations and exact-recovery property, we explore a totally novel application for (generalized) LCTD, i.e., image steganography. Experimental results on synthetic data validate our theory, and results on image steganography show that our method outperforms the state-of-the-art methods.
Chao Li 0013, Mohammad Emtiyaz Khan, Zhun Sun, Gang Niu 0001, Bo Han 0003, Shengli Xie 0001, Qibin Zhao
AAAI3
2020 Evolutionary Topology Search for Tensor Network Decomposition
abstract
Tensor network (TN) decomposition is a promising framework to represent extremely high-dimensional problems with few parameters. However, it is challenging to search the (near-)optimal topological structures for TN decomposition, since the number of candidate solutions exponentially grows with increasing the order of a tensor. In this paper, we claim that the issue can be practically tackled by evolutionary algorithms in an affordable manner. We encode the complex topological structures into binary strings, and develop a simple genetic meta-algorithm to search the optimal topology on Hamming space. The experimental results by both synthetic and real-world data demonstrate that our method can effectively discover the ground-truth topology or even better structures with a small number of generations, and significantly boost the representational power of TN decomposition compared with well-known tensor-train (TT) or tensor-ring (TR) models.
Chao Li 0013, Zhun Sun
ICML2
2020 Application-Oblivious L7 Parsing Using Recurrent Neural Networks
abstract
Extracting fields from layer 7 protocols such as HTTP, known as L7 parsing, is the key to many critical network applications. However, existing L7 parsing techniques center around protocol specifications, thereby incurring large human efforts in specifying data format and high computational/memory costs that poorly scale with the explosive number of L7 protocols. To this end, this paper introduces a new framework namedcontent-based L7 parsing, where the content instead of the format becomes the first class citizen. Under this framework, users only need to label what content they are interested in, and the parser learns an extraction model from the users’ labeling behaviors. Since the parser is specification-independent, both the human effort and computational/memory costs can be dramatically reduced. To realize content-based L7 parsing, we propose REPLAY which builds on recurrent neural network (RNN) and addresses a series of technical challenges like large labeling overhead and slow parsing speed. We prototype REPLAY on GPUs, and show it can achieve a precision of 98% and a recall of 97%, with a throughput as high as 12Gbps for diverse extraction tasks.
Hao Li 0011, Zhengda Bian, Peng Zhang 0011, Zhun Sun, Chengchen Hu, Qiang Fu 0011, Tian Pan 0001, Jia Lv
IEEE/ACM Trans. Netw.4
2019 Guaranteed Matrix Completion Under Multiple Linear Transformations
abstract
Low-rank matrix completion (LRMC) is a classical model in both computer vision (CV) and machine learning, and has been successfully applied to various real applications. In the recent CV tasks, the completion is usually employed on the variants of data, such as "non-local" or filtered, rather than their original forms. This fact makes that the theoretical analysis of the conventional LRMC is no longer suitable in these applications. To tackle this problem, we propose a more general framework for LRMC, in which the linear transformations of the data are taken into account. We rigorously prove the identifiability of the proposed model and show an upper bound of the reconstruction error. Furthermore, we derive an efficient completion algorithm by using augmented Lagrangian multipliers and the sketching trick. In the experiments, we apply the proposed method to the classical image inpainting problem and achieve the state-of-the-art results.
Chao Li 0013, Wei He 0003, Longhao Yuan, Zhun Sun, Qibin Zhao
CVPR4
2019 Dual Residual Networks Leveraging the Potential of Paired Operations for Image Restoration
abstract
In this paper, we study design of deep neural networks for tasks of image restoration. We propose a novel style of residual connections dubbed "dual residual connection", which exploits the potential of paired operations, e.g., up- and down-sampling or convolution with large- and small-size kernels. We design a modular block implementing this connection style; it is equipped with two containers to which arbitrary paired operations are inserted. Adopting the "unraveled" view of the residual networks proposed by Veit et al., we point out that a stack of the proposed modular blocks allows the first operation in a block interact with the second operation in any subsequent blocks. Specifying the two operations in each of the stacked blocks, we build a complete network for each individual task of image restoration. We experimentally evaluate the proposed approach on five image restoration tasks using nine datasets. The results show that the proposed networks with properly chosen paired operations outperform previous methods on almost all of the tasks and datasets.
Xing Liu 0010, Masanori Suganuma, Zhun Sun, Takayuki Okatani
CVPR3
2019 Improving Head Pose Estimation with a Combined Loss and Bounding Box Margin Adjustment
abstract
We address a problem of estimating pose of a person's head from its RGB image. The employment of CNNs for the problem has contributed to significant improvement in accuracy in recent works. However, we show that the following two methods, despite their simplicity, can attain further improvement: (i) proper adjustment of the margin of bounding box of a detected face, and (ii) choice of loss functions. We show that the integration of these two methods achieve the new state-of-the-art on standard benchmark datasets for in-the-wild head pose estimation. The Tensorflow implementation of our work is available at https://github.com/MingzhenShao/HeadPose.
Mingzhen Shao, Zhun Sun, Mete Ozay, Takayuki Okatani
FG2
2019 Low-rank Embedding of Kernels in Convolutional Neural Networks under Random Shuffling
abstract
Although the convolutional neural networks (CNNs) have become popular for various image processing and computer vision tasks recently, it remains a challenging problem to reduce the storage cost of the parameters for resource-limited platforms. In the previous studies, tensor decomposition (TD) has achieved promising compression performance by embedding the kernel of a convolutional layer into a low-rank subspace. However the employment of TD is naively on the kernel or its specified variants. Unlike the conventional approaches, this paper shows that the kernel can be embedded into more general or even random low-rank subspaces. We demonstrate this by compressing the convolutional layers via randomly-shuffled tensor decomposition (RsTD) for a standard classification task using CIFAR-10. In addition, we analyze how the spatial similarity of the training data influences the low-rank structure of the kernels. The experimental results show that the CNN can be significantly compressed even if the kernels are randomly shuffled. Furthermore, the RsTD-based method yields more stable classification accuracy than the conventional TD-based methods in a large range of compression ratios.
Chao Li 0013, Zhun Sun, Jinshi Yu, Qibin Zhao
ICASSP2
2018 Feature Quantization for Defending Against Distortion of Images
abstract
In this work, we address the problem of improving robustness of convolutional neural networks (CNNs) to image distortion. We argue that higher moment statistics of feature distributions can be shifted due to image distortion, and the shift leads to performance decrease and cannot be reduced by ordinary normalization methods as observed in our experimental analyses. In order to mitigate this effect, we propose an approach base on feature quantization. To be specific, we propose to employ three different types of additional non-linearity in CNNs: i) a floor function with scalable resolution, ii) a power function with learnable exponents, and iii) a power function with data-dependent exponents. In the experiments, we observe that CNNs which employ the proposed methods obtain better performance in both generalization performance and robustness for various distortion types for large scale benchmark datasets. For instance, a ResNet-50 model equipped with the proposed method (+HPOW) obtains 6.95%, 5.26% and 5.61% better accuracy on the ILSVRC-12 classification tasks using images distorted with motion blur, salt and pepper and mixed distortions.
Zhun Sun, Mete Ozay, Yan Zhang 0055, Xing Liu 0010, Takayuki Okatani
CVPR1
2016 Design of Kernels in Convolutional Neural Networks for Image Classification
Zhun Sun, Mete Ozay, Takayuki Okatani
ECCV (7)1