Zhengxing Sun

dblp:21/2848 · DBLP profile ↗
← Back
94ranked-venue papers
1as first author
26since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 23 since 2021Artificial intelligence and machine learning · 19 · 7 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Weakly Supervised Object Detection Framework based on Classification-Localization Consistency
abstract
The inconsistency between classification and localization brings a challenge in object detection. Fully supervised object detection (FSOD) benefits from bounding-box regression networks to alleviate that, which is absent in weakly supervised object detection (WSOD). Consequently, there is a significant performance gap between the two paradigms. To bridge the performance and technical gaps between WSOD and FSOD, this paper proposes a novel weakly supervised object detection framework based on classification-localization consistency. We propose Max Score Pooling (MSP), which compels features relevant to classification to align with features relevant to localization, thereby achieving consistency between classification and localization. Additionally, we propose a Proposal Fusion Mechanism (PFM) to generate pseudo-supervision for training the bounding box regression network, further reducing the impact of classification-localization inconsistency. Extensive experiments are conducted on PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO 2017 datasets, demonstrating our framework’s superior performance.
Yihuan Zhu, Simiao Wang, Mingyu Lu, Zhengxing Sun
ICME4
2025 Text Prompted Spatiotemporal Sequence Prediction with Text-Vision Prompt Refiner and Masked Diffusion Transformers
abstract
Classical spatiotemporal sequence prediction tasks are designed to forecast future image sequences based on historical observations. However, the inherent unpredictability of future events often renders this process uncontrollable due to infinite possibilities in nature, limiting broader applicability of this technology. In this study, we explore the utilization of text prompts to constrain probabilistic space of future outcomes, resulting more controllable future prediction complying with user intent. We primarily address two critical challenges in this research setting: (i) text-vision misalignment, where embeddings extracted by text pre-trained models are not strictly aligned with visual embeddings, leading to predictions semantically irrelevant to text prompts. (ii) Spatiotemporal modeling distortion, where the fixed observation interval during training causes the model to produce unrealistic results when reasoning longer time dimensions. To tackle these issues, we propose a text-prompted spatiotemporal sequence prediction (TPS2P) model, leveraging historical observations and textual prompts to predict probabilistic future outcomes. In this model, a text-vision prompt refiner (TV-Refiner) is introduced to provide aligned textual and historical visual embeddings for integrating the denoising diffusion prediction process. Additionally, a spatiotemporal-masked diffusion transformer (StMDiT) is proposed by exploiting masked attention in constituting spatial and temporal self-attention modules within latent diffusion processes, enabling the model to observe more sequences of varying spatiotemporal patterns during training. We conduct extensive experiments on Something-Something V2 (Sthv2) and BridgeData datasets. Reported results demonstrate that our TPS2P predicts more accurate and high-quality future sequences, more user-intent compliant by textual controllability.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun
ACM Multimedia2
2025 Spatiotemporal semantic structural representation learning for image sequence prediction
Yechao Xu, Zhengxing Sun, Yunhan Sun
Neurocomputing2
2024 SasWOT: Real-Time Semantic Segmentation Architecture Search WithOut Training
abstract
In this paper, we present SasWOT, the first training-free Semantic segmentation Architecture Search (SAS) framework via an auto-discovery proxy. Semantic segmentation is widely used in many real-time applications. For fast inference and memory efficiency, Previous SAS seeks the optimal segmenter by differentiable or RL Search. However, the significant computational costs of these training-based SAS limit their practical usage. To improve the search efficiency, we explore the training-free route but empirically observe that the existing zero-cost proxies designed on the classification task are sub-optimal on the segmentation benchmark. To address this challenge, we develop a customized proxy search framework for SAS tasks to augment its predictive capabilities. Specifically, we design the proxy search space based on the some observations: (1) different inputs of segmenter statistics can be well combined; (2) some basic operators can effectively improve the correlation. Thus, we build computational graphs with multiple statistics as inputs and different advanced basis arithmetic as the primary operations to represent candidate proxies. Then, we employ an evolutionary algorithm to crossover and mutate the superior candidates in the population based on correlation evaluation. Finally, based on the searched proxy, we perform the segmenter search without candidate training. In this way, SasWOT not only enables automated proxy optimization for SAS tasks but also achieves significant search acceleration before the retrain stage. Extensive experiments on Cityscapes and CamVid datasets demonstrate that SasWOT achieves superior trade-off between accuracy and speed over several state-of-the-art techniques. More remarkably, on Cityscapes dataset, SasWOT achieves the performance of 71.3% mIoU with the speed of 162 FPS.
Lujun Li 0001, Yuli Wu 0001, Zhengxing Sun
AAAI4
2024 In-WSOD: Integrality Weakly Supervised Object Detection with Classification and Localization Consistency
Yihuan Zhu, Simiao Wang, Zhengxing Sun
ICONIP (7)3
2024 LDCNet: Long-Distance Context Modeling for Large-Scale 3D Point Cloud Scene Semantic Segmentation
abstract
Large-scale point cloud semantic segmentation is a challenging task in 3D computer vision. A key challenge is how to resolve ambiguities arising from locally high inter-class similarity. In this study, we introduce a solution by modeling long-distance contextual information to understand the scene's overall layout. The context sensitivity of previous methods is typically constrained to small blocks(e.g. 2m x 2m) and cannot be directly extended to the entire scene. For this reason, we propose Long-Distance Context Modeling Network(LDCNet). Our key insight is that keypoints are enough for inferring the layout of a scene. Therefore, we represent the entire scene using keypoints along with local descriptors and model long-distance context on these keypoints. Finally, we propagate the long-distance context information from keypoints back to non-keypoints. This allows our method to model long-distance context effectively. We conducted experiments on six datasets, demonstrating that our approach can effectively mitigate ambiguities. Our method performs well on large, irregular objects and exhibits good generalization for typical scenarios.
Shoutong Luo, Zhengxing Sun, Yi Wang 0125, Yunhan Sun
ACM Multimedia2
2024 Style creation: multiple styles transfer with incremental learning and distillation loss
Zhengxing Sun, Chengfeng Ruan
Multim. Tools Appl.2
2024 High-to-low-level feature matching and complementary information fusion for reference-based image super-resolution
Shuang Wang 0009, Zhengxing Sun, Qian Li 0014
Vis. Comput.2
2023 Image super-resolution based on self-similarity generative adversarial networks
abstract
Abstract Self‐attention has been successfully leveraged for long‐range feature‐wise similarities in deep learning super‐resolution (SR) methods. However, most of the SR methods only explore the features on the original scale, but do not take full advantage of self‐similarities features on different scales especially in generative adversarial networks (GAN). In this paper, self‐similarity generative adversarial networks (SSGAN) are proposed as the SR framework. The framework establishes the multi‐scale feature correlation by adding two modules to the generative network: downscale attention block (DAB) and upscale attention block (UAB). Specifically, DAB is designed to restore the repetitive details from the corresponding downsampled image, which achieves multi‐scale feature restoration through self‐similarity. And UAB improves the baseline up‐sampling operations and captures low‐resolution to high‐resolution feature mapping, which enhances the cross‐scale repetitive features to reconstruct the high‐resolution image. Experimental results demonstrate that the proposed SSGAN achieve better visual performance especially in the similar pattern details.
Shuang Wang 0009, Zhengxing Sun, Qian Li 0014
IET Image Process.2
2023 Fine-grained traffic video vehicle recognition based orientation estimation and temporal information
Anqi Hu, Zhengxing Sun, Qian Li 0014, Yechao Xu, Yihuan Zhu
Multim. Tools Appl.2
2022 Multi-view 3D Reconstruction from Video with Transformer
abstract
Multi-view 3D reconstruction is the base for many other applications in computer vision. Video provides multi-view images and temporal information, which can help us better complete the reconstruction goal. Redundant information handling in video and multi-view feature extraction and fusion become the key issues in the shape prior extraction for reconstruction. In this paper, inspired by the recent great success in Transformer models, we propose a transformer-based 3D reconstruction network. We formulate the multi-view 3D reconstruction into three parts: frame encoder, fusion module, and shape decoder. We apply several special used tokens and perform the fusion progressively in the encoder phase, called patch-level progressive fusion module. These tokens describe which part of the object the frame should focus on and the local structural detail progressively. Then we further design a transformer fusion module to aggregate the structure information. Finally, multi-head attention is utilized to build the transformer-based decoder to reuse the shallow features from encoder. In experiments not only can ours method achieve competitive performance, but it also has low model complexity and computation cost.
Yijie Zhong 0001, Zhengxing Sun, Yunhan Sun, Shoutong Luo
ICIP2
2022 Learning Semantic Segmentation on Unlabeled Real-World Indoor Point Clouds via Synthetic Data
abstract
The data-hungry nature of deep learning and the high cost of annotating point-level labels for point clouds make it difficult to apply semantic segmentation methods to unlabeled real-world indoor scenes. Therefore, label-efficient point cloud segmentation has become a promising research topic. We noticed that the online housing design platforms can provide a large number of synthetic indoor 3D scenes, which are created with semantic labels. In this paper, we propose to learn semantic segmentation on synthetic point clouds and adapt the model for unlabeled real-world data. The main challenge is that directly using models trained on synthetic data for real-world data produces poor results due to the large domain gap between synthetic and real-world data. We design a point cloud style transfer network and a feature discrimination network to reduce the domain gap in both the input space and the feature space. Experiments show that our approach significantly improves the performance on real-world data for models learned from synthetic data.
Youcheng Song, Zhengxing Sun, Yunjie Wu, Yunhan Sun, Shoutong Luo, Qian Li 0014
ICPR2
2022 Weakly Supervised Fine-grained Recognition based on Combined Learning for Small Data and Coarse Label
abstract
Learning with weak supervision already becomes one of the research trends in fine-grained image recognition. These methods aim to learn feature representation in the case of less manual cost or expert knowledge. Most existing weakly supervised methods are based on incomplete annotation or inexact annotation, which is difficult to perform well limited by supervision information. Therefore, using these two kind of annotations for training at the same time could mine more relevance while the annotating burden will not increase much. In this paper, we propose a combined learning framework by coarse-grained large data and fine-grained small data for weakly supervised fine-grained recognition. Combined learning contains two significant modules: 1) a discriminant module, which maintains the structure information consistent between coarse label and fine label by attention map and part sampling, 2) a cluster division strategy, which mines the detail differences between fine categories by feature subtraction. Experiment results show that our method outperforms weakly supervised methods and achieves the performance close to fully supervised methods in CUB-200-2011 and Stanford Cars datasets.
Anqi Hu, Zhengxing Sun, Qian Li 0014
ICMR2
2022 Active Patterns Perceived for Stochastic Video Prediction
abstract
Predicting future scenes based on historical frames is challenging, especially when it comes to the complex uncertainty in nature. We observe that there is a divergence between spatial-temporal variations of active patterns and non-active patterns in a video, where these patterns constitute visual content and the former ones implicate more violent movement. This divergence enables active patterns the higher potential to act with more severe future uncertainty. Meanwhile, the existence of non-active patterns provides an opportunity for machines to examine some underlying rules with a mutual constraint between non-active patterns and active patterns. In order to solve this divergence, we provide a method called active patterns-perceived stochastic video prediction (ASVP) which allows active patterns to be perceived by neural networks during training. Our method starts with separating active patterns along with non-active ones from a video. Then, both scene-based prediction and active pattern-perceived prediction are conducted to respectively capture the variations within the whole scene and active patterns. Specially for active pattern-perceived prediction, a conditional generative adversarial network (CGAN) is exploited to model active patterns as conditions, with a variational autoencoder (VAE) for predicting the complex dynamics of active patterns. Additionally, a mutual constraint is designed to improve the learning procedure for the network to better understand underlying interacting rules among these patterns. Extensive experiments are conducted on both KTH human action and BAIR action-free robot pushing datasets with comparison to state-of-the-art works. Experimental results demonstrate the competitive performance of the proposed method as we expected. The released code and models are at https://github.com/tolearnmuch/ASVP.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun, Shoutong Luo
ACM Multimedia2
2022 Category-Sensitive Incremental Learning for Image-Based 3D Shape Reconstruction
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun
MMM (1)2
2022 Resolution-switchable 3D Semantic Scene Completion
abstract
Abstract Semantic scene completion (SSC) aims to recover the complete geometric structure as well as the semantic segmentation results from partial observations. Previous works could only perform this task at a fixed resolution. To handle this problem, we propose a new method that can generate results at different resolutions without redesigning and retraining. The basic idea is to decouple the direct connection between resolution and network structure. To achieve this, we convert feature volume generated by SSC encoders into a resolution adaptive feature and decode this feature via point. We also design a resolution‐adapted point sampling strategy for testing and a category‐based point sampling strategy for training to further handle this problem. The encoder of our method can be replaced by existing SSC encoders. We can achieve better results at other resolutions while maintaining the same accuracy as the original resolution results. Code and data are available at https://github.com/lstcutong/ReS-SSC .
Shoutong Luo, Zhengxing Sun, Yunhan Sun, Yi Wang 0125
Comput. Graph. Forum2
2022 Video supervised for 3D reconstruction from single image
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun
Multim. Tools Appl.2
2022 Multilayered stitch generating for random-needle embroidery
Zhengxing Sun
Vis. Comput.2
2022 Learning indoor point cloud semantic segmentation from image-level labels
Youcheng Song, Zhengxing Sun, Qian Li 0014, Yunjie Wu, Yunhan Sun, Shoutong Luo
Vis. Comput.2
2021 Shape-Pose Ambiguity in Learning 3D Reconstruction from Images
Yunjie Wu, Zhengxing Sun, Youcheng Song, Yunhan Sun, Yijie Zhong 0001
AAAI2
2021 AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer
abstract
Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distributions, or adaptively normalize deep content feature according to the style such that their global statistics are matched. Although effective, leaving shallow feature unexplored and without locally considering feature statistics, they are prone to unnatural output with unpleasing local distortions. To alleviate this problem, in this paper, we propose a novel attention and normalization module, named Adaptive Attention Normalization (AdaAttN), to adaptively perform attentive normalization on per-point basis. Specifically, spatial attention score is learnt from both shallow and deep features of content and style images. Then perpoint weighted statistics are calculated by regarding a style feature point as a distribution of attention-weighted output of all style feature points. Finally, the content feature is normalized so that they demonstrate the same local feature statistics as the calculated per-point weighted style feature statistics. Besides, a novel local feature loss is derived based on AdaAttN to enhance local visual quality. We also extend AdaAttN to be ready for video style transfer with slight modifications. Experiments demonstrate that our method achieves state-of-the-art arbitrary image/video style transfer. Codes and models are available on https://github.com/wzmsltw/AdaAttN.
Songhua Liu, Dongliang He, Fu Li 0003, Xin Li 0106, Zhengxing Sun, Qian Li 0014, Errui Ding
ICCV7
2021 Semi-supervised point cloud segmentation using self-training with label confidence prediction
Zhengxing Sun, Yunjie Wu, Youcheng Song
Neurocomputing2
2021 Dyeing creation: a textile pattern discovery and fabric image generation method
Shuang Wang 0009, Zhengxing Sun
Multim. Tools Appl.2
2021 Crowd aware summarization of surveillance videos by deep reinforcement learning
Zhengxing Sun
Multim. Tools Appl.2
2021 A structural-constraint 3D point clouds segmentation adversarial method
Zhengxing Sun
Vis. Comput.2
2021 A self-supervised method of single-image depth estimation by feeding forward information using max-pooling layers
Yunhan Sun, Suqin Bai, Zhengxing Sun, Zhaohui Tian
Vis. Comput.4
2020 Slicenet: Slice-Wise 3D Shapes Reconstruction from Single Image
abstract
3D object reconstruction from a single image is a highly ill-posed problem, requiring strong prior knowledge of 3D shapes. Deep learning methods are popular for this task. Especially, most works utilized 3D deconvolution to generate 3D shapes. However, the resolution of results is limited by the high resource consumption of 3D deconvolution. In this paper, we propose SliceNet, sequentially generating 2D slices of 3D shapes with shared 2D deconvolution parameters. To capture relations between slices, the RNN is also introduced. Our model has three main advantages: First, the introduction of RNN allows the CNN to focus more on local geometry details,improving the results’ fine-grained plausibility. Second, replacing 3D deconvolution with 2D deconvolution reducs much consumption of memory, enabling higher resolution of final results. Third, an slice-aware attention mechanism is designed to provide dynamic information for each slice’s generation, which helps modeling the difference between multiple slices, making the learning process easier. Experiments on both synthesized data and real data illustrate the effectiveness of our method.
Yunjie Wu, Zhengxing Sun, Youcheng Song, Yunhan Sun
ICASSP2
2020 Stable Video Style Transfer Based on Partial Convolution with Depth-Aware Supervision
abstract
As a very important research issue in digital media art, neural learning based video style transfer has attracted more and more attention. A lot of recent works import optical flow method to original image style transfer framework to preserve frame-coherency and prevent flicker. However, these methods highly rely on paired video datasets of content video and stylized video, which are often difficult to obtain. Another limitation of existing methods is that while maintaining inter-frame coherency, they will introduce strong ghosting artifacts. In order to address these problems, this paper has following contributions: (1).presents a novel training framework for video style transfer without dependency on video dataset of target style; (2).firstly focuses on the ghosting problem existing in most previous works and uses partial convolution-based strategy to utilize inter-frame context and correlation, together with additional depth loss as a constrain to the generated frames to suppress ghosting artifacts and preserve stability at the same time. Extensive experiments demonstrate that our method can produce natural and stable video frames with target style. Qualitative and quantitative comparisons also show that the proposed approach outperforms previous works in terms of overall image quality and inter-frame stability. To facilitate future research, we publish our experiment code at \urlhttps://github.com/Huage001/Artistic-Video-Partial-Conv-Depth-Loss.
Songhua Liu, Shoutong Luo, Zhengxing Sun
ACM Multimedia4
2020 Single View Depth Estimation via Dense Convolution Network with Self-supervision
Yunhan Sun, Suqin Bai, Zhengxing Sun
MMM (2)5
2020 Structure-aware shape correspondence network for 3D shape synthesis
Xufeng Lang, Zhengxing Sun
Comput. Aided Geom. Des.2
2020 Qualitative photo collage by quartet analysis and active learning
Yuan Gan, Yan Zhang 0057, Zhengxing Sun, Hao (Richard) Zhang
Comput. Graph.3
2020 DFR: Differentiable Function Rendering for Learning 3D Generation from Images
abstract
Abstract Learning‐based 3D generation is a popular research field in computer graphics. Recently, some works adapted implicit function defined by a neural network to represent 3D objects and have become the current state‐of‐the‐art. However, training the network requires precise ground truth 3D data and heavy pre‐processing, which is unrealistic. To tackle this problem, we propose the DFR, a differentiable process for rendering implicit function representation of 3D objects into 2D images. Briefly, our method is to simulate the physical imaging process by casting multiple rays through the image plane to the function space, aggregating all information along with each ray, and performing a differentiable shading according to every ray's state. Some strategies are also proposed to optimize the rendering pipeline, making it efficient both in time and memory to support training a network. With DFR, we can perform many 3D modeling tasks with only 2D supervision. We conduct several experiments for various applications. The quantitative and qualitative evaluations both demonstrate the effectiveness of our method.
Yunjie Wu, Zhengxing Sun
Comput. Graph. Forum2
2020 Saliency based multiple object cosegmentation by ensemble MIML learning
Bo Li 0061, Zhengxing Sun, Shuang Wang 0009
Multim. Tools Appl.2
2020 Progressive decomposition: a method of coarse-to-fine image parsing using stacked networks
Yunhan Sun, Jiagao Hu, Zhengxing Sun
Multim. Tools Appl.4
2019 SuperVAE: Superpixelwise Variational Autoencoder for Salient Object Detection
abstract
Image saliency detection has recently witnessed rapid progress due to deep neural networks. However, there still exist many important problems in the existing deep learning based methods. Pixel-wise convolutional neural network (CNN) methods suffer from blurry boundaries due to the convolutional and pooling operations. While region-based deep learning methods lack spatial consistency since they deal with each region independently. In this paper, we propose a novel salient object detection framework using a superpixelwise variational autoencoder (SuperVAE) network. We first use VAE to model the image background and then separate salient objects from the background through the reconstruction residuals. To better capture semantic and spatial contexts information, we also propose a perceptual loss to take advantage from deep pre-trained CNNs to train our SuperVAE network. Without the supervision of mask-level annotated data, our method generates high quality saliency results which can better preserve object boundaries and maintain the spatial consistency. Extensive experiments on five wildly-used benchmark datasets show that the proposed method achieves superior or competitive performance compared to other algorithms including the very recent state-of-the-art supervised methods.
Bo Li 0061, Zhengxing Sun
AAAI2
2019 Qualitative Organization of Photo Collections via Quartet Analysis and Active Learning
Yuan Gan, Yan Zhang 0057, Zhengxing Sun, Hao (Richard) Zhang
Graphics Interface3
2019 Two-B-real Net: Two-branch Network for Real-time Salient Object Detection
abstract
As a hot topic in computer vision, recent researches on salient object detection (SOD) have focused on using the over-designed deep convolutional neural networks (CNNs) to improve the detection accuracy. However, these complex architectures constraint themselves to low speed and drag them on wide-ranging applications. In this paper, we simplify the over-designed networks and propose the Two-Branch Network for Real-time Salient Object Detection (Two-B-Real Net). Particularly, the Perceptual Branch and the Objectness Branch in our network can efficiently capture detailed information and distinctive objectness simultaneously. And we also design novel attention mechanisms to guide the network to focus on most saliency-related features and generate more accurate results. Extensive evaluations show that the proposed algorithm achieves the leading accuracy performance with real-time speed (125fps) which is significantly faster than the existing methods.
Bo Li 0061, Zhengxing Sun, Lv Tang, Anqi Hu
ICASSP2
2019 PPSAN: Perceptual-aware 3D Point Cloud Segmentation via Adversarial Learning
abstract
Point cloud segmentation is a key problem of 3D multimedia signal processing. Existing methods usually use a single network structure which is trained by a per-point loss. These methods mainly focus on the geometric similarity between the prediction results and the ground truth, ignoring visual perception difference. In this paper, we present a segmentation adversarial network to overcome the drawbacks above. A discriminator is introduced to provide a perceptual loss to increase the rationality judgment of prediction and guide the further optimization of the segmentator. In order to perfectly capture the structural information of parts in the same category of objects, condition settings are employed to add a global constraint. Experimental results show the proposed methods can correct the common errors in point cloud segmentation and obtain more accurate and better segmentation of visual perceptual.
Hongyan Li 0007, Zhengxing Sun, Yunjie Wu, Bo Li 0061
ICASSP2
2019 Group-Wise Deep Object Co-Segmentation With Co-Attention Recurrent Neural Network
abstract
Effective feature representations which should not only express the images individual properties, but also reflect the interaction among group images are essentially crucial for real-world co-segmentation. This paper proposes a novel end-to-end deep learning approach for group-wise object co-segmentation with a recurrent network architecture. Specifically, the semantic features extracted from a pre-trained CNN of each image are first processed by single image representation branch to learn the unique properties. Meanwhile, a specially designed Co-Attention Recurrent Unit (CARU) recurrently explores all images to generate the final group representation by using the co-attention between images, and simultaneously suppresses noisy information. The group feature which contains synergetic information is broadcasted to each individual image and fused with multi-scale fine-resolution features to facilitate the inferring of co-segmentation. Moreover, we propose a groupwise training objective to utilize the co-object similarity and figure-ground distinctness as the additional supervision. The whole modules are collaboratively optimized in an end-to-end manner, further improving the robustness of the approach. Comprehensive experiments on three benchmarks can demonstrate the superiority of our approach in comparison with the state-of-the-art methods.
Zhengxing Sun, Qian Li 0014, Yunjie Wu, Anqi Hu
ICCV2
2019 Detecting Robust Co-Saliency with Recurrent Co-Attention Neural Network
abstract
Effective feature representations which should not only express the images individual properties, but also reflect the interaction among group images are essentially crucial for robust co-saliency detection. This paper proposes a novel deep learning co-saliency detection approach which simultaneously learns single image properties and robust group feature in a recurrent manner. Specifically, our network first extracts the semantic features of each image. Then, a specially designed Recurrent Co-Attention Unit (RCAU) will explore all images in the group recurrently to generate the final group representation using the co-attention between images, and meanwhile suppresses noisy information. The group feature which contains complementary synergetic information is later merged with the single image features which express the unique properties to infer robust co-saliency. We also propose a novel co-perceptual loss to make full use of interactive relationships of whole images in the training group as the supervision in our end-to-end training process. Extensive experimental results demonstrate the superiority of our approach in comparison with the state-of-the-art methods.
Bo Li 0061, Zhengxing Sun, Lv Tang, Yunhan Sun
IJCAI2
2019 Co-saliency Detection Based on Hierarchical Consistency
abstract
As an interesting and emerging topic, co-saliency detection aims at discovering common and salient objects in a group of related images, which is useful to variety of visual media applications. Although a number of approaches have been proposed to address this problem, many of them are designed with the misleading assumption, suboptimal image representation, or heavy supervision cost and thus still suffer from certain limitations, which reduces their capability in the real-world scenarios. To alleviate these limitations, we propose a novel unsupervised co-saliency detection method, which successively explores the hierarchical consistency in the image group including background consistency, high-level and low-level objects consistency in a unified framework. We first design a novel superpixel-wise variational autoencoder (SVAE) network to precisely distinguish the salient objects from the background collection based on the reconstruction errors. Then, we propose a two-stage clustering strategy to explore the multi-level salient objects consistency by using high-level and low-level features separately. Finally, the co-saliency results are refined by applying a CRF based refinement method with the multi-level salient objects consistency. Extensive experiments on three widely datasets show that our method achieves superior or competitive performance compared to the state-of-the-art methods.
Zhengxing Sun, Qian Li 0014
ACM Multimedia2
2019 Coarse-to-fine segmentation for indoor scenes with progressive supervision
Youcheng Song, Zhengxing Sun, Yunjie Wu, Hongyan Li 0007
Comput. Aided Geom. Des.2
2019 Direction-aware neural style transfer with texture enhancement
Zhengxing Sun, Yan Zhang 0007, Qian Li 0014
Neurocomputing2
2019 StitchGeneration: Modeling and creation of random-needle embroidery based on Markov chain model
Zhengxing Sun
Multim. Tools Appl.2
2018 Progressive Refinement: A Method of Coarse-to-Fine Image Parsing Using Stacked Network
abstract
To parse images into fine-grained semantic parts, the complex fine-grained elements will put it in trouble when using off-the-shelf semantic segmentation networks. In this paper, for image parsing task, we propose to parse images from coarse to fine with progressively refined semantic classes. It is achieved by stacking the segmentation layers in a segmentation network several times. The former segmentation module parses images at a coarser-grained level, and the result will be feed to the following one to provide effective contextual clues for the finer-grained parsing. To recover the details of small structures, we add skip connections from shallow layers of the network to fine-grained parsing modules. As for the network training, we merge classes in groundtruth to get coarse-to-fine label maps, and train the stacked network with these hierarchical supervision end-to-end. Our coarse-to-fine stacked framework can be injected into many advanced neural networks to improve the parsing results. Extensive evaluations on several public datasets including face parsing and human parsing well demonstrate the superiority of our method.
Jiagao Hu, Zhengxing Sun, Yunhan Sun
ICME2
2018 Direction-aware Neural Style Transfer
abstract
Neural learning methods have been shown to be effective in style transfer. These methods, which are called NST, aim to synthesize a new image that retains the high-level structure of a content image while keeps the low-level features of a style image. However, these models using convolutional structures only extract local statistical features of style images and semantic features of content images. Since the absence of low-level features in the content image, these methods would synthesize images that look unnatural and full of traces of machines. In this paper, we find that direction, that is, the orientation of each painting stroke, can capture the soul of image style preferably and thus generates much more natural and vivid stylizations. According to this observation, we propose a Direction-aware Neural Style Transfer (DaNST) with two major innovations. First, a novel direction field loss is proposed to steer the direction of strokes in the synthesized image. And to build this loss function, we propose novel direction field loss networks to generate and compare the direction fields of content image and synthesized image. By incorporating the direction field loss in neural style transfer, we obtain a new optimization objective. Through minimizing this objective, we can produce synthesized images that better follow the direction field of the content image. Second, our method provides a simple interaction mechanism to control the generated direction fields, and further control the texture direction in synthesized images. Experiments show that our method outperforms state-of-the-art in most styles such as oil painting and mosaic.
Zhengxing Sun, Weihang Yuan
ACM Multimedia2
2018 Iterative Active Classification of Large Image Collection
Mofei Song, Zhengxing Sun, Bo Li 0061, Jiagao Hu
MMM (1)2
2018 ShapeCreator: 3D Shape Generation from Isomorphic Datasets Based on Autoencoder
Yunjie Wu, Zhengxing Sun, Youcheng Song, Hongyan Li 0007
MMM (2)2
2018 Stitch-Based Image Stylization for Thread Art Using Sparse Modeling
Ke-Wei Yang 0002, Zhengxing Sun, Shuang Wang 0009, Bo Li 0061
MMM (1)2
2018 3D shape segmentation via shape fully convolutional networks
Pengyu Wang 0004, Yuan Gan, Panpan Shui, Fenggen Yu, Yan Zhang 0057, Song-Le Chen, Zhengxing Sun
Comput. Graph.7
2018 Corrigendum to "3D shape segmentation via shape fully convolutional networks" [Computers & Graphics 70 (2018) 128-139]
Pengyu Wang 0004, Yuan Gan, Panpan Shui, Fenggen Yu, Yan Zhang 0057, Song-Le Chen, Zhengxing Sun
Comput. Graph.7
2018 3D shape segmentation via shape fully convolutional networks
Pengyu Wang 0004, Yuan Gan, Panpan Shui, Fenggen Yu, Yan Zhang 0057, Song-Le Chen, Zhengxing Sun
Comput. Graph.7
2018 Accumulative image categorization: a personal photo classification method for progressive collection
Jiagao Hu, Zhengxing Sun, Yunhan Sun
Multim. Tools Appl.2
2018 Paint with stitches: a style definition and image-based rendering method for random-needle embroidery
Ke-Wei Yang 0002, Zhengxing Sun
Multim. Tools Appl.2
2018 3D reconstruction framework via combining one 3D scanner and multiple stereo trackers
Zhengxing Sun, Suqin Bai
Vis. Comput.2
2017 Active Classification of Large 3D Shape Collection
abstract
To efficiently and accurately classify a large 3D shape collection, this paper proposes a novel interactive system by incorporating active learning, online learning and user intervention. Given a shape collection, our system iteratively alternates the interactive annotation and verification until all the shapes are classified. The main advantage is that it provides faster interactive classification rates than alternative approaches. Our system achieves this goal by a unified active learning algorithm that selects the shapes to be annotated or verified, which requires a probability model for simulating the time cost of human input during manual intervention. After manually classifying these selected shapes, we use an extended soft confidence-weighted learning method to update the classifier incrementally and efficiently for the subsequent active selection and shape classification in turn. Experimental results demonstrated the effectiveness of the proposed method.
Mofei Song, Zhengxing Sun
ICTAI2
2017 Online User Modeling for Interactive Streaming Image Classification
Jiagao Hu, Zhengxing Sun, Bo Li 0061, Ke-Wei Yang 0002
MMM (2)2
2017 Unsupervised Multiple Object Cosegmentation via Ensemble MIML Learning
Weichen Yang, Zhengxing Sun, Bo Li 0061, Jiagao Hu, Ke-Wei Yang 0002
MMM (2)2
2017 Accumulative categorization: Online 3D shape classification for progressive collections
Mofei Song, Zhengxing Sun, Hongyan Li 0007
Graph. Model.2
2017 Iterative samples labeling for sketch recognition
Kai Liu 0022, Zhengxing Sun, Mofei Song, Bo Li 0061
Multim. Tools Appl.2
2017 Image palette: painting style transfer via brushstroke control synthesis
Zheng Miao, Yan Zhang 0007, Zhibin Zheng, Zhengxing Sun
Multim. Tools Appl.4
2016 PicMarker: Data-Driven Image Categorization Based on Iterative Clustering
Jiagao Hu, Zhengxing Sun, Bo Li 0061, Shuang Wang 0009
ACCV (4)2
2016 Relevance feedback for human motion retrieval using a boosting approach
Song-Le Chen, Zhengxing Sun, Yan Zhang 0007, Qian Li 0014
Multim. Tools Appl.2
2016 Dynamic node selection in camera networks based on approximate reinforcement learning
Qian Li 0014, Zhengxing Sun, Song-Le Chen, Shi-ming Xia
Multim. Tools Appl.2
2016 An improved artificial bee colony algorithm based on the strategy of global reconnaissance
Zhengxing Sun, Junlou Li, Mofei Song, Xufeng Lang
Soft Comput.2
2016 Large-scale three-dimensional measurement based on LED marker tracking
Zhengxing Sun
Vis. Comput.2
2015 Model-driven indoor scenes modeling from a single image
Zicheng Liu 0005, Yan Zhang 0007, Wentao Wu 0003, Kai Liu 0022, Zhengxing Sun
Graphics Interface5
2015 Scalable Organization of Collections of Motion Capture Data via Quantitative and Qualitative Analysis
abstract
This paper proposes a scalable method for organizing the collection of motion capture data for overview and exploration, and it mainly addresses three core problems, including data abstraction, neighborhood construction and data visualization. To alleviate the contradiction between limited visual space and the ever-increasing size of real-word datasets, hierarchical affinity propagation (HAP) is adopt to perform data abstraction on low-level pose features to generate multi-layers of data aggregations in consistent with coarse to fine abstraction levels of human cognition. To construct a meaningful neighborhood for user choosing a browsing path and positioning themselves, quartet analysis-based phylogenetic tree is created upon high-level pose features to produce more reliable neighbors for different aggregations of the specific abstraction level. To provide a convenient interactive environment for user navigation, a phylogenetic tree-centric visualization strategy in three-dimensional space is present. Experimental results on HDM05 motion capture dataset verify the effectiveness of the proposed method.
Song-Le Chen, Zhengxing Sun, Yan Zhang 0007
ICMR2
2015 Online 3D Shape Segmentation by Blended Learning
Fei-qian Zhang, Zhengxing Sun, Mofei Song, Xufeng Lang
MMM (1)2
2015 Progressive 3D shape segmentation using online learning
Fei-qian Zhang, Zhengxing Sun, Mofei Song, Xufeng Lang
Comput. Aided Des.2
2015 Iterative 3D shape classification by online metric learning
Mofei Song, Zhengxing Sun, Kai Liu 0022, Xufeng Lang
Comput. Aided Geom. Des.2
2015 Singe image-based data-driven indoor scenes modeling
Yan Zhang 0007, Zicheng Liu 0005, Zheng Miao, Wentao Wu 0003, Kai Liu 0022, Zhengxing Sun
Comput. Graph.6
2014 A model synthesis method based on single building facade
Yan Zhang 0007, Wentao Wu 0003, Mofei Song, Zhengxing Sun
Graph. Model.5
2014 A controllable stitch layout strategy for random needle embroidery
abstract
Random needle embroidery (RNE) is a graceful art enrolled in the world intangible cultural heritage. In this paper, we study the stitch layout problem and propose a controllable stitch layout strategy for RNE. Using our method, a user can easily change the layout styles by adjusting several high-level layout parameters. This approach has three main features: firstly, a stitch layout rule containing low-level stitch attributes and high-level layout parameters is designed; secondly, a stitch neighborhood graph is built for each region to model the spatial relationship among stitches; thirdly, different stitch attributes (orientations, lengths, and colors) are controlled using different reaction-diffusion processes based on a stitch neighborhood graph. Moreover, our method supports the user in changing the stitch orientation layout by drawing guide curves interactively. The experimental results show its capability for reflecting various stitch layout styles and flexibility for user interaction.
Zhengxing Sun, Ke-Wei Yang 0002
J. Zhejiang Univ. Sci. C2
2014 Generative tracking of 3D human motion in latent space by sequential clonal selection algorithm
Zhengxing Sun
Multim. Tools Appl.2
2013 3D Shapes Co-segmentation by Combining Fuzzy C-Means with Random Walks
abstract
Co-segmentation of 3D shapes has been receiving increasing attention, and treated as clustering problem in a descriptor space by a few unsupervised approaches to achieve proper co-segmentation of shapes with large variability. However, most of the existing algorithms are performed on segment level and heavily dependent on the per-object segmentation. Accordingly, we propose a co-segmentation method based on combination of Fuzzy C-Means (FCM) and Random Walks together. The novelty of our method is twofold. As an efficient soft clustering algorithm, FCM is firstly used to cluster directly all the facets in the set in terms of their shape descriptors. The clusters of facets are created as candidates of the consistent parts of shapes. Random Walks model is then incorporated into the iterations of FCM clustering to adjust the assignment of facets in each candidate according to the minima rule of shape segmentation. The results of co-segmentation are refined through the iterations of FCM until its convergence conditions are satisfied. Experiments prove that the method proposed in this paper can not only get more stable results without per-object segmentation, but also improve the accuracy of co-segmentation.
Fei-qian Zhang, Zhengxing Sun, Mofei Song, Xufeng Lang, Hai Yan
CAD/Graphics2
2013 Best view selection of 3D models based on unsupervised feature learning and discrimination ability
abstract
In this poster, an approach for best view selection of 3D models is proposed, which is based on the framework that formulates the selection as a problem of evaluating views' discrimination ability. Firstly, different views' features are extracted by unsupervised feature learning. Then classifiers are trained to evaluate each view's discrimination ability. A view with the best classifier has the best discrimination ability, and it is chosen as the best view of the 3D model. At last, experiments show that 3D models of same class have similar best views.
Zhengxing Sun, Mofei Song, Yejia Zhang
VINCI2
2013 Intent-driven model synthesis
abstract
This paper presents an intent-driven model synthesis method. The method introduces an interactive straight prismatic construction space to realize the structure and shape variation of the example model simultaneously. The construction space defines the global size and the local shape feature of the desired model. Users can draw a closed curve and some skeleton lines by a sketch-based interface to design the construction space. Our algorithm first uses a quadrangulation algorithm to create a subdivision plane with the same contour as the closed curve. And the drawn skeleton lines control the local orientation of split units in the plane. Then it creates the construction space by sweeping the subdivision plane. Finally, it fills the construction space with the deformed model pieces while maintaining the generalized adjacent constraints, which are defined according to the example model. We demonstrate the effectiveness of the approach on large-scale complex models such as architecture, mountains.
Mofei Song, Fei-qian Zhang, Zhengxing Sun, Yan Zhang 0007
VINCI3
2013 Synthesis of 3D models by Petri net
abstract
This paper presents a synthesis method for 3D models using Petri net. Feature structure units from the example model are extracted, along with their constraints, through structure analysis, to create a new model using an inference method based on Petri net. Our method has two main advantages: first, 3D model pieces are delineated as the feature structure units and Petri net is used to record their shape features and their constraints in order to outline the model, including extending and deforming operations; second, a construction space generating algorithm is presented to convert the curve drawn by the user into local shape controlling parameters, and the free form deformation (FFD) algorithm is used in the inference process to deform the feature structure units. Experimental results showed that the proposed method can create large-scale complex scenes or models and allow users to effectively control the model result.
Mofei Song, Zhengxing Sun, Yan Zhang 0007, Fei-qian Zhang
J. Zhejiang Univ. Sci. C2
2013 Extracting 3D model feature lines based on conditional random fields
abstract
We propose a 3D model feature line extraction method using templates for guidance. The 3D model is first projected into a depth map, and a set of candidate feature points are extracted. Then, a conditional random fields (CRF) model is established to match the sketch points and the candidate feature points. Using sketch strokes, the candidate feature points can then be connected to obtain the feature lines, and using a CRF-matching model, the 2D image shape similarity features and 3D model geometric features can be effectively integrated. Finally, a relational metric based on shape and topological similarity is proposed to evaluate the matching results, and an iterative matching process is applied to obtain the globally optimized model feature lines. Experimental results showed that the proposed method can extract sound 3D model feature lines which correspond to the initial sketch template.
Yaoye Zhang, Zhengxing Sun, Kai Liu 0022, Mofei Song, Fei-qian Zhang
J. Zhejiang Univ. Sci. C2
2012 Image-based irregular needling embroidery rendering
abstract
Irregular needling embroidery has a very high artistic value as a Changzhou craftwork. Artists use stitches of rich colors and apply different types of stitch categories from course to fine on multiple layers to express objects. This paper proposes an irregular needling embroidery rendering method based on image. We first decompose the image into different regions, vector fields and image detail information to guild the subsequent simulation process. Then we hierarchically render the image with a series of layers and propose a multilayer rendering technique. Additionally, we build a stitch dictionary and present a method to simulate the threads. Experimental results demonstrate that the method proposed in this paper transfers reference images quickly and efficiently into their irregular needling embroidery simulations.
Ke-Wei Yang 0002, Zhengxing Sun
VINCI3
2009 Semi-automatic Roof Reconstruction
abstract
A semi-automatic 3D roof reconstruction method is proposed in this paper. It consists of two components: automatic recognition of 2D plane drawings and interactively “pulling” or “pushing” the recognized results. Only a limited number of reconstruction operations are needed to generate various types of 3D roofs, making the method efficient.
Feng Su, Zhengxing Sun
ICDAR4
2009 On-line hand-drawn electric circuit diagram recognition using 2D dynamic programming
Guihuan Feng, Christian Viard-Gaudin, Zhengxing Sun
Pattern Recognit.3
2008 Texture synthesis based on Direction Empirical Mode Decomposition
Yan Zhang 0007, Zhengxing Sun
Comput. Graph.2
2008 Sketch retrieval and relevance feedback with biased SVM classification
Shuang Liang 0001, Zhengxing Sun
Pattern Recognit. Lett.2
2007 A Sketch-Based Cooperative Diagramming Tool for Conceptual Design
abstract
A number of diagrams are used for representing design concepts and constructing function structure in the stage of conceptual design. This paper considers sketches as the language of communication in cooperative conceptual design and presents a tool that enables distributed participants to cooperatively construct diagrams in the form of freehand sketches. The tool is designed and implemented as a multi-agent system, which contains a set of agents to fulfill the different requirements, such as sketch-based interaction, sketch recognition, communication and coordination. The operating strategies of these agents are also addressed, including communication between agents, sketch recognition algorithm, semantics consistency maintenance. The proposed approach is implemented in flowchart drawing domain and a runtime example is given in this paper that shows satisfactory results.
Zhengxing Sun
CAD/Graphics2
2006 COSINE: A Sketch-Based Interactive Environment for Cooperative Design
abstract
Sketch plays an important role in the brainstorming of ideas. During conceptual design stage, designers use generally 2D graphical tools such as pencil and paper to draw a lot of sketchy documents for conveying ideas and communicating with other designers. Furthermore, conceptual design for complicated system involves a group of geographically distributed participants. However, most of current graphical design tools do not support such kind of collaborative sketching. This paper introduces a novel sketch-based interactive environment (COSINE) for cooperative design. Three key problems are addressed in this research: (1) how to represent and store sketch document for sketch-based cooperative design; (2) how does the system infer the designers' intention and help them to finish the design process; (3) how to implement the communication between client and server. Solutions to these problems are proposed in this paper. A prototype system for cooperative UML class diagram design based on COSINE is also outlined
Zhengxing Sun, Wentao Zheng
CSCWD2
2005 Informal User Interface for Graphical Computing
Zhengxing Sun
ACII1
2005 An Online Multi-stroke Sketch Recognition Method Integrated with Stroke Segmentation
Jianfeng Yin, Zhengxing Sun
ACII2
2004 An online composite graphics recognition approach based on matching of spatial relation graphs
Xiaogang Xu 0004, Zhengxing Sun, Binbin Peng, Xiangyu Jin, Wenyin Liu
Int. J. Document Anal. Recognit.2
2004 An Svm-Based Incremental Learning Algorithm For User Adaptation Of Sketch Recognition
abstract
User adaptation is a critical problem in the design of human-computer interaction systems. Many pattern recognition problems, such as handwriting/sketching recognition and speech recognition, are user dependent, since different users' handwritings, drawing styles, and accents are different. Therefore, the classifiers for these problems should provide the functionality of user adaptation so as to let each particular user experience better recognition accuracy according to his input habit/style. However, the user adaptation functionality requires the classifiers to have the incremental learning ability, by which the classifiers can adapt to the user quickly without too much computation cost. In this paper, an SVM-based incremental learning algorithm is presented to solve this problem for sketch recognition. Our algorithm utilizes only the support vectors instead of all the historical samples, and selects some important samples from all newly added samples as training data. The importance of a sample is measured according to its distance to the hyper-plane of the SVM classifier. Theoretical analysis, experimentation, and evaluation of our algorithm in our online graphics recognition system SmartSketchpad, are presented to show the effectiveness of this algorithm. According to our experiments, this algorithm can reduce both the training time and the required storage space for the training dataset to a large extent with very little loss of precision.
Binbin Peng, Wenyin Liu, Yin Liu 0001, Guanglin Huang, Zhengxing Sun, Xiangyu Jin
Int. J. Pattern Recognit. Artif. Intell.5
2002 On-Line Graphics Recognition
abstract
A novel and fast shape classification and regularization algorithm for on-line sketchy graphics recognition is proposed. We divided the on-line graphics recognition process into four stages: preprocessing, shape classification, shape fitting, and regularization. The attraction force model is proposed to combine progressively the vertices on the input sketchy stroke and reduce the total number of vertices before the type of shape can be determined After that, the shape is fitted and gradually rectified to a regular one, thus the regularized shape fits the user-intended one precisely. Experimental results show that this algorithm can rapidly yield good recognition precision (averagely above 90%) and a fine regularization effect. Consequently, it is especially suitable for weak computation environments such as PDAs, which solely depend on a pen-based user interface.
Xiangyu Jin, Wenyin Liu, Jianyong Sun, Zhengxing Sun
PG4
2002 Adaptive efficient video transmission over the Internet based on congestion control and RS coding
abstract
An approach based on adaptive congestion control and adaptive error recovery with RS (Reed-Solomon) coding method is presented for efficient video transmission over the Internet. Featured by weighted moving average rate control and TCP-friendliness, AVSP, a novel adaptive video streaming protocol, is designed with adjustable rate control parameters so as to respond quickly to the QoS status fluctuation during video transmission over the Internet. Combined with congestion control policy, an adaptive RS coding error recovery scheme with variable parameters is presented to enhance the robustness of MPEG video transmission over the Internet with restriction to the total system bandwidth.
Fuyan Zhang, Zhengxing Sun
Sci. China Ser. F Inf. Sci.3
2002 Greylevel Difference Classification Algorithm in Fractal Image Compression
Yisong Chen, Jian Lu 0001, Zhengxing Sun, Fuyan Zhang
J. Comput. Sci. Technol.3