Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Bin Sun 0002

dblp:01/5401-2 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0001-9239-0402ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 24% Efficient and distributed learning · 19% Video understanding and tracking · 19%
Computer graphics and multimedia
2 papers
Image and video processing · 57% Visual content generation and editing · 22% Geometric modeling and processing · 22%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
convolutional neural network
0.712023
Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution · AAAI 2023
Machine learning › Efficient and distributed learning › model compression
efficient architecture design
0.712023
Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution · AAAI 2023
Machine learning › Representation and self-supervised learning › visual representation
image representation
0.712023
Image as Set of Points · ICLR 2023
Computer vision › Image recognition and object detection
point set representation
0.712023
Image as Set of Points · ICLR 2023
Image and video processing › super-resolution
image super-resolution
0.712023
Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution · AAAI 2023
Image and video processing › super-resolution › image super-resolution
lightweight super-resolution
0.712023
Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution · AAAI 2023
Visual content generation and editing
image vectorization
0.612022
Towards Layer-wise Image Vectorization · CVPR 2022
Computer vision › Video understanding and tracking
action anticipation
0.312018
Action Prediction From Videos via Memorizing Hard-to-Predict Samples · AAAI 2018
Computer vision › Video understanding and tracking › action anticipation
early action prediction
0.312018
Action Prediction From Videos via Memorizing Hard-to-Predict Samples · AAAI 2018
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.212023
Image as Set of Points · ICLR 2023
Image and video processing
image restoration
0.212023
Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution · AAAI 2023

Methods — techniques the papers use, named apart from their topics

pixel-unshuffle · 1.3grouped convolution · 1.3depthwise separable convolution · 1.3point set representation · 0.7bezier path optimization · 0.6memory module · 0.3bidirectional LSTM · 0.3LSTM · 0.3CNN · 0.3
YearPublicationVenuePosition
2024 α-Former: Local-Feature-Aware (L-FA) Transformer
Zhi Xu 0013, Bin Sun 0002, Yun Fu 0001
UAI2
2023 Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution
abstract
Convolutional neural network (CNN) has achieved great success on image super-resolution (SR). However, most deep CNN-based SR models take massive computations to obtain high performance. Downsampling features for multi-resolution fusion is an efficient and effective way to improve the performance of visual recognition. Still, it is counter-intuitive in the SR task, which needs to project a low-resolution input to high-resolution. In this paper, we propose a novel Hybrid Pixel-Unshuffled Network (HPUN) by introducing an efficient and effective downsampling module into the SR task. The network contains pixel-unshuffled downsampling and Self-Residual Depthwise Separable Convolutions. Specifically, we utilize pixel-unshuffle operation to downsample the input features and use grouped convolution to reduce the channels. Besides, we enhance the depthwise convolution's performance by adding the input feature to its output. The comparison findings demonstrate that, with fewer parameters and computational costs, our HPUN achieves and surpasses the state-of-the-art performance on SISR. All results are provided in the github https://github.com/Sun1992/HPUN.
Bin Sun 0002, Yulun Zhang 0001, Songyao Jiang, Yun Fu 0001
AAAI1
2023 Image as Set of Points
Xu Ma 0005, Yuqian Zhou, Huan Wang 0014, Can Qin, Bin Sun 0002, Chang Liu 0022, Yun Fu 0001
ICLR5
2023 LRPRNet: Lightweight Deep Network by Low-Rank Pointwise Residual Convolution
abstract
Deep learning has become popular in recent years primarily due to powerful computing devices such as graphics processing units (GPUs). However, it is challenging to deploy these deep models to multimedia devices, smartphones, or embedded systems with limited resources. To reduce the computation and memory costs, we propose a novel lightweight deep learning module by low-rank pointwise residual (LRPR) convolution, called LRPRNet. Essentially, LRPR aims at using a low-rank approximation in pointwise convolution to further reduce the module size while keeping depthwise convolutions as the residual module to rectify the LRPR module. This is critical when the low-rankness undermines the convolution process. Moreover, our LRPR is quite general and can be directly applied to many existing network architectures such as MobileNetv1, ShuffleNetv2, MixNet, and so on. Experiments on visual recognition tasks, including image classification and face alignment on popular benchmarks, show that our LRPRNet achieves competitive performance but with a significant reduction of Flops and memory cost compared to the state-of-the-art deep lightweight models.
Bin Sun 0002, Jun Li 0027, Ming Shao, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Towards Layer-wise Image Vectorization
abstract
Image rasterization is a mature technique in computer graphics, while image vectorization, the reverse path of rasterization, remains a major challenge. Recent advanced deep learning-based models achieve vectorization and semantic interpolation of vector graphs and demonstrate a better topology of generating new figures. However, deep models cannot be easily generalized to out-of-domain testing data. The generated SVGs also contain complex and redundant shapes that are not quite convenient for further editing. Specifically, the crucial layer-wise topology and fundamental semantics in images are still not well understood and thus not fully explored. In this work, we propose Layer-wise Image Vectorization, namely LIVE, to convert raster images to SVGs and simultaneously maintain its image topology. LIVE can generate compact SVG forms with layer-wise structures that are semantically consistent with human perspective. We progressively add new bezier paths and optimize these paths with the layer-wise framework, newly designed loss functions, and component-wise path initialization technique. Our experiments demonstrate that LIVE presents more plausible vectorized forms than prior works and can be generalized to new images. With the help of this newly learned topology, LIVE initiates human editable SVGs for both designers and other downstream applications. Codes are made available at https://github.com/Picsart-AI-Research/LIVE-Layerwise-Image-Vectorization.
Xu Ma 0005, Yuqian Zhou, Xingqian Xu, Bin Sun 0002, Valerii Filev, Nikita Orlov, Yun Fu 0001, Humphrey Shi
CVPR4
2020 Block Mobilenet: Align Large-Pose Faces with <1MB Model Size
abstract
3D face alignment methods based on deep models have become very popular due to their empirical success. However, high time and space complexities make these methods difficult to be applied to mobile devices and embedded devices. To decrease the time and space complexity, we propose a novel Depthwise Separable Block (DSB) which consists of a depthwise block and a pointwise block. The depthwise block is constructed by stacking depthwise convolution layers and concatenating the low layer, and the pointwise block has only pointwise convolution layers stacking together. Moreover, we develop a light-weight Block-Mobilenet by using our DSBs to reconstruct Mobilenet. It is worth noting that our Block-Mobilenet successfully reduces network parameters from MB to KB. Experiments on four popular datasets verify that Block Mobilenet has better overall performance (mean NME on 68 points: 3.81%; speed on CPU: 91 FPS; storage size: 876 KB) than the state-of-the-art methods.
Bin Sun 0002, Jun Li 0027, Yun Fu 0001
FG1
2020 EV-Action: Electromyography-Vision Multi-Modal Action Dataset
abstract
Multi-modal human action analysis is a critical and attractive research topic. However, the majority of the existing datasets only provide visual modalities (i.e., RGB, depth and skeleton). To make up this, we introduce a new, largescale EV-Action dataset in this work, which consists of RGB, depth, electromyography (EMG), and two skeleton modalities. Compared with the conventional datasets, EV-Action dataset has two major improvements: (1) we deploy a motion capturing system to obtain high quality skeleton modality, which provides more comprehensive motion information including skeleton, trajectory, acceleration with higher accuracy, sampling frequency, and more skeleton markers. (2) we introduce an EMG modality which is usually used as an effective indicator in the biomechanics area, also it has yet to be well explored in motion related research. To the best of our knowledge, this is the first action dataset with EMG modality. The details of EVAction dataset are clarified, meanwhile, a simple yet effective framework for EMG-based action recognition is proposed. Moreover, state-of-the-art baselines are applied to evaluate the effectiveness of all the modalities. The obtained result clearly shows the validity of EMG modality in human action analysis tasks. We hope this dataset can make significant contributions to human motion analysis, computer vision, machine learning, biomechanics, and other interdisciplinary fields.
Lichen Wang, Bin Sun 0002, Joseph P. Robinson, Taotao Jing, Yun Fu 0001
FG2
2018 Action Prediction From Videos via Memorizing Hard-to-Predict Samples
abstract
Action prediction based on video is an important problem in computer vision field with many applications, such as preventing accidents and criminal activities. It's challenging to predict actions at the early stage because of the large variations between early observed videos and complete ones. Besides, intra-class variations cause confusions to the predictors as well. In this paper, we propose a mem-LSTM model to predict actions in the early stage, in which a memory module is introduced to record several "hard-to-predict" samples and a variety of early observations. Our method uses Convolution Neural Network (CNN) and Long Short-Term Memory (LSTM) to model partial observed video input. We augment LSTM with a memory module to remember challenging video instances. With the memory module, our mem-LSTM model not only achieves impressive performance in the early stage but also makes predictions without the prior knowledge of observation ratio. Information in future frames is also utilized using a bi-directional layer of LSTM. Experiments on UCF-101 and Sports-1M datasets show that our method outperforms state-of-the-art methods.
Yu Kong 0001, Shangqian Gao, Bin Sun 0002, Yun Fu 0001
AAAI3
2018 Deep Evolutionary 3D Diffusion Heat Maps for Large-pose Face Alignment
Bin Sun 0002, Ming Shao, Si-Yu Xia, Yun Fu 0001
BMVC1