Chuanxin Tang

dblp:159/3894 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2023 Look Before You Match: Instance Understanding Matters in Video Object Segmentation
abstract
Exploring dense matching between the current frame and past frames for long-range context modeling, memory-based methods have demonstrated impressive results in video object segmentation (VOS) recently. Nevertheless, due to the lack of instance understanding ability, the above approaches are oftentimes brittle to large appearance variations or viewpoint changes resulted from the movement of objects and cameras. In this paper, we argue that instance understanding matters in VOS, and integrating it with memory-based matching can enjoy the synergy, which is intuitively sensible from the definition of VOS task, i.e., identifying and segmenting object instances within the video. Towards this goal, we present a two-branch network for VOS, where the query-based instance segmentation (IS) branch delves into the instance details of the current frame and the VOS branch performs spatial-temporal matching with the memory bank. We employ the well-learned object queries from IS branch to inject instance-specific information into the query key, with which the instance-augmented matching is further performed. In addition, we introduce a multi-path fusion block to effectively combine the memory readout with multi-scale features from the instance segmentation decoder, which incorporates high-resolution instance-aware features to produce final segmentation results. Our method achieves state-of-the-art performance on DAVIS 2016/2017 val (92.6% and 87.1%), DAVIS 2017 test-dev (82.8%), and YouTube-VOS 2018/2019 val (86.3% and 86.3%), outperforming alternative methods by clear margins.
Dongdong Chen 0001, Zuxuan Wu, Chong Luo 0001, Chuanxin Tang, Xiyang Dai, Yujia Xie, Lu Yuan 0001, Yu-Gang Jiang 0001
CVPR5
2023 Streaming Video Model
abstract
Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recognition, use a video backbone to directly extract spatiotemporal features, while frame-based video tasks, such as multiple object tracking (MOT), rely on single fixed-image backbone to extract spatial features. In contrast, we propose to unify video understanding tasks into one novel streaming video architecture, referred to as Streaming Vision Transformer (S-ViT). S-ViT first produces frame-level features with a memory-enabled temporally-aware spatial encoder to serve the frame-based video tasks. Then the frame features are input into a task-related temporal decoder to obtain spatiotemporal features for sequence-based tasks. The efficiency and efficacy of S-ViT is demonstrated by the state-of-the-art accuracy in the sequence-based action recognition task and the competitive advantage over conventional architecture in the frame-based MOT task. We believe that the concept of streaming video model and the implementation of S-ViT are solid steps towards a unified deep learning architecture for video understanding. Code will be available at https://github.com/yuzhms/Streaming-Video-Model.
Chong Luo 0001, Chuanxin Tang, Dongdong Chen 0001, Noel Codella, Zhengjun Zha
CVPR3
2023 Filler Word Detection with Hard Category Mining and Inter-Category Focal Loss
abstract
Filler words like "um" or "uh" are common in spontaneous speech. It is desirable to automatically detect and remove them in recordings, as they affect the fluency, confidence, and professionalism of speech. Previous studies and our preliminary experiments reveal that the biggest challenge in filler word detection is that fillers can be easily confused with other hard categories like "a" or "I". In this paper, we propose a novel filler word detection method that effectively addresses this challenge by adding auxiliary categories dynamically and applying an additional inter-category focal loss. The auxiliary categories force the model to explicitly model the confusing words by mining hard categories. In addition, inter-category focal loss adaptively adjusts the penalty weight between "filler" and "non-filler" categories to deal with other confusing words left in the "non-filler" category. Our system achieves the best results, with a huge improvement compared to other methods on the PodcastFillers dataset.
Zhiyuan Zhao 0001, Chuanxin Tang, Dacheng Yin, Chong Luo 0001
ICASSP3
2023 TridentSE: Guiding Speech Enhancement with 32 Global Tokens
Dacheng Yin, Zhiyuan Zhao 0001, Chuanxin Tang, Zhiwei Xiong, Chong Luo 0001
INTERSPEECH3
2022 Sparse MLP for Image Recognition: Is Self-Attention Really Necessary?
abstract
Transformers have sprung up in the field of computer vision. In this work, we explore whether the core self-attention module in Transformer is the key to achieving excellent performance in image recognition. To this end, we build an attention-free network called sMLPNet based on the existing MLP-based vision models. Specifically, we replace the MLP module in the token-mixing step with a novel sparse MLP (sMLP) module. For 2D image tokens, sMLP applies 1D MLP along the axial directions and the parameters are shared among rows or columns. By sparse connection and weight sharing, sMLP module significantly reduces the number of model parameters and computational complexity, avoiding the common over-fitting problem that plagues the performance of MLP-like models. When only trained on the ImageNet-1K dataset, the proposed sMLPNet achieves 81.9% top-1 accuracy with only 24M parameters, which is much better than most CNNs and vision Transformers under the same model size constraint. When scaling up to 66M parameters, sMLPNet achieves 83.4% top-1 accuracy, which is on par with the state-of-the-art Swin Transformer. The success of sMLPNet suggests that the self-attention mechanism is not necessarily a silver bullet in computer vision. The code and models are publicly available at https://github.com/microsoft/SPACH.
Chuanxin Tang, Guangting Wang, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001
AAAI1
2022 When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
abstract
Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. However, is the attention mechanism truly an indispensable part of ViT? Can it be replaced by some other alternatives? To demystify the role of attention mechanism, we simplify it into an extremely simple case: ZERO FLOP and ZERO parameter. Concretely, we revisit the shift operation. It does not contain any parameter or arithmetic calculation. The only operation is to exchange a small portion of the channels between neighboring features. Based on this simple operation, we construct a new backbone network, namely ShiftViT, where the attention layers in ViT are substituted by shift operations. Surprisingly, ShiftViT works quite well in several mainstream tasks, e.g., classification, detection, and segmentation. The performance is on par with or even better than the strong baseline Swin Transformer. These results suggest that the attention mechanism might not be the vital factor that makes ViT successful. It can be even replaced by a zero-parameter operation. We should pay more attentions to the remaining parts of ViT in the future work. Code is available at github.com/microsoft/SPACH.
Guangting Wang, Chuanxin Tang, Chong Luo 0001, Wenjun Zeng 0001
AAAI3
2022 RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech Insertion
abstract
This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrarylength speech insertion and even full sentence generation.In the proposed paradigm, global and local factors in speech are explicitly decomposed and separately manipulated to achieve high speaker similarity and continuous prosody.Specifically, we proposed to represent the global factors by multiple tokens, which are extracted by cross-attention operation and then injected back by link-attention operation.Due to the rich representation of global factors, we manage to achieve high speaker similarity in a zero-shot manner.In addition, we introduce a prosody smoothing task to make the local prosody factor context-aware and therefore achieve satisfactory prosody continuity.We further achieve high voice quality with an adversarial training stage.In the subjective test, our method achieves state-of-the-art performance in both naturalness and similarity.Audio samples can be found at https://ydcustc.github.io/retrieverTTS-demo/.
Dacheng Yin, Chuanxin Tang, Xiaoqiang Wang 0006, Zhiyuan Zhao 0001, Zhiwei Xiong, Sheng Zhao 0002, Chong Luo 0001
INTERSPEECH2
2022 An Anchor-Free Detector for Continuous Speech Keyword Spotting
abstract
Continuous Speech Keyword Spotting (CSKWS) is a task to detect predefined keywords in a continuous speech.In this paper, we regard CSKWS as a one-dimensional object detection task and propose a novel anchor-free detector, named AF-KWS, to solve the problem.AF-KWS directly regresses the center locations and lengths of the keywords through a single-stage deep neural network.In particular, AF-KWS is tailored for this speech task as we introduce an auxiliary unknown class to exclude other words from non-speech or silent background.We have built two benchmark datasets named LibriTop-20 and continuous meeting analysis keywords (CMAK) dataset for CSKWS.Evaluations on these two datasets show that our proposed AF-KWS outperforms reference schemes by a large margin, and therefore provides a decent baseline for future research.
Zhiyuan Zhao 0001, Chuanxin Tang, Chengdong Yao, Chong Luo 0001
INTERSPEECH2
2021 Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration
abstract
Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.Existing methods adopt a two-stage approach: synthesize the input text using a generic text-to-speech (TTS) engine and then transform the voice to the desired voice using voice conversion (VC).A major problem of this framework is that VC is a challenging problem which usually needs a moderate amount of parallel training data to work satisfactorily.In this paper, we propose a one-stage context-aware framework to generate natural and coherent target speech without any training data of the target speaker.In particular, we manage to perform accurate zero-shot duration prediction for the inserted text.The predicted duration is used to regulate both text embedding and speech embedding.Then, based on the aligned cross-modality input, we directly generate the mel-spectrogram of the edited speech with a transformer-based decoder.Subjective listening tests show that despite the lack of training data for the speaker, our method has achieved satisfactory results.It outperforms a recent zero-shot TTS engine by a large margin.
Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Dacheng Yin, Wenjun Zeng 0001
Interspeech1
2020 Joint Time-Frequency and Time Domain Learning for Speech Enhancement
abstract
For single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framework takes advantage of the knowledge we have about spectrogram and avoids some of the drawbacks that T-F-domain methods have been suffering from. In TFT-Net, we design an innovative dual-path attention block (DAB) to fully exploit correlations along the time and frequency axes. We further discover that a sample-independent DAB (SDAB) achieves a good tradeoff between enhanced speech quality and complexity. Ablation studies show that both the cross-domain design and the SDAB block bring large performance gain. When logarithmic MSE is used as the training criteria, TFT-Net achieves the highest SDR and SSNR among state-of-the-art methods on two major speech enhancement benchmarks.
Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Wenxuan Xie, Wenjun Zeng 0001
IJCAI1
2014 A new frame interpolation method with pixel-level motion vector field
abstract
In this paper, a new frame interpolation method with pixel-level motion vector field (MVF) is proposed. Given that existing methods cannot handle occlusions and blocking artifacts well, there are three contributions in our method: (i) applying the pixel-level motion vectors (MVs) estimated by optical flow algorithm to eliminate blocking artifacts (ii) motion post-processing to keep spatial consistency (iii) robust warping method to address collisions and holes caused by occlusions. The method could remove blocking artifacts and alleviate the artifacts caused by occlusions. Experimental results show that the proposed method outperforms existing methods both in terms of objective and subjective performances, especially for sequences with complex motions.
Chuanxin Tang, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
VCIP1