Tiancheng Shen

dblp:215/5560 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video Editing
Tiancheng Shen, Xiangtai Li, Zhijie Lin 0001, Jiyang Liu, Jiashi Feng, Ming-Hsuan Yang 0001, Jun Hao Liew
ICCV1
2025 Rethinking Evaluation Metrics of Open-Vocabulary Segmentation
abstract
This paper highlights a problem of evaluation metrics adopted in the open-vocabulary segmentation. The evaluation process relies heavily on closed-set metrics on zero-shot or cross-dataset pipelines without considering the similarity between predicted and ground truth categories. We first survey eleven similarity measurements between two categorical words using WordNet linguistics statistics, text embedding, or language models by comprehensive quantitative analysis and user study to tackle this issue. Based on those explored measurements, we design novel evaluation metrics, Open mIoU, Open AP, and Open PQ, tailored for three open-vocabulary segmentation tasks. We benchmark the proposed evaluation metrics on twelve open-vocabulary methods in three segmentation tasks. Despite the relative subjectivity of similarity distance, we demonstrate that our metrics can still well evaluate the open ability of the existing open-vocabulary segmentation methods. We hope our work can bring the community new thinking about evaluating model ability for open-vocabulary segmentation.
Hao Zhou 0014, Lu Qi 0001, Tiancheng Shen, Hai Huang 0004, Xu Yang 0004, Xiangtai Li, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 High Quality Entity Segmentation
abstract
Dense image segmentation tasks (e.g., semantic, panoptic) are useful for image editing, but existing methods can hardly generalize well in an in-the-wild setting where there are unrestricted image domains, classes, and image resolution & quality variations. Motivated by these observations, we construct a new entity segmentation dataset, with a strong focus on high-quality dense segmentation in the wild. The dataset contains images spanning diverse image domains and entities, along with plentiful high-resolution images and high-quality mask annotations for training and testing. Given the high-quality and -resolution nature of the dataset, we propose CropFormer which is designed to tackle the intractability of instance-level segmentation on high-resolution images. It improves mask prediction by fusing high-res image crops that provides more fine-grained image details and the full image. CropFormer is the first query-based Transformer architecture that can effectively fuse mask predictions from multiple image views, by learning queries that effectively associate the same entities across the full image and its crop. With CropFormer, we achieve a significant AP gain of 1.9 on the challenging entity segmentation task. Furthermore, CropFormer consistently improves the accuracy of traditional segmentation tasks and datasets. The dataset and code are released at http://luqi.info/entityv2.github.io/.
Lu Qi 0001, Jason Kuen, Tiancheng Shen, Jiuxiang Gu, Wenbo Li 0001, Weidong Guo, Jiaya Jia, Zhe Lin 0001, Ming-Hsuan Yang 0001
ICCV3
2022 EfficientNeRF - Efficient Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) has been wildly applied to various tasks for its high-quality representation of 3D scenes. It takes long per-scene training time and per-image testing time. In this paper, we present EfficientNeRF as an efficient NeRF-based method to represent 3D scene and synthesize novel-view images. Although several ways exist to accelerate the training or testing process, it is still difficult to much reduce time for both phases simultaneously. We analyze the density and weight distribution of the sampled points then propose valid and pivotal sampling at the coarse and fine stage, respectively, to significantly improve sampling efficiency. In addition, we design a novel data structure to cache the whole scene during testing to accelerate the rendering speed. Overall, our method can reduce over 88% of training time, reach rendering speed of over 200 FPS, while still achieving competitive accuracy. Experiments prove that our method promotes the practicality of NeRF in the real world and enables many applications. The code is available in https://github.com/dvlabresearch/EfficientNeRF.
Tao Hu 0011, Shu Liu 0005, Tiancheng Shen, Jiaya Jia
CVPR4
2022 High Quality Segmentation for Ultra High-resolution Images
abstract
To segment 4K or 6K ultra high-resolution images needs extra computation consideration in image segmentation. Common strategies, such as downsampling, patch cropping, and cascade model, cannot address well the balance issue between accuracy and computation cost. Motivated by the fact that humans distinguish among objects continuously from coarse to precise levels, we propose the Continuous Refinement Model (CRM) for the ultra high-resolution segmentation refinement task. CRM continuously aligns the feature map with the refinement target and aggregates features to reconstruct these image details. Besides, our CRM shows its significant generalization ability to fill the resolution gap between low-resolution training images and ultra high-resolution testing ones. We present quantitative performance evaluation and visualization to show that our proposed method is fast and effective on image segmentation refinement. Code is available at https://github.com/dvlab-research/Entity/tree/main/CRM.
Tiancheng Shen, Yuechen Zhang, Lu Qi 0001, Jason Kuen, Xingyu Xie, Jianlong Wu, Zhe Lin 0001, Jiaya Jia
CVPR1
2021 PDO-eS2CNNs: Partial Differential Operator Based Equivariant Spherical CNNs
abstract
Spherical signals exist in many applications, e.g., planetary data, LiDAR scans and digitalization of 3D objects, calling for models that can process spherical data effectively. It does not perform well when simply projecting spherical data into the 2D plane and then using planar convolution neural networks (CNNs), because of the distortion from projection and ineffective translation equivariance. Actually, good principles of designing spherical CNNs are avoiding distortions and converting the shift equivariance property in planar CNNs to rotation equivariance in the spherical domain. In this work, we use partial differential operators (PDOs) to design a spherical equivariant CNN, PDO-eS2CNN, which is exactly rotation equivariant in the continuous domain. We then discretize PDO-eS2CNNs, and analyze the equivariance error resulted from discretization. This is the first time that the equivariance error is theoretically analyzed in the spherical domain. In experiments, PDO-eS2CNNs show greater parameter efficiency and outperform other spherical CNNs significantly on several tasks.
Zhengyang Shen, Tiancheng Shen, Zhouchen Lin, Jinwen Ma
AAAI2
2021 Recurrent learning with clique structures for prostate sparse-view CT artifacts reduction
abstract
Abstract In recent years, convolutional neural networks have achieved great success in streak artifacts reduction. However, there is no special method designed for the artifacts reduction of the prostate. To solve the problem, the artifacts reduction CliqueNet (ARCliqueNet) to reconstruct dense‐view computed tomography images form sparse‐view computed tomography images is proposed. In detail, first, the proposed ARCliqueNet extracts a set of feature maps from the prostate sparse‐view CT image by Clique Block. Second, the feature maps are sent to ASPP with memory to be refined. Thenanother Clique Block is applied to the output of ASPP with memory and reconstruct the dense‐view CT images. Later on, reconstructed dense‐view CT images are used as new input of the original network. This process is repeated recurrently with memory delivering information between these recurrent stages. The final reconstructed dense‐view CT images are the output of the last recurrent stage. Our proposed ARCliqueNet outperforms the SOTA (state‐of‐the‐art) general artifacts reduction methods on the prostate dataset in terms of PSNR (peak signal‐to‐noise ratio) and SSIM (structural similarity). Therefore, we can draw the conclusion that Clique structures, ASPP with memory and recurrent learning are useful for prostate sparse‐view CT Artifacts here.
Tiancheng Shen, Zhouchen Lin, Mingbin Zhang
IET Image Process.1
2020 PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms
abstract
With the rapid expansion of digital music formats, it's indispensable to recommend users with their favorite music. For music recommendation, users' personality and emotion greatly affect their music preference, respectively in a long-term and short-term manner, while rich social media data provides effective feedback on these information. In this paper, aiming at music recommendation on social media platforms, we propose a Personality and Emotion Integrated Attentive model (PEIA), which fully utilizes social media data to comprehensively model users' long-term taste (personality) and short-term preference (emotion). Specifically, it takes full advantage of personality-oriented user features, emotion-oriented user features and music features of multi-faceted attributes. Hierarchical attention is employed to distinguish the important factors when incorporating the latent representations of users' personality and emotion. Extensive experiments on a large real-world dataset of 171,254 users demonstrate the effectiveness of our PEIA model which achieves an NDCG of 0.5369, outperforming the state-of-the-art methods. We also perform detailed parameter analysis and feature contribution analysis, which further verify our scheme and demonstrate the significance of co-modeling of user personality and emotion in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Yihui Ma, Yaohua Bu, Hanjie Wang, Tat-Seng Chua, Wendy Hall 0001
AAAI1
2020 Dynamical System Inspired Adaptive Time Stepping Controller for Residual Network Families
abstract
The correspondence between residual networks and dynamical systems motivates researchers to unravel the physics of ResNets with well-developed tools in numeral methods of ODE systems. The Runge-Kutta-Fehlberg method is an adaptive time stepping that renders a good trade-off between the stability and efficiency. Can we also have an adaptive time stepping for ResNets to ensure both stability and performance? In this study, we analyze the effects of time stepping on the Euler method and ResNets. We establish a stability condition for ResNets with step sizes and weight parameters, and point out the effects of step sizes on the stability and performance. Inspired by our analyses, we develop an adaptive time stepping controller that is dependent on the parameters of the current step, and aware of previous steps. The controller is jointly optimized with the network training so that variable step sizes and evolution time can be adaptively adjusted. We conduct experiments on ImageNet and CIFAR to demonstrate the effectiveness. It is shown that our proposed method is able to improve both stability and accuracy without introducing additional overhead in inference phase.
Jianlong Wu, Xia Li 0005, Tiancheng Shen, Zhouchen Lin
AAAI5
2020 Spatial Pyramid Based Graph Reasoning for Semantic Segmentation
abstract
The convolution operation suffers from a limited receptive filed, while global modeling is fundamental to dense prediction tasks, such as semantic segmentation. In this paper, we apply graph convolution into the semantic segmentation task and propose an improved Laplacian. The graph reasoning is directly performed in the original feature space organized as a spatial pyramid. Different from existing methods, our Laplacian is data-dependent and we introduce an attention diagonal matrix to learn a better distance metric. It gets rid of projecting and re-projecting processes, which makes our proposed method a light-weight module that can be easily plugged into current computer vision architectures. More importantly, performing graph reasoning directly in the feature space retains spatial relationships and makes spatial pyramid possible to explore multiple long-range contextual patterns from different scales. Experiments on Cityscapes, COCO Stuff, PASCAL Context and PASCAL VOC demonstrate the effectiveness of our proposed methods on semantic segmentation. We achieve comparable performance with advantages in computational and memory overhead.
Xia Li 0005, Qijie Zhao, Tiancheng Shen, Zhouchen Lin, Hong Liu 0008
CVPR4
2020 Enhancing Music Recommendation with Social Media Content: an Attentive Multimodal Autoencoder Approach
abstract
Music recommendation methods predict users' music preference primarily based on historical ratings. Meanwhile, manifold personal factors of users are also important for the problem, and research efforts have been made to improve the recommendation performance with auxiliary user information. As an important indicator of users' personal traits and states, the numerous social media content (e.g., texts, images and short videos), however, is still hardly exploited. In this work, we systematically study the utilization of multimodal social media content for music recommendation. We define groups of both targeted handcrafted features and generic deep features for each modality, and further propose an Attentive Multimodal Autoencoder approach (AMAE) to learn cross-modal latent representations from the extracted features. Attention mechanism is also employed to integrate users' global and contextual music preference with alterable weights. Experiments demonstrate remarkable improvement of recommendation performance (+2.40% in Hit Ratio and +3.30% in NDCG), manifesting the effectiveness of our AMAE approach, as well as the significance of incorporating social media content data in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Hanjie Wang
IJCNN1
2019 R ^2 2 -Net: Recurrent and Recursive Network for Sparse-View CT Artifacts Removal
Tiancheng Shen, Xia Li 0005, Zhisheng Zhong, Jianlong Wu, Zhouchen Lin
MICCAI (6)1
2018 Convolutional Neural Networks With Alternately Updated Clique
abstract
Improving information flow in deep networks helps to ease the training difficulties and utilize parameters more efficiently. Here we propose a new convolutional neural network architecture with alternately updated clique (CliqueNet). In contrast to prior networks, there are both forward and backward connections between any two layers in the same block. The layers are constructed as a loop and are updated alternately. The CliqueNet has some unique properties. For each layer, it is both the input and output of any other layer in the same block, so that the information flow among layers is maximized. During propagation, the newly updated layers are concatenated to re-update previously updated layer, and parameters are reused for multiple times. This recurrent feedback structure is able to bring higher level visual information back to refine low-level filters and achieve spatial attention. We analyze the features generated at different stages and observe that using refined features leads to a better result. We adopt a multiscale feature strategy that effectively avoids the progressive growth of parameters. Experiments on image recognition datasets including CIFAR-10, CIFAR-100, SVHN and ImageNet show that our proposed models achieve the state-of-the-art performance with fewer parameters.
Zhisheng Zhong, Tiancheng Shen, Zhouchen Lin
CVPR3
2018 Cross-Domain Depression Detection via Harvesting Social Media
abstract
Depression detection is a significant issue for human well-being. In previous studies, online detection has proven effective in Twitter, enabling proactive care for depressed users. Owing to cultural differences, replicating the method to other social media platforms, such as Chinese Weibo, however, might lead to poor performance because of insufficient available labeled (self-reported depression) data for model training. In this paper, we study an interesting but challenging problem of enhancing detection in a certain target domain (e.g. Weibo) with ample Twitter data as the source domain. We first systematically analyze the depression-related feature patterns across domains and summarize two major detection challenges, namely isomerism and divergency. We further propose a cross-domain Deep Neural Network model with Feature Adaptive Transformation & Combination strategy (DNN-FATC) that transfers the relevant information across heterogeneous domains. Experiments demonstrate improved performance compared to existing heterogeneous transfer methods or training directly in the target domain (over 3.4% improvement in F1), indicating the potential of our model to enable depression detection via social media for more countries with different cultural settings.
Tiancheng Shen, Jia Jia 0001, Guangyao Shen, Fuli Feng, Xiangnan He 0001, Huan-Bo Luan, Jie Tang 0001, Thanassis Tiropanis, Tat-Seng Chua, Wendy Hall 0001
IJCAI1
2018 Joint Sub-bands Learning with Clique Structures for Wavelet Domain Super-Resolution
abstract
Convolutional neural networks (CNNs) have recently achieved great success in single-image super-resolution (SISR). However, these methods tend to produce over-smoothed outputs and miss some textural details. To solve these problems, we propose the Super-Resolution CliqueNet (SRCliqueNet) to reconstruct the high resolution (HR) image with better textural details in the wavelet domain. The proposed SRCliqueNet firstly extracts a set of feature maps from the low resolution (LR) image by the clique blocks group. Then we send the set of feature maps to the clique up-sampling module to reconstruct the HR image. The clique up-sampling module consists of four sub-nets which predict the high resolution wavelet coefficients of four sub-bands. Since we consider the edge feature properties of four sub-bands, the four sub-nets are connected to the others so that they can learn the coefficients of four sub-bands jointly. Finally we apply inverse discrete wavelet transform (IDWT) to the output of four sub-nets at the end of the clique up-sampling module to increase the resolution and reconstruct the HR image. Extensive quantitative and qualitative experiments on benchmark datasets show that our method achieves superior performance over the state-of-the-art methods.
Zhisheng Zhong, Tiancheng Shen, Zhouchen Lin, Chao Zhang 0001
NeurIPS2