Tong Shao

dblp:53/2244 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
abstract
Brain decoding is a key neuroscience field that reconstructs the visual stimuli from brain activity with fMRI, which helps illuminate how the brain represents the world. fMRI-to-image reconstruction has achieved impressive progress by leveraging diffusion models. However, brain signals infused with prior knowledge and associations exhibit a significant information asymmetry when compared to raw visual features, still posing challenges for decoding fMRI representations under the supervision of images. Consequently, the reconstructed images often lack fine-grained visual fidelity, such as missing attributes and distorted spatial relationships. To tackle this challenge, we propose BrainCognizer, a novel brain decoding model inspired by human visual cognition, which explores multilevel semantics and correlations without fine-tuning of generative models. Specifically, BrainCognizer introduces two modules: the Cognitive Integration Module which incorporates prior human knowledge to extract hierarchical region semantics; and the Cognitive Correlation Module which captures contextual semantic relationships across regions. Incorporating these two modules enhances intra-region semantic consistency and maintains interregion contextual associations, thereby facilitating fine-grained brain decoding. Moreover, we quantitatively interpret our components from a neuroscience perspective and analyze the associations between different visual patterns and brain functions. Extensive quantitative and qualitative experiments demonstrate that BrainCognizer outperforms state-of-the-art approaches on multiple evaluation metrics. Our code is released publicly at https://github.com/Grace160/BrainCognizer.
Guoying Sun, Weiyu Guo, Tong Shao, Yang Yang 0002, Haijin Zeng, Jingyong Su
BIBM3
2025 Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicing
abstract
Quad Bayer demosaicing is the central challenge for enabling the widespread application of Hybrid Event-based Vision Sensors (HybridEVS). Although existing learning-based methods that leverage long-range dependency modeling have achieved promising results, their complexity severely limits deployment on mobile devices for real-world applications. To address these limitations, we propose a lightweight Mamba-based binary neural network designed for efficient and high-performing demosaicing of HybridEVS RAW images. First, to effectively capture both global and local dependencies, we introduce a hybrid Binarized Mamba-Transformer architecture that combines the strengths of the Mamba and Swin Transformer architectures. Next, to significantly reduce computational complexity, we propose a binarized Mamba (Bi-Mamba), which binarizes all projections while retaining the core Selective Scan in full precision. Bi-Mamba also incorporates additional global visual information to enhance global context and mitigate precision loss. We conduct quantitative and qualitative experiments to demonstrate the effectiveness of BMTNet in both performance and computational efficiency, providing a lightweight demosaicing solution suited for real-world edge devices. Our codes and models are available at https://github.com/Clausy9/BMTNet.
Haijin Zeng, Yunfan Lu, Tong Shao, Yongyong Chen, Jingyong Su
CVPR4
2025 BFRA: A Bi-Level Feature Relation Alignment Method for Cross-Domain Few-Shot Learning
abstract
While existing Few-Shot Learning (FSL) techniques demonstrate strong performance on uniform datasets, they encounter domain shift challenges when presented with domainagnostic queries in real-world scenarios. So we investigate it in Cross-Domain Few-Shot Learning (CD-FSL) and propose to learn more universal feature representations to enhance generalization on unseen domains. Toward this issue, we pinpoint two issues in current multi-model fusion approaches: 1) the entanglement of domain and class information, and 2) feature overlap across distinct domains. To address these challenges, we introduce a Bi-level Feature Relation Alignment method, BFRA, which facilitates the acquisition of more versatile features by decoupling domain-class relationships and aligning feature relations. Through the segregation of domain and class feature learning, we devise a smoothing layer prior for domain feature alignment to mitigate inter-domain discrepancies. This approach enables our model to acquire domain-consistent features, diminishing interference in subsequent class feature alignment procedures. During the class feature alignment, we notice that class feature representations from various in-domain models may intersect, leading to a diminished distinction between classes. To address this, we adopt a topological perspective to train our target model, by aligning feature relations instead of features between our target model and multiple in-domain models. The integration of these components results in the establishment of a bi-level feature relation alignment framework aimed at acquiring more universal features. Furthermore, we partially fine-tune the plug-in layer-wise affine adapter on domain-agnostic queries to expedite adaptation without impacting the known domains. Experiments of 21 datasets on meta-dataset and BSCD-FSL benchmark demonstrate the effectiveness of our method. The code are made publicly available at https://github.com/leaves162/BFRA.
Tong Shao, Zhuotao Tian, Jingyong Su
IEEE Trans. Circuits Syst. Video Technol.1
2024 Residual Block Fusion in Low Complexity Neural Network-Based In-loop Filtering for Video Compression
abstract
In this paper, a novel low complexity residual block fusion (RBF) based split luma chroma architecture is proposed to improve coding efficiency of neural network-based in-loop filter in video compression. The residual block in this architecture consists of a 1x1 convolution layer with wide activation and a regular 3x3 convolutional layer decomposed into 1x1 pointwise convolutions and 1x3/3x1 separable convolutions via Canonical Polyadic (CP) decomposition to reduce complexity. By adjusting the location of the skip connection in each residual block, the fusion of adjacent 1x1 pointwise convolutions is performed. The RBF backbone consists of a new wide activation that directly starts with PReLU and is followed by a 1x1 convolution, while the 1x1 layers after CP decomposition are fully fused. This new fusion design reduces the complexity from 17.05 kMac/Pixel to 16.56 kMac/Pixel and the number of convolutional layers by 13%. The experimental results show that new RBF architecture’s BDRate is {-0.11%, -0.31%, -0.33%} under All Intra (AI) and {-0.14%, 0.66%, 1.56%} under Random Access (RA) compared to existing residual block design, while the BD-Rate of the proposed RBF loop filer compared to VTM anchor is {-4.77%, -9.14%, -9.13%} under AI and {-5.46%, -9.31%, -9.20%} under RA. The actual decoding time is reduced by around 5% after residual block fusion. The BD-Rate and kMac/Pixel plot also shows superior trade-off between complexity and coding gain compared to state-of-the-art filters.
Tong Shao, Jay N. Shingala, Ajay Shyam, Peng Yin 0002, Ajat Suneja, Siddarth P. Badya, Arjun Arora, Sean McCarthy
DCC1
2024 Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation
Tong Shao, Zhuotao Tian, Hang Zhao 0019, Jingyong Su
ECCV (86)1
2024 Exploring Bimanual Haptic Feedback for Spatial Search in Virtual Reality
abstract
Spatial search tasks are common and crucial in many Virtual Reality (VR) applications. Traditional methods to enhance the performance of spatial search often employ sensory cues such as visual, auditory, or haptic feedback. However, the design and use of bimanual haptic feedback with two VR controllers for spatial search in VR remains largely unexplored. In this work, we explored bimanual haptic feedback with various combinations of haptic properties, where four types of bimanual haptic feedback were designed, for spatial search tasks in VR. Two experiments were designed to evaluate the effectiveness of bimanual haptic feedback on spatial direction guidance and search in VR. The results from the first experiment reveal that our proposed bimanual haptic schemes significantly enhanced the recognition of spatial directions in terms of accuracy and speed compared to spatial audio feedback. The second experiment's findings suggest that the performance of bimanual haptic feedback was comparable to or even better than the visual arrow, especially in reducing the angle of head movement and enhancing searching targets behind the participants, which was supported by subjective feedback as well. Based on these findings, we have derived a set of design recommendations for spatial search using bimanual haptic feedback in VR.
Boyu Gao 0003, Tong Shao, Huawei Tu, Qizi Ma, Zitao Liu 0001, Teng Han
IEEE Trans. Vis. Comput. Graph.2
2023 fmLRE: A Low-Resource Relation Extraction Model Based on Feature Mapping Similarity Calculation
abstract
Low-resource relation extraction (LRE) aims to extract relations from limited labeled corpora. Existing work takes advantages of self-training or distant supervision to expand the limited labeled data in the data-driven approaches, while the selection bias of pseudo labels may cause the error accumulation in subsequent relation classification. To address this issue, this paper proposes fmLRE, an iterative feedback method based on feature mapping similarity calculation to improve the accuracy of pseudo labels. First, it calculates the similarities between pseudo-label and real-label data of the same category in a feature mapping space based on semantic features of labeled dataset after feature projection. Then, it fine-tunes initial model according to the iterative process of reinforcement learning. Finally, the similarity is used as a threshold for screening high-precision pseudo-labels and the basis for setting different rewards, which also acts as a penalty term for the loss function of relation classifier. Experimental results demonstrate that fmLRE achieves the state-of-the-art performance compared with strong baselines on two public datasets.
Peng Wang 0004, Tong Shao, Ke Ji, Wenjun Ke 0002
AAAI2
2023 A Low Complexity Convolutional Neural Network with Fused CP Decomposition for In-Loop Filtering in Video Coding
abstract
In this paper, a novel low complexity convolutional neural network with fused CP decomposition is proposed for in-loop filtering in video coding. Based on the baseline model in JVET-X0140, the regular 3x3 convolutional layers are replaced by pointwise convolutions and separable convolutions via CP decomposition. We further propose to fuse the 1x1 pointwise convolutional layers among the decomposed layers with their adjacent regular 1x1 convolutional layers, resulting in one single 1x1 convolutional layer. The two procedures reduce the model complexity from 33.6 KMAC/Pixel to 16.265 KMAC/Pixel. Experimental results show that the model has 4.45% BD-Rate luma gain over VTM NNVC-2.0. It demonstrates the (0.56%, -0.63%, -1.89%) loss of (Y, U, V) for RA and (0.51%, 0.21%, 0.39%) for AI, while the CPU decoding time is reduced by 19% for RA and 24% for AI, proving the great ability of the fused CP decomposition model to reduce complexity while maintaining good trade-off. The BD-Rate and KMAC/Pixel plot also shows the superior trade-off between complexity and coding gain compared to state-of the-art filters.
Tong Shao, Jay N. Shingala, Peng Yin 0002, Arjun Arora, Ajay Shyam, Sean McCarthy
DCC1
2022 Hybrid Conditional Deep Inverse Tone Mapping
abstract
Emerging modern displays are capable to render ultra-high definition (UHD) media contents with high dynamic range (HDR) and wide color gamut (WCG). Although more and more native contents as such have been getting produced, the total amount is still in severe lack. Considering the massive amount of legacy contents with standard dynamic range (SDR) which may be exploitable, the urgent demand for proper conversion techniques thus springs up. In this paper, we try to tackle the conversion task from SDR to HDR-WCG for media contents and consumer displays. We propose a deep learning based SDR-to-HDR solution, Hybrid Conditional Deep Inverse Tone Mapping (HyCondITM), which is an end-to-end trainable framework including global transform, local adjustment, and detail refinement in a single unified pipeline. We present a hybrid condition network that can simultaneously extract both global and local priors for guidance to achieve scene-adaptive and spatially-variant manipulations. Experiments show that our method achieves state-of-the-art performance in both quantitative comparisons and visual quality, out-performing the previous methods.
Tong Shao, Deming Zhai, Junjun Jiang, Xianming Liu 0005
ACM Multimedia1
2022 PTR-CNN for in-loop filtering in video coding
Tong Shao, Dapeng Oliver Wu, Chia-Yang Tsai, Zhijun Lei, Ioannis Katsavounidis
J. Vis. Commun. Image Represent.1
2021 A Sentiment and Style Controllable Approach for Chinese Poetry Generation
abstract
Sentiment and style control are two vital aspects in automatic poetry generation. Excellent Chinese classical poetry should express a certain emotion and embody a specific style at the same time. Existing work still has deficiencies in controlling sentiment and style simultaneously. To address above issues, in this paper, we propose a novel approach for Chinese classical poetry generation, which can generate sentiment-controllable and style-controllable poems. First, it classifies hundreds of thousands of poems by style, sentiment, format, and primary keyword. Then, it utilizes masking self-attention mechanism to associate multiple tags and verses. Besides, it can generate metrical rhyming verses with distinctive sentiment and style characteristics according to the tag-set and secondary keywords. Finally, this approach is applied in Chang Qing Yin, which can collaborate with users to polish generated poems, providing alternatives automatically. Experimental results show that our approach performs well in sentiment and style control, and quality of generated poems outperforms several strong baselines.
Yizhan Shao, Tong Shao, Peng Wang 0004
CIKM2
2021 Application of Bayesian phylogenetic inference modelling for evolutionary genetic analysis and dynamic changes in 2019-nCoV
abstract
The novel coronavirus (2019-nCoV) has recently caused a large-scale outbreak of viral pneumonia both in China and worldwide. In this study, we obtained the entire genome sequence of 777 new coronavirus strains as of 29 February 2020 from a public gene bank. Bioinformatics analysis of these strains indicated that the mutation rate of these new coronaviruses is not high at present, similar to the mutation rate of the severe acute respiratory syndrome (SARS) virus. The similarities of 2019-nCoV and SARS virus suggested that the S and ORF6 proteins shared a low similarity, while the E protein shared the higher similarity. The 2019-nCoV sequence has similar potential phosphorylation sites and glycosylation sites on the surface protein and the ORF1ab polyprotein as the SARS virus; however, there are differences in potential modification sites between the Chinese strain and some American strains. At the same time, we proposed two possible recombination sites for 2019-nCoV. Based on the results of the skyline, we speculate that the activity of the gene population of 2019-nCoV may be before the end of 2019. As the scope of the 2019-nCoV infection further expands, it may produce different adaptive evolutions due to different environments. Finally, evolutionary genetic analysis can be a useful resource for studying the spread and virulence of 2019-nCoV, which are essential aspects of preventive and precise medicine.
Tong Shao, Wenfang Wang, Meiyu Duan, Zhuoyuan Xin, Baoyue Liu, Fengfeng Zhou
Briefings Bioinform.1
2021 An Efficient Scheme of Cloud Data Assured Deletion
abstract
Abstract With the rapid development of cloud storage technology, cloud data assured deletion has received extensive attention. While ensuring the deletion of cloud data, users have also placed increasing demands on cloud data assured deletion, such as improving the execution efficiency of various stages of a cloud data assured deletion system and performing fine-grained access and deletion operations. In this paper, we propose an efficient scheme of cloud data assured deletion. The scheme replaces complicated bilinear pairing with simple scalar multiplication on elliptic curves to realize ciphertext policy attribute-based encryption of cloud data, while solving the security problem of shared data. In addition, the efficiency of encryption and decryption is improved, and fine-grained access of ciphertext is realized. The scheme designs an attribute key management system that employs a dual-server to solve system flaws caused by single point failure. The scheme is proven to be secure, based on the decisional Diffie-Hellman assumption in the standard model; therefore, it has stronger security. The theoretical analysis and experimental results show that the scheme guarantees security and significantly improves the efficiency of each stage of cloud data assured deletion.
Yuechi Tian, Tong Shao
Mob. Networks Appl.2
2015 Inter-picture prediction based on 3D point cloud model
abstract
The explosive growth of images in the cloud calls for more efficient image compression that exploits redundancy between images. Traditional solutions to image set compression are sensitive to the diversity in the clustered image sets in the cloud. In this paper, we propose a more efficient inter-picture prediction based on 3D point cloud model. Leveraging the 3D points and camera parameters, we design a patch-wise warping algorithm to generate prediction. Preliminary experimental results show the proposed scheme significantly outperforms the JPEG and HEVC intra as long as high quality reference is provided. The average bitrate reduction is 15.1% in comparison with HEVC intra on the tested 30 images of Notre Dame. Meanwhile, the perceptual quality of the reconstructed images is improved.
Tong Shao, Dong Liu 0002, Houqiang Li
ICIP1