VLDB 2026 Research / reviewers in the wild / expert
Sheng Shi
dblp:166/2843
· DBLP profile ↗
17ranked-venue papers
9as first author
11since 2021 · last 2025
0009-0002-3573-8104ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contrastive and self-supervised learning for open-set damage classification in structural health monitoring with incomplete and imbalanced vibration data
Sheng Shi, Dongsheng Du, Oya Mercan, Erol Kalkan, Jafarali Parol |
Expert Syst. Appl. | 1 |
| 2025 | Window normalization: Enhancing point cloud understanding by unifying inconsistent point densities
Sheng Shi, Jiahui Li 0009, Wuming Jiang, Xiangde Zhang |
Image Vis. Comput. | 2 |
| 2024 | A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar DisordersabstractPrevious research on the diagnosis of Bipolar disorder has mainly focused on resting-state functional magnetic resonance imaging. However, their accuracy can not meet the requirements of clinical diagnosis. Efficient multimodal fusion strategies have great potential for applications in multimodal data and can further improve the performance of medical diagnosis models. In this work, we utilize both sMRI and fMRI data and propose a novel multimodal diagnosis model for bipolar disorder. The proposed Patch Pyramid Feature Extraction Module extracts sMRI features, and the spatio-temporal pyramid structure extracts the fMRI features. Finally, they are fused by a fusion module to output diagnosis results with a classifier. Extensive experiments show that our proposed method outperforms others in balanced accuracy from 0.657 to 0.732 on the OpenfMRI dataset, and achieves the state of the art. Sheng Shi, Shan An, Fengmei Fan, Wenshu Ge, Feng Yu 0003, Zhiren Wang |
ICASSP | 2 |
| 2024 | An innovative prediction algorithm based on grey modeling theory and the marine predators algorithm for short-term carbon dioxide emissions in China
Wen-Ze Wu, Wanli Xie, Sheng Shi, Hegui Zhu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Mitigating biases in long-tailed recognition via semantic-guided feature transfer
Sheng Shi, Peng Wang 0095, Xinfeng Zhang 0001, Jianping Fan 0007 |
Neurocomputing | 1 |
| 2023 | AGAIN: Adversarial Training with Attribution Span Enlargement and Hybrid Feature FusionabstractThe deep neural networks (DNNs) trained by adversarial training (AT) usually suffered from significant robust generalization gap, i.e., DNNs achieve high training robustness but low test robustness. In this paper, we propose a generic method to boost the robust generalization of AT methods from the novel perspective of attribution span. To this end, compared with standard DNNs, we discover that the generalization gap of adversarially trained DNNs is caused by the smaller attribution span on the input image. In other words, adversarially trained DNNs tend to focus on specific visual concepts on training images, causing its limitation on test robustness. In this way, to enhance the robustness, we propose an effective method to enlarge the learned attribution span. Besides, we use hybrid feature statistics for feature fusion to enrich the diversity of features. Extensive experiments show that our method can effectively improves robustness of adversarially trained DNNs, outperforming previous SOTA methods. Furthermore, we provide a theoretical analysis of our method to prove its effectiveness. Shenglin Yin, Kelu Yao, Sheng Shi, Yangzhou Du |
CVPR | 3 |
| 2023 | Long-Tailed Recognition with Causal Invariant TransformationabstractStandard classification models rely on the assumption that all the classes of interest are equally represented in training datasets. However, visual phenomena exhibit a long-tailed distribution, such that many standard approaches fail to properly model and result in a considerable degeneration on accuracy. The recent methods have produced encouraging results, but their efforts only seek to simulate the statistical relationship between data and labels and compensate for imbalanced data-related issues, without addressing the underlying causal mechanisms. In this paper, a comprehensive structural causal model is developed to excavate the intrinsic causal mechanism between data and labels. Specifically, we assume that each input is constructed from a mix of causal factors and non-causal factors, and only the causal factors cause the classification judgments. In order to extract such causal factors from inputs and then reconstruct the invariant causal mechanisms, we propose a Causal Invariant Transformation algorithm for Long-tailed recognition (CITL), which generates diverse data to avoid the over-fitting on the tail classes and enforces the learnt representations to maintain the causal factors and eliminate the non-causal factors. Our extensive experimental results on several widely used datasets have demonstrated the effectiveness of our proposed CITL approach. Yahong Zhang, Sheng Shi, Yixin Wang 0003, Wenli Ouyang, WeiFan, Jianping Fan 0007 |
ICASSP | 2 |
| 2023 | Learnable Query Guided Representation Learning for Treatment Effect EstimationabstractThe estimation of Individual Treatment Effect (ITE) is a challenging problem in causal inference, due to the missing counterfactual data and the selection bias. In this paper, we propose a novel representation learning framework via the Learnable Query based transformer for Treatment Effect Estimation (LQTEE). A certain number of queries are learned in a sample-agnostic way and extract global critical features from covariate and treatment data separately. We also propose the hierarchical propensity score regularized adversarial loss to obtain balanced covariate representations, and the mutual orthogonal constraint to force queries to focus on diverse parts of covariates, thus the impact of instrumental variables can be adaptively reduced. Treatment representation learning enables our estimator to support general-purpose treatments, and more importantly, it can reveal the underlying patterns of data-generation process efficiently. Extensive experiments show that our ITE estimator significantly outperforms the state-of-the-art methods. Yixin Wang 0003, Yahong Zhang, Wenli Ouyang, Sheng Shi, Jianping Fan 0007 |
IJCNN | 5 |
| 2022 | Temporal Correlation Network for Video Polyp SegmentationabstractAccurate polyp segmentation from colonoscopy images is essential for identifying colorectal cancer. Recently, segmentation methods based on convolutional neural networks and transformers have represented excellent performance for image polyp segmentation. However, these methods are mostly designed for individual images rather than the entire video datasets, which results in the absence of sequential relationships among lesion images and neglects the significant intrinsic property of continuous video. In this work, we propose a temporal correlation network (TC-Net) for video polyp segmentation. In TC-Net, the temporal correlation is unprecedentedly modeled based on the relationship between the original video and the captured frames to be adaptable for video polyp segmentation, and the network is also calibrated for the corresponding time correlation output. Furthermore, we design a dual-track learning strategy for the optimization method in TC-Net to ensure the independence of TC-Net during the learning process to adequately exploit the optimization effect of temporal correlation. The network’s effectiveness is demonstrated by extensive experiments on five publicly available biomedical datasets, and TC-Net achieves state-of-the-art (SOTA) performance. Dehui Qiu, Senlin Lin, Sheng Shi, Shengtao Zhu, Fa Zhang 0001 |
BIBM | 5 |
| 2022 | U-GAT-VC: Unsupervised Generative Attentional Networks for Non-Parallel Voice ConversionabstractNon-parallel voice conversion (VC) is a technique of transfer-ring voice from one style to another without using a parallel corpus in model training. Various methods are proposed to approach non-parallel VC using deep neural networks. Among them, CycleGAN-VC and its variants have been widely accepted as benchmark methods. However, there is still a gap to bridge between the real target and converted voice and an increased number of parameters leads to slow convergence in training process. Inspired by recent advancements in unsupervised image translation, we propose a new end-to-end unsupervised framework U-GAT-VC that adopts a novel inter- and intra-attention mechanism to guide the voice conversion to focus on more important regions in spectrograms. We also introduce disentangle perceptual loss in our model to capture high-level spectral features. Subjective and objective evaluations shows our proposed model outperforms CycleGAN-VC2/3 in terms of conversion quality and voice naturalness. Sheng Shi, Jiahao Shao, Yifei Hao, Yangzhou Du, Jianping Fan 0007 |
ICASSP | 1 |
| 2021 | Uncertainty-Based Visual Guidance for Interactive Medical Volume Segmentation Editing
Sheng Shi, Bowei Zhou, Yibo Song |
ICIG (2) | 1 |
| 2020 | Kernel-based LIME with feature dependency samplingabstractWhile deep learning makes significant achievements in Artificial Intelligence (AI), the lack of transparency has limited its broad application in various vertical domains. Explainability is not only a gateway between AI and society, but also a powerful feature to detect flaws of the models and bias of the data. Local Interpretable Model-agnostic Explanation (LIME) is a widely-accepted technique that explains the predictions of any classifier faithfully by learning an interpretable model locally around the predicted instance. However, the sampling operation in the standard implementation of LIME is defective. Perturbed samples are generated from an uniform distribution, ignoring the complicated correlation between features. Moreover, as the local decision boundary is non-linear for most complex networks, linear approximation may produce serious errors. This paper proposes an high-interpretability and high-fidelity local explanation method, known as Kernel-based LIME with Feature Dependency Sampling (KLFDS). KLFDS enhances interpretability by feature sampling with intrinsic dependency. Besides, KLFDS improves the local explanation fidelity by approximating nonlinear boundary of local decision. We evaluate our method on image classification tasks and results show that KLFDS's explanation of the black-box model achieves much better performance than original LIME. Sheng Shi, Yangzhou Du |
ICPR | 1 |
| 2020 | Algorithm Bias Detection and Mitigation in Lenovo Face Recognition Engine
Sheng Shi, Shanshan Wei, Zhongchao Shi, Yangzhou Du, Jianping Fan 0007, Yolanda Conyers |
NLPCC (2) | 1 |
| 2019 | A Modified LIME and Its Application to Explain Service Supply Chain Forecasting
Sheng Shi, Qiang Chou |
NLPCC (2) | 3 |
| 2017 | A new two-dimensional Fourier transform algorithm based on image sparsityabstractWith the coming age of big data, the image signals play more and more important role in our life due to the extraordinary advance of network communication technology, and the corresponding high efficiency image processing techniques are demanded urgently. The Fourier transform is an important image processing tool which is used in a wide range of applications. Traditional Fourier transform algorithm computes on the value of each point of image, regardless of their properties in frequency domain. However, most image signals possess sparsity in frequency domain. In this paper, we present a new fast two-dimensional Fourier transform based on image sparsity. With hash function including a series of procedures such as random spectrum permutation, filtering and subsampling in frequency domain, the algorithm could identify and estimate the k largest coefficients quickly. In most sparse cases, the resulting algorithm performs faster than state-of-the-art fast Fourier transform algorithm, FFTW. Sheng Shi, Runkai Yang, Haihang You |
ICASSP | 1 |
| 2015 | Image compressive sensing using overlapped block projection and reconstructionabstractCompressive sensing allows a signal to be sampled at sub-Nyquist rate and still get recovered exactly, if the signal is sparse in some domain. Block compressive sensing (BCS) is advocated for practical image compressive sensing, since it processes image at block level and significantly reduces the memory requirement for storing projection matrix. However, existing BCS methods process blocks separately, which breaks the continuity between blocks and usually produces blocking artifacts. This paper proposes a new image compressive sensing scheme using overlapped-block projection and reconstruction (OBPR), in which the sampling is performed on overlapped blocks. During reconstruction, the sparsity constraint in transform domain is also enforced on the overlapped blocks. An augmented Lagrangian method is used to solve the optimization problem efficiently. Experimental results show that the proposed OBPR scheme achieves significantly better results than the existing BCS schemes in reconstruction quality. Sheng Shi, Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
ISCAS | 1 |
| 2015 | Study on subjective quality assessment of Screen Content ImagesabstractWith the coming age of big data, the cloud technology, referred to as the computations or applications through the Internet, is dramatically developed. The screen content has become one of the most common data form due to the extraordinary advance of network communication technology, and the JCT-VC has started to develop new standard focusing on improving the efficiency of screen content based on High Efficiency Video Coding (HEVC). Nevertheless, the research on the quality assessment of screen content is still quite limited at the current stage. In this paper, we present a study on subjective quality assessment of the Screen Content Images (SCIs) and investigate whether the existing objective Image Quality Assessment (IQA) methods can effectively evaluate the quality of distorted SCIs. We construct a new Screen Content Database (SCD) including 24 source SCIs and 492 compressed ones with two codecs including HEVC as well as the HEVC extension. The Single Comparison (SC) method is employed for the subjective viewing to guarantee the reliability of the results. In our experiment, the correlations of eight popular IQA methods with the obtained Mean Opinion Score (MOS) values are evaluated. The result indicates that visual information fidelity method can achieve highest consistency with human visual perception. Sheng Shi, Xiang Zhang 0004, Shiqi Wang 0001, Ruiqin Xiong, Siwei Ma 0001 |
PCS | 1 |