VLDB 2026 Research / reviewers in the wild / expert
Yuchen Yuan
dblp:169/4969
· DBLP profile ↗
30ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) are susceptible to the implicit reasoning risk, wherein innocuous unimodal inputs synergistically assemble into risky multimodal data that produce harmful outputs. We attribute this vulnerability to the difficulty of MLLMs maintaining safety alignment through long-chain reasoning.To address this issue, we introduce Safe-Semantics-but-Unsafe-Interpretation (SSUI), the first dataset featuring interpretable reasoning paths tailored for such a cross-modal challenge.A novel training framework, Safety-aware Reasoning Path Optimization (SRPO), is also designed based on the SSUI dataset to align the MLLM's internal reasoning process with human safety values. Experimental results show that our SRPO-trained models achieve state-of-the-art results on key safety benchmarks, including the proposed Reasoning Path Benchmark (RSBench), significantly outperforming both open-source and top-tier commercial MLLMs. Shujuan Liu, Jian Zhao 0006, Ziyan Shi, Yusheng Zhao, Yuchen Yuan, Chi Zhang 0012, Xuelong Li 0001 |
AAAI | 6 |
| 2026 | Visual Attention Reasoning via Hierarchical Search and Self-VerificationabstractWei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang, Ming Zhu, Haichuan Tang, Xuelong Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuchen Yuan, Haichuan Tang, Xuelong Li 0001 |
ACL (1) | 3 |
| 2026 | SAM-driven cross prompting with adaptive sampling consistency for semi-supervised medical image segmentationabstractSemi-supervised learning (SSL) has achieved notable progress in medical image segmentation. To achieve effective SSL, a model needs to be able to efficiently learn from limited labeled data and effectively exploit knowledge from abundant unlabeled data. Recent developments in visual foundation models, such as the Segment Anything Model (SAM), have demonstrated remarkable adaptability with improved sample efficiency. To seamlessly harness foundation models in SSL, we propose a SAM-driven cross prompting framework with adaptive sampling and prompt consistency for semi-supervised medical image segmentation, named CPAC-SAM. Our method employs SAM's unique prompt design and innovates a cross prompting strategy within a dual-branch framework to automatically generate prompts and supervision across two decoder branches, enabling effective learning from both scarce labeled and valuable unlabeled data. To ensure the quality of prompts for unlabeled data and provide meaningful supervision in the cross prompting scheme, we propose an innovative prototype-guided grid sampling strategy with adaptive intervals to simultaneously improve the reliability of the prompt selection area and ensure both adequate prompt density and complete target coverage. We further design a novel prompt consistency regularization to reduce SAM's prompt sensitivity and to enhance the output invariance under different prompts. We validate our method on five medical image segmentation tasks, encompassing both 2D and 3D scenarios. The extensive experiments with different labeled-data ratios and modalities demonstrate the superiority of our proposed method over the state-of-the-art SSL methods, with more than 4.1% and 3.8% Dice improvement on the breast cancer segmentation task and left atrium segmentation task, respectively. Our code is available at: https://github.com/JuzhengMiao/CPAC-SAM. Juzheng Miao, Cheng Chen 0013, Yuchen Yuan, Quanzheng Li, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2026 | Beyond similarity: Mutual information-guided retrieval for in-context learning in VQA
Zezhong Lv, Jian Zhao 0006, Yan Wang 0122, Yuchen Yuan, Yuchu Jiang, Wenqi Ren, Xuelong Li 0001 |
Pattern Recognit. | 6 |
| 2025 | Sequence-Independent Continual Test-Time Adaptation with Mixture of Incremental Experts for Cross-Domain Segmentation
Dunyuan Xu, Yuchen Yuan, Xikai Yang, Jingyang Zhang, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (16) | 2 |
| 2025 | Medical Large Vision Language Models with Multi-image Visual Ability
Xikai Yang, Juzheng Miao, Yuchen Yuan, Qi Dou 0001, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (5) | 3 |
| 2025 | DP-BICNN: A Bidirectional Information Compensation Neural Network Coupled With Data-Driven and Physical Information for Sea Surface Temperature PredictionabstractAccurate prediction of Sea Surface Temperature (SST) plays a crucial role in climate research, resource development, marine disaster prevention, and environmental protection. Traditional numerical models, while demonstrating excellent predictive accuracy, heavily rely on precise initial and boundary conditions. In contrast, data-driven approaches compensate for these shortcomings with their flexibility and lower computational costs; however, the lack of physical constraints and interpretability limits their effectiveness. Therefore, enhancing model efficiency and interpretability while ensuring accuracy has become a key challenge in current research. This paper proposes a coupled dual-stream SST prediction model that integrates numerical simulation and data-driven methods to improve predictive accuracy and physical consistency. The model includes a bidirectional information compensation module, a physical equation constraint module, and an anomaly compensation module. First, the bidirectional information compensation module constructs physical information flow and complex process flow using Conv-LSTM networks and designs an interaction mechanism between the two streams to enhance the model’s ability to capture complex processes that cannot be explained by partial differential equations. Next, the physical equation constraint module employs data assimilation techniques based on partial differential equations to simulate the dynamics of energy transfer in fluids, ensuring that the prediction results adhere to physical laws. Finally, the anomaly compensation module integrates the prediction results of physical processes and complex processes to output the final SST prediction. Experimental results demonstrate that, compared to existing methods, this model achieves superior predictive accuracy across multiple datasets. Jie Nie, Yuchen Yuan, Zhiqiang Wei 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Multi-Scale Spatio-Temporal Transformer-Based Imbalanced Longitudinal Learning for Glaucoma Forecasting From Irregular Time Series ImagesabstractGlaucoma is one of the major eye diseases that leads to progressive optic nerve fiber damage and irreversible blindness, afflicting millions of individuals. Glaucoma forecast is a good solution to early screening and intervention of potential patients, which is helpful to prevent further deterioration of the disease. It leverages a series of historical fundus images of an eye and forecasts the likelihood of glaucoma occurrence in the future. However, the irregular sampling nature and the imbalanced class distribution are two challenges in the development of disease forecasting approaches. To this end, we introduce the Multi-scale Spatio-temporal Transformer Network (MST-former) based on the transformer architecture tailored for sequential image inputs, which can effectively learn representative semantic information from sequential images on both temporal and spatial dimensions. Specifically, we employ a multi-scale structure to extract features at various resolutions, which can largely exploit rich spatial information encoded in each image. Besides, we design a time distance matrix to scale time attention in a non-linear manner, which could effectively deal with the irregularly sampled data. Furthermore, we introduce a temperature-controlled Balanced Softmax Cross-entropy loss to address the class imbalance issue. Extensive experiments on the Sequential fundus Images for Glaucoma Forecast (SIGF) dataset demonstrate the superiority of the proposed MST-former method, achieving an AUC of 96.6% for glaucoma forecasting. Besides, our method shows excellent generalization capability on the Alzheimer's Disease Neuroimaging Initiative (ADNI) MRI dataset, with an accuracy of 88.2% for mild cognitive impairment and Alzheimer's disease prediction, outperforming the compared method by a large margin. A series of ablation studies further verify the contribution of our proposed components in addressing the irregular sampled and class imbalanced problems. Xikai Yang, Xi Wang 0013, Yuchen Yuan, Jinpeng Li 0004, Guangyong Chen, Ning Li Wang, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | PAL: Boosting Skin Lesion Segmentation via Probabilistic Attribute LearningabstractSkin lesion segmentation is vital for the early detection, diagnosis, and treatment of melanoma, yet it remains challenging due to significant variations in lesion attributes (e.g., color, size, shape), ambiguous boundaries, and noise interference. Recent advancements have focused on capturing contextual information and incorporating boundary priors to handle challenging lesions. However, there has been limited exploration on the explicit analysis of the inherent patterns of skin lesions, a crucial aspect of the knowledge-driven decision-making process used by clinical experts. In this work, we introduce a novel approach called Probabilistic Attribute Learning (PAL), which leverages knowledge of lesion patterns to achieve enhanced performance on challenging lesions. Recognizing that the lesion patterns exhibited in each image can be properly depicted by disentangled attributes, we begin by explicitly estimating the distributions of these attributes as distinct Gaussian distributions, with mean and variance indicating the most likely pattern of that attribute and its variation. Using Monte Carlo Sampling, we iteratively draw multiple samples from these distributions to capture various potential patterns for each attribute. These samples are then merged through an effective attribute fusion technique, resulting in diverse representations that comprehensively depict the lesion class. By performing pixel-class proximity matching between each pixel-wise representation and the diverse class-wise representations, we significantly enhance the model's robustness. Extensive experiments on two public skin lesion datasets and one unified polyp lesion dataset demonstrate the effectiveness and strong generalization ability of our method. Codes are available at https://github.com/IsYuchenYuan/PAL. Yuchen Yuan, Xi Wang 0013, Jinpeng Li 0004, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Effective Semi-Supervised Medical Image Segmentation With Probabilistic Representations and Prototype LearningabstractLabel scarcity, class imbalance and data uncertainty are three primary challenges that are commonly encountered in the semi-supervised medical image segmentation. In this work, we focus on the data uncertainty issue that is overlooked by previous literature. To address this issue, we propose a probabilistic prototype-based classifier that introduces uncertainty estimation into the entire pixel classification process, including probabilistic representation formulation, probabilistic pixel-prototype proximity matching, and distribution prototype update, leveraging principles from probability theory. By explicitly modeling data uncertainty at the pixel level, model robustness of our proposed framework to tricky pixels, such as ambiguous boundaries and noises, is greatly enhanced when compared to its deterministic counterpart and other uncertainty-aware strategy. Empirical evaluations on three publicly available datasets that exhibit severe boundary ambiguity show the superiority of our method over several competitors. Moreover, our method also demonstrates a stronger model robustness to simulated noisy data. Code is available at https://github.com/IsYuchenYuan/PPC. Yuchen Yuan, Xi Wang 0013, Xikai Yang, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2024 | JAPO: learning join and pushdown order for cloud-native join optimization
Yuchen Yuan, Xiaoyue Feng, Jie Song 0001 |
Frontiers Comput. Sci. | 1 |
| 2024 | Introducing on-chain graph data to consortium blockchain for commercial transactions
Yuchen Yuan, Jie Song 0001, Yu Gu 0002, Qiang Qu 0001, Yongjie Bai |
Frontiers Comput. Sci. | 2 |
| 2024 | dSPG: A New Discriminant Superpixel Graph Regularizer and Convolutional Network for Hyperspectral Image ClassificationabstractSupervised hyperspectral image classification suffers from the overfitting problem when limited labels are available. Graph-based semisupervised classifiers can tackle this problem by building connections between labeled and unlabeled samples. In this work, we prove the following two propositions for an optimal graph: 1) the interclass connection weights must be 0 and 2) for a given class, a subset must contain labeled samples or be connected to the remaining subset. In a semisupervised scenario, it is very difficult to ensure that the aforementioned propositions hold. Here, we introduce a new discriminant superpixel graph (dSPG) to build a suboptimal graph, which combines a newly proposed within-superpixel graph, aimed at disconnecting pixels belonging to different classes in a superpixel (so as to decrease interclass connection weights) and a between-superpixel graph that connects spectral adjacent superpixels (to increase the intraclass subset connections). We further propose a dSPG regularizer for hyperspectral image classification and a dSPG-guided graph convolutional network (dSPGCN) to extract discriminant features. Experimental results on real hyperspectral datasets demonstrate the good performance of our newly proposed dSPG for semisupervised hyperspectral image classification. The source codes for this study are available athttps://github.com/yulong112/dSPG. Jun Li 0009, Lin He 0001, Antonio Plaza, Lizhe Wang 0001, Zhonghui Tang, Li Zhuo 0002, Yuchen Yuan |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | Semi-supervised Class Imbalanced Deep Learning for Cardiac MRI Segmentation
Yuchen Yuan, Xi Wang 0013, Xikai Yang, Ruijiang Li, Pheng-Ann Heng |
MICCAI (4) | 1 |
| 2023 | Improving vessel connectivity in retinal vessel segmentation via adversarial learning
Yuchen Yuan, Lituan Wang, Lei Zhang 0005 |
Knowl. Based Syst. | 1 |
| 2023 | MemBridge: Video-Language Pre-Training With Memory-Augmented Inter-Modality BridgeabstractVideo-language pre-training has attracted considerable attention recently for its promising performance on various downstream tasks. Most existing methods utilize the modality-specific or modality-joint representation architectures for the cross-modality pre-training. Different from previous methods, this paper presents a novel architecture named Memory-augmented Inter-Modality Bridge (MemBridge), which uses the learnable intermediate modality representations as the bridge for the interaction between videos and language. Specifically, in the transformer-based cross-modality encoder, we introduce the learnable bridge tokens as the interaction approach, which means the video and language tokens can only perceive information from bridge tokens and themselves. Moreover, a memory bank is proposed to store abundant modality interaction information for adaptively generating bridge tokens according to different cases, enhancing the capacity and robustness of the inter-modality bridge. Through pre-training, MemBridge explicitly models the representations for more sufficient inter-modality interaction. Comprehensive experiments show that our approach achieves competitive performance with previous methods on various downstream tasks including video-text retrieval, video captioning, and video question answering on multiple datasets, demonstrating the effectiveness of the proposed method. The code has been available at https://github.com/jahhaoyang/MemBridge. Xiangyang Li 0002, Mao Zheng, Xiaoqian Guo, Yuchen Yuan, Zifeng Chai, Shuqiang Jiang |
IEEE Trans. Image Process. | 7 |
| 2023 | Focus and Align: Learning Tube Tokens for Video-Language Pre-TrainingabstractVideo-language pre-training (VLP) has attracted increasing attention for cross-modality understanding tasks. To enhance visual representations, recent works attempt to adopt transformer-based architectures as video encoders. These works usually focus on the visual representations of the sampled frames. Compared with frame representations, frame patches incorporate more fine-grained spatio-temporal information, which could lead to a better understanding of video contents. However, how to exploit the spatio-temporal information within frame patches for VLP has been less investigated. In this work, we propose a method to learn tube tokens to model the key spatio-temporal information from frame patches. To this end, multiple semantic centers are introduced to focus on the underlying patterns of frame patches. Based on each semantic center, the spatio-temporal information within frame patches is integrated into a unique tube token. Complementary to frame representations, tube tokens provide detailed clues of video contents. Furthermore, to better align the generated tube tokens and the contents of descriptions, a local alignment mechanism is introduced. The experiments based on a variety of downstream tasks demonstrate the effectiveness of the proposed method. Xiangyang Li 0002, Mao Zheng, Xiaoqian Guo, Zifeng Chai, Yuchen Yuan, Shuqiang Jiang |
IEEE Trans. Multim. | 8 |
| 2022 | Spatial transcriptomics prediction from histology jointly through Transformer and graph neural networksabstractThe rapid development of spatial transcriptomics allows the measurement of RNA abundance at a high spatial resolution, making it possible to simultaneously profile gene expression, spatial locations of cells or spots, and the corresponding hematoxylin and eosin-stained histology images. It turns promising to predict gene expression from histology images that are relatively easy and cheap to obtain. For this purpose, several methods are devised, but they have not fully captured the internal relations of the 2D vision features or spatial dependency between spots. Here, we developed Hist2ST, a deep learning-based model to predict RNA-seq expression from histology images. Around each sequenced spot, the corresponding histology image is cropped into an image patch and fed into a convolutional module to extract 2D vision features. Meanwhile, the spatial relations with the whole image and neighbored patches are captured through Transformer and graph neural network modules, respectively. These learned features are then used to predict the gene expression by following the zero-inflated negative binomial distribution. To alleviate the impact by the small spatial transcriptomics data, a self-distillation mechanism is employed for efficient learning of the model. By comprehensive tests on cancer and normal datasets, Hist2ST was shown to outperform existing methods in terms of both gene expression prediction and spatial region identification. Further pathway analyses indicated that our model could reserve biological information. Thus, Hist2ST enables generating spatial transcriptomics data from histology images for elucidating molecular signatures of tissues. Yuansong Zeng, Zhuoyi Wei, Weijiang Yu, Yuchen Yuan, Bingling Li, Zhonghui Tang, Yutong Lu, Yuedong Yang |
Briefings Bioinform. | 5 |
| 2022 | Multi-Level Attention Network for Retinal Vessel SegmentationabstractAutomatic vessel segmentation in the fundus images plays an important role in the screening, diagnosis, treatment, and evaluation of various cardiovascular and ophthalmologic diseases. However, due to the limited well-annotated data, varying size of vessels, and intricate vessel structures, retinal vessel segmentation has become a long-standing challenge. In this paper, a novel deep learning model called AACA-MLA-D-UNet is proposed to fully utilize the low-level detailed information and the complementary information encoded in different layers to accurately distinguish the vessels from the background with low model complexity. The architecture of the proposed model is based on U-Net, and the dropout dense block is proposed to preserve maximum vessel information between convolution layers and mitigate the over-fitting problem. The adaptive atrous channel attention module is embedded in the contracting path to sort the importance of each feature channel automatically. After that, the multi-level attention module is proposed to integrate the multi-level features extracted from the expanding path, and use them to refine the features at each individual layer via attention mechanism. The proposed method has been validated on the three publicly available databases, i.e. the DRIVE, STARE, and CHASE _ DB1. The experimental results demonstrate that the proposed method can achieve better or comparable performance on retinal vessel segmentation with lower model complexity. Furthermore, the proposed method can also deal with some challenging cases and has strong generalization ability. Yuchen Yuan, Lei Zhang 0005, Lituan Wang, Haiying Huang 0004 |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Deep Density-Aware Count RegressorabstractWe seek to improve crowd counting as we perceive limits of currently prevalent density map estimation approach on both prediction accuracy and time efficiency. We leverage multilevel pixelation of density map as it helps improve SNR of training data and therefore, reduce prediction error. To achieve a better model, we introduce multilayer gradient fusion for training a density-aware global count regressor. More specifically, on training stage, a backbone network receives gradients from multiple branches to learn the density information, whereas those branches are to be detached to accelerate inference. By taking advantages of such method, our model improves benchmark results on public datasets and exhibits itself to be a new solution to crowd counting problems in practice. Zhuojun Chen, Yuchen Yuan, Dongping Liao, Jiancheng Lv 0001 |
ECAI | 3 |
| 2020 | HANet: Hybrid Attention-aware Network for Crowd CountingabstractAn essential yet challenging issue in crowd counting is the diverse background variations under complicated real-life environments, which makes attention based methods favorable in recent years. However, most existing methods only rely on first-order attention schemes (e.g. 2D position-wise attention), while ignoring the higher-order information within the congested scenes completely. In this paper, we propose a hybrid attention-aware network (HANet) with a high-order attention module (HAM) and an adaptive compensation loss (ACLoss) to tackle this problem. On the one hand, the HAM applies 3D attention to capture the subtle discriminative features around each people in the crowd. On the other hand, with the distributed supervision, the ACLoss exploits the prior knowledge from higher-level stages to guide the density map prediction at a lower level. The proposed HANet is then established with HAM and ACLoss working as different roles and promoting each other. Extensive experimental results show the superiority of our HANet against the state-of-the-arts on three challenging benchmarks. Xinxing Su, Yuchen Yuan, Xiangbo Su, Zhikang Zou, Shilei Wen, Pan Zhou 0001 |
ICPR | 2 |
| 2019 | Perspective-Guided Convolution Networks for Crowd CountingabstractIn this paper, we propose a novel perspective-guided convolution (PGC) for convolutional neural network (CNN) based crowd counting (i.e. PGCNet), which aims to overcome the dramatic intra-scene scale variations of people due to the perspective effect. While most state-of-the-arts adopt multi-scale or multi-column architectures to address such issue, they generally fail in modeling continuous scale variations since only discrete representative scales are considered. PGCNet, on the other hand, utilizes perspective information to guide the spatially variant smoothing of feature maps before feeding them to the successive convolutions. An effective perspective estimation branch is also introduced to PGCNet, which can be trained in either supervised setting or weakly-supervised setting when the branch has been pre-trained. Our PGCNet is single-column with moderate increase in computation, and extensive experimental results on four benchmark datasets show the improvements of our method against the state-of-the-arts. Additionally, we also introduce Crowd Surveillance, a large scale dataset for crowd counting that contains 13,000+ high-resolution images with challenging scenarios. Code is available at https://github.com/Zhaoyi-Yan/PGCNet. Zhaoyi Yan, Yuchen Yuan, Wangmeng Zuo, Xiao Tan 0001, Yezhen Wang, Shilei Wen, Errui Ding |
ICCV | 2 |
| 2019 | Recognizing Part Attributes With Insufficient DataabstractRecognizing the attributes of objects and their parts is central to many computer vision applications. Although great progress has been made to apply object-level recognition, recognizing the attributes of parts remains less applicable since the training data for part attributes recognition is usually scarce especially for internet-scale applications. Furthermore, most existing part attribute recognition methods rely on the part annotations which are more expensive to obtain. In order to solve the data insufficiency problem and get rid of dependence on the part annotation, we introduce a novel Concept Sharing Network (CSN) for part attribute recognition. A great advantage of CSN is its capability of recognizing the part attribute (a combination of part location and appearance pattern) that has insufficient or zero training data, by learning the part location and appearance pattern respectively from the training data that usually mix them in a single label. Extensive experiments on CUB, Celeb A, and a newly proposed human attribute dataset demonstrate the effectiveness of CSN and its advantages over other methods, especially for the attributes with few training samples. Further experiments show that CSN can also perform zero-shot part attribute recognition. Xiangyun Zhao, Yi Yang 0007, Feng Zhou 0002, Xiao Tan 0001, Yuchen Yuan, Sid Ying-Ze Bao, Ying Wu 0001 |
ICCV | 5 |
| 2018 | Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition
Ming Sun 0008, Yuchen Yuan, Feng Zhou 0002, Errui Ding |
ECCV (16) | 2 |
| 2018 | Compact Generalized Non-local NetworkabstractThe non-local module is designed for capturing long-range spatio-temporal dependencies in images and videos. Although having shown excellent performance, it lacks the mechanism to model the interactions between positions across channels, which are of vital importance in recognizing fine-grained objects and actions. To address this limitation, we generalize the non-local module and take the correlations between the positions of any two channels into account. This extension utilizes the compact representation for multiple kernel functions with Taylor expansion that makes the generalized non-local module in a fast and low-complexity computation flow. Moreover, we implement our generalized non-local method within channel groups to ease the optimization. Experimental results illustrate the clear-cut improvements and practical applicability of the generalized non-local module on both fine-grained object recognition and video classification. Code is available at: https://github.com/KaiyuYue/cgnl-network.pytorch. Kaiyu Yue, Ming Sun 0008, Yuchen Yuan, Feng Zhou 0002, Errui Ding, Fuxin Xu |
NeurIPS | 3 |
| 2018 | Dense and Sparse Labeling With Multidimensional Features for Saliency DetectionabstractConventional low-level feature-based saliency detection methods tend to use nonrobust prior knowledge and do not perform well in complex or low-contrast images. In this paper, to address these issues in existing methods, we propose a novel deep neural network (DNN)-based dense and sparse labeling (DSL) framework for saliency detection. DSL consists of three major steps, namely, dense labeling (DL), sparse labeling (SL), and deep convolutional (DC) network. The DL and SL steps conduct initial saliency estimations with macro object contours and low-level image features, respectively, which effectively approximate the location of the salient object and generate accurate guidance channels for the DC step; the DC step, on the other hand, takes in the results of DL and SL, establishes a six-channeled input data structure (including local superpixel information), and conducts accurate final saliency classification. Our DSL framework exploits the saliency estimation guidance from both macro object contours and local low-level features, as well as utilizing the DNN for high-level saliency feature extraction. Extensive experiments are conducted on six well-recognized public data sets against 16 state-of-the-art saliency detection methods, including ten conventional feature-based methods and six learning-based methods. The results demonstrate the superior performance of DSL on various challenging cases in terms of both accuracy and robustness. Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Reversion Correction and Regularized Random Walk Ranking for Saliency DetectionabstractIn recent saliency detection research, many graph-based algorithms have applied boundary priors as background queries, which may generate completely "reversed" saliency maps if the salient objects are on the image boundaries. Moreover, these algorithms usually depend heavily on pre-processed superpixel segmentation, which may lead to notable degradation in image detail features. In this paper, a novel saliency detection method is proposed to overcome the above issues. First, we propose a saliency reversion correction process, which locates and removes the boundary-adjacent foreground superpixels, and thereby increases the accuracy and robustness of the boundary prior-based saliency estimations. Second, we propose a regularized random walk ranking model, which introduces prior saliency estimation to every pixel in the image by taking both region and pixel image features into account, thus leading to pixel-detailed and superpixel-independent saliency maps. Experiments are conducted on four well-recognized data sets; the results indicate the superiority of our proposed method against 14 state-of-the-art methods, and demonstrate its general extensibility as a saliency optimization algorithm. We further evaluate our method on a new data set comprised of images that we define as boundary adjacent object saliency, on which our method performs better than the comparison methods. Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Image Process. | 1 |
| 2016 | Adaptive background search and foreground estimation for saliency detection via comprehensive autoencoderabstractIn saliency object detection, inappropriate boundary-background priors is known to degrade performance in challenging image datasets, and even may lead to `inverse' results when saliency regions are attached to the image boundaries. This is an active field where many works have proposed various techniques to lessen such degradation by inappropriate boundary-background priors. Although the use of boundary-background priors has shown to be capable of improving the detection, inherently, these techniques confront serious challenges in background suppression. To overcome this limitation, we propose an adaptive background extractor to search background seeds without the need of boundary-background priors. With the adaptive background seeds, the saliency objects can be then extracted via our proposed hierarchical foreground estimation model. We evaluate our adaptive Background Search and Foreground Estimation (BSFE) algorithm in comparison with six state-of-the-art methods on four well-recognized public datasets. The experimental results demonstrate that our BSFE algorithm outperforms compared methods in majority of the datasets and in particular achieves double-winners in terms of F-measure and mean absolute error on two challenging datasets. Ke Yan 0005, ChangYang Li, Xiuying Wang 0001, Yuchen Yuan, Jinman Kim, David Dagan Feng |
ICIP | 5 |
| 2016 | DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutationsabstractBACKGROUND: With the developments of DNA sequencing technology, large amounts of sequencing data have become available in recent years and provide unprecedented opportunities for advanced association studies between somatic point mutations and cancer types/subtypes, which may contribute to more accurate somatic point mutation based cancer classification (SMCC). However in existing SMCC methods, issues like high data sparsity, small volume of sample size, and the application of simple linear classifiers, are major obstacles in improving the classification performance. RESULTS: To address the obstacles in existing SMCC studies, we propose DeepGene, an advanced deep neural network (DNN) based classifier, that consists of three steps: firstly, the clustered gene filtering (CGF) concentrates the gene data by mutation occurrence frequency, filtering out the majority of irrelevant genes; secondly, the indexed sparsity reduction (ISR) converts the gene data into indexes of its non-zero elements, thereby significantly suppressing the impact of data sparsity; finally, the data after CGF and ISR is fed into a DNN classifier, which extracts high-level features for accurate classification. Experimental results on our curated TCGA-DeepGene dataset, which is a reformulated subset of the TCGA dataset containing 12 selected types of cancer, show that CGF, ISR and DNN all contribute in improving the overall classification performance. We further compare DeepGene with three widely adopted classifiers and demonstrate that DeepGene has at least 24% performance improvement in terms of testing accuracy. CONCLUSIONS: Based on deep learning and somatic point mutation data, we devise DeepGene, an advanced cancer type classifier, which addresses the obstacles in existing SMCC studies. Experiments indicate that DeepGene outperforms three widely adopted existing classifiers, which is mainly attributed to its deep learning module that is able to extract the high level features between combinatorial somatic point mutations and cancer types. Yuchen Yuan, Yi Shi 0007, ChangYang Li, Jinman Kim, Tom Weidong Cai, Zeguang Han, David Dagan Feng |
BMC Bioinform. | 1 |
| 2015 | Robust saliency detection via regularized random walks rankingabstractIn the field of saliency detection, many graph-based algorithms heavily depend on the accuracy of the pre-processed superpixel segmentation, which leads to significant sacrifice of detail information from the input image. In this paper, we propose a novel bottom-up saliency detection approach that takes advantage of both region-based features and image details. To provide more accurate saliency estimations, we first optimize the image boundary selection by the proposed erroneous boundary removal. By taking the image details and region-based estimations into account, we then propose the regularized random walks ranking to formulate pixel-wised saliency maps from the superpixel-based background and foreground saliency estimations. Experiment results on two public datasets indicate the significantly improved accuracy and robustness of the proposed algorithm in comparison with 12 state-of-the-art saliency detection approaches. ChangYang Li, Yuchen Yuan, Tom Weidong Cai, Yong Xia 0001, David Dagan Feng |
CVPR | 2 |