VLDB 2026 Research / reviewers in the wild / expert
Jun Feng 0003
dblp:00/4883-3 · also Judy Jun Feng
· DBLP profile ↗
93ranked-venue papers
6as first author
61since 2021 · last 2026
0000-0002-0706-2103ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 16 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Syn-T5: Syntax-aware fine-tuning for aspect sentiment triplet extraction
Wang Zou, Jun Feng 0003, Yaqiong Xing, Xiaodi Zhao |
Inf. Process. Manag. | 3 |
| 2026 | TABNet: A Triplet Augmentation Self-recovery framework with Boundary-aware Pseudo-labels for scribble-based medical image segmentation
Peilin Zhang, Shaoxuan Wu, Jun Feng 0003, Zhuo Jin, Zhizezhang Gao, Jingkun Chen, Yaqiong Xing, Xiao Zhang 0028 |
Image Vis. Comput. | 3 |
| 2026 | GiTNet: A graph-based trajectory-informed network for gaze-supervised medical image segmentation
Shaoxuan Wu, Xiao Zhang 0028, Jingkun Chen, Yaqiong Xing, Jun Feng 0003 |
Medical Image Anal. | 6 |
| 2026 | Toward Safe Driving: Efficient Detection of Small Blurred Signs in Real-World ScenariosabstractAccurate traffic sign recognition is critical for safe driving, as over half of traffic accidents stem from drivers’ negligence of traffic signs. Thus, developing robust traffic sign detection methods is essential to improve road safety. While existing object detection methods have achieved remarkable success, their performance in traffic sign detection is often limited by small object sizes and low-resolution appearances. To address these issues, this study proposes a novel traffic sign detection framework with three innovative components for precise localization and classification: 1) a hierarchical feature aggregation module that emphasizes high-level semantic information for traffic sign localization; 2) a cross-layer semantic residual network that enhances recognition of small and blurred signs via hierarchical feature interaction and fusion; 3) a lightweight feature alignment unit that bridges semantic gaps between cross-level representations. These components jointly tackle the challenges of detecting small and blurred traffic signs in real-world driving scenarios. Experiments were conducted on three public datasets (TT100K, CCTSDB2021, GTSDB). Results show the proposed model outperforms other methods with comparable parameters. Additionally, dynamic motion blur augmentation was applied to datasets to simulate real driving scenarios, and experiments confirm the proposed method achieves state-of-the-art performance under such challenging conditions. Code is publicly available athttps://github.com/Mo7nex/SAttFusion-YOLO Dekui Wang, Jun Feng 0003, Qirong Bo, Yaqiong Xing, Wei Zhou 0012, Xingxing Hao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | A Coarse-to-Fine Progressive Ensemble Framework for Coronary Artery LabelingabstractAutomatic coronary artery labeling is essential for accurate vascular identification and the diagnosis of coronary disease. The task requires delineating the full vasculature and classifying each segment; however, preserving global topology and local demarcation line precision is difficult due to complex anatomy and blurry contours. We propose a coarse-to-fine ensemble framework with two modules: a Coarse-to-fine Topology Extraction (CTE) network using topology priors for global continuity, and a Progressive Vessel Labeling (PVL) module with multibranch fusion for segmentation and classification. Experiments on the ARCADE dataset achieve a mean F1-score of 0.6028, outperforming state-of-the-art methods and enhancing topological integrity and labeling accuracy. Code: https://github.com/IPMINWU/PGSMODEL. Guansheng Peng, Zhuo Jin, Shaoxuan Wu, Yuhao Dong, Xiao Zhang 0028, Jun Feng 0003 |
BIBM | 7 |
| 2025 | Structural Points Dependency-Aware Template-Free Learning for Cardiac Mesh ReconstructionabstractHigh-fidelity, patient-specific cardiac mesh reconstruction underpins diagnosis, surgical planning, and hemodynamic simulation. Accurate and topologically coherent reconstruction remains challenging due to large inter-individual anatomical variability and complex cardiac morphology. We propose a template-free framework, TFSG, that integrates an Adaptive Structural Point Generation (ASG) module and a Structural Consistency Constraint (SCC). ASG extracts patientspecific anatomical landmarks from a point-cloud representation to guide deformation, while SCC enforces multi-level consistency (point distance, normal alignment and structural-point relations) to suppress topological and structural errors. Experiments on the CARE2025 WHS dataset show TFSG improves segmentation and mesh reconstruction quality compared to prior methods. Code: https://github.com/IPMI-NWU/TFSG. Shaoxuan Wu, Peilin Zhang, Yuhao Dong, Xiao Zhang 0028, Jun Feng 0003 |
BIBM | 7 |
| 2025 | Graph-Based Neighbor-Aware Network for Gaze-Supervised Medical Image Segmentation
Shaoxuan Wu, Jingkun Chen, Zhuo Jin, Peilin Zhang, Zhizezhang Gao, Jun Feng 0003, Xiao Zhang 0028, Dinggang Shen |
MICCAI (4) | 6 |
| 2025 | Rodecon-net: Medical Image Segmentation via Robust Decoupling and Contrast-enhanced FusionabstractMedical image segmentation is crucial for clinical decision-making, treatment planning, and disease tracking. Nonetheless, it confronts two significant challenges: the presence of ''soft boundaries'' between the foreground and background exacerbated by poor illumination and low contrast, and the misleading co-occurrence of salient and non-salient objects during the training phase, which complicates the model's accuracy in distinguishing relevant features. To overcome these challenges, we introduce RoDeCon-Net, a novel framework engineered to enhance medical image segmentation. RoDeCon-Net incorporates a Feature Decoupling Unit (FDU) that dynamically separates encoded features into foreground, background, and uncertain regions, using advanced attention mechanisms to refine feature distinction and reduce uncertainty. Additionally, our Contrast-driven Feature Alignment Unit (CFAU) and Cross-layer Feature Cascade Unit (CFCU) synergize to reinforce feature contrasts and promote effective multi-level feature fusion, thus improving the detection of salient objects amidst complex backgrounds and handling various object scales within images. Comprehensive evaluations of RoDeCon-Net on five diverse medical image datasets validate its superior performance and versatility, showcasing its potential to set new benchmarks in medical image segmentation. Our code is available on https://github.com/ILoveACM-MM/RoDeCon-Net. Yongquan Xue, Zhaoru Guo, Zhaozhao Su, Chong Peng 0001, Jun Feng 0003, Pan Zhou 0001, Marcin Pietron, Panpan Zheng |
ACM Multimedia | 5 |
| 2025 | Towards a Quantitative Competency Model for CS1 via Five-Channel Learning Sequences
Zhizezhang Gao, Can Cui 0016, Haochen Yan, Jun Feng 0003 |
SIGCSE (1) | 6 |
| 2025 | Context-enhanced framework for medical image report generation using multimodal contexts
Hongzhao Li, Hongyu Wang 0007, Jun Feng 0003 |
Knowl. Based Syst. | 5 |
| 2025 | HELPNet: Hierarchical perturbations consistency and entropy-guided ensemble for scribble supervised medical image segmentation
Xiao Zhang 0028, Shaoxuan Wu, Peilin Zhang, Zhuo Jin, Xiaosong Xiong, Qirong Bu, Jingkun Chen, Jun Feng 0003 |
Medical Image Anal. | 8 |
| 2025 | A vision and language hierarchical alignment for multimodal aspect-based sentiment analysis
Wang Zou, Qiang Lu 0006, Xuxin Wang, Jun Feng 0003 |
Pattern Recognit. | 5 |
| 2025 | Exploring Unbiased Activation Maps for Weakly Supervised Tissue Segmentation of Histopathological ImagesabstractTissue segmentation in histopathological images plays a crucial role in computational pathology, owing to its significant potential to indicate the prognosis of cancer patients. Presently, numerous Weakly Supervised Semantic Segmentation (WSSS) methods strive to utilize image-level labels to achieve pixel-level segmentation, aiming to minimize the need for detailed annotations. Most of these methods rely on Class Activation Maps (CAM) extracted from classification models, frequently leading to poor coverage of objects. The major cause is attributed to the strong inductive bias of the classification model, focusing primarily on discriminative feature of objects, rather than non-discriminative features. Inspired by this, we propose a simple yet effective method that introduces a self-supervised task by exploiting both the discriminative and non-discriminative features, and generate Unbiased Activation Maps (UAM) to encompass the whole object. Specifically, our method entails clustering all spatial features of an object class to derive semantic centers. Each center then works as a spatial filter that amplifies similar feature and suppresses dissimilar feature, and extract high-quality pseudo-labels (some noise at object boundaries). Moreover, we further propose a Noise-Reduced (NR) Learning method to train the segmentation network towards credible signals and lessen the impact of false predictions. Comprehensive experimental results on two public histopathology image datasets demonstrate the superior performance of our method over the state-of-the-art weakly supervised segmentation methods. Yuxin Kang, Hansheng Li, Xiaoshuang Shi, Xiao Zhang 0028, Yaqiong Xing, Yuting Wen, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 9 |
| 2024 | Prompt-Guided Generation of Structured Chest X-Ray Report Using a Pre-trained LLMabstractMedical report generation automates radiology descriptions from images, easing the burden on physicians and minimizing errors. However, current methods lack structured outputs and physician interactivity for clear, clinically relevant reports. Our method introduces a prompt-guided approach to generate structured chest X-ray reports using a pre-trained large language model (LLM). First, we identify anatomical regions in chest X-rays to generate focused sentences that center on key visual elements, thereby establishing a structured report foundation with anatomy-based sentences. We also convert the detected anatomy into textual prompts conveying anatomical comprehension to the LLM. Additionally, the clinical context prompts guide the LLM to emphasize interactivity and clinical requirements. By integrating anatomy-focused sentences and anatomy/clinical prompts, the pre-trained LLM can generate structured chest X-ray reports tailored to prompted anatomical regions and clinical contexts. We evaluate using language generation and clinical effectiveness metrics, demonstrating strong performance. Hongzhao Li, Hongyu Wang 0007, Jun Feng 0003 |
ICME | 5 |
| 2024 | Gaze-Directed Vision GNN for Mitigating Shortcut Learning in Medical Image
Shaoxuan Wu, Xiao Zhang 0028, Zhuo Jin, Hansheng Li, Jun Feng 0003 |
MICCAI (1) | 6 |
| 2024 | GO: A two-step generative optimization method for point cloud registration
Yan Zhao 0042, Jiahui Deng, Feihong Liu, Wen Tang 0004, Jun Feng 0003 |
Comput. Graph. | 5 |
| 2024 | Coordinated-joint translation fusion framework with sentiment-interactive graph convolutional networks for multimodal sentiment analysis
Qiang Lu 0006, Zhizezhang Gao, Jun Feng 0003, Hao Zhang 0202 |
Inf. Process. Manag. | 5 |
| 2024 | An Anatomy- and Topology-Preserving Framework for Coronary Artery SegmentationabstractCoronary artery segmentation is critical for coronary artery disease diagnosis but challenging due to its tortuous course with numerous small branches and inter-subject variations. Most existing studies ignore important anatomical information and vascular topologies, leading to less desirable segmentation performance that usually cannot satisfy clinical demands. To deal with these challenges, in this paper we propose an anatomy- and topology-preserving two-stage framework for coronary artery segmentation. The proposed framework consists of an anatomical dependency encoding (ADE) module and a hierarchical topology learning (HTL) module for coarse-to-fine segmentation, respectively. Specifically, the ADE module segments four heart chambers and aorta, and thus five distance field maps are obtained to encode distance between chamber surfaces and coarsely segmented coronary artery. Meanwhile, ADE also performs coronary artery detection to crop region-of-interest and eliminate foreground-background imbalance. The follow-up HTL module performs fine segmentation by exploiting three hierarchical vascular topologies, i.e., key points, centerlines, and neighbor connectivity using a multi-task learning scheme. In addition, we adopt a bottom-up attention interaction (BAI) module to integrate the feature representations extracted across hierarchical topologies. Extensive experiments on public and in-house datasets show that the proposed framework achieves state-of-the-art performance for coronary artery segmentation. Xiao Zhang 0028, Kaicong Sun, Dijia Wu, Xiaosong Xiong, Jiameng Liu, Linlin Yao, Shufang Li, Jun Feng 0003, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2024 | Sentiment Analysis: Comprehensive Reviews, Recent Advances, and Open ChallengesabstractSentiment analysis (SA) aims to understand the attitudes and views of opinion holders with computers. Previous studies have achieved significant breakthroughs and extensive applications in the past decade, such as public opinion analysis and intelligent voice service. With the rapid development of deep learning, SA based on various modalities has become a research hotspot. However, only individual modality has been analyzed separately, lacking a systematic carding of comprehensive SA methods. Meanwhile, few surveys covering the topic of multimodal SA (MSA) have been explored yet. In this article, we first take the modality as the thread to design a novel framework of SA tasks to provide researchers with a comprehensive understanding of relevant advances in SA. Then, we introduce the general workflows and recent advances of single-modal in detail, discuss the similarities and differences of single-modal SA in data processing and modeling to guide MSA, and summarize the commonly used datasets to provide guidance on data and methods for researchers according to different task types. Next, a new taxonomy is proposed to fill the research gaps in MSA, which is divided into multimodal representation learning and multimodal data fusion. The similarities and differences between these two methods and the latest advances are described in detail, such as dynamic interaction between multimodalities, and the multimodal fusion technologies are further expanded. Moreover, we explore the advanced studies on multimodal alignment, chatbots, and Chat Generative Pre-trained Transformer (ChatGPT) in SA. Finally, we discuss the open research challenges of MSA and provide four potential aspects to improve future works, such as cross-modal contrastive learning and multimodal pretraining models. Qiang Lu 0006, Zhizezhang Gao, Jun Feng 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Addressing Sparse Annotation: a Novel Semantic Energy Loss for Tumor Cell Detection from Histopathologic ImagesabstractTumor cell detection plays a vital role in immunohistochemistry (IHC) quantitative analysis. While recent remarkable developments in fully-supervised deep learning have greatly contributed to the efficiency of this task, the necessity for manually annotating all cells of specific detection types remains impractical. Obviously, if we directly use full supervision to train these datasets, it can cause error in loss calculation due to the misclassification of unannotated cells as background. To address this issue, we observe that although some cells are omitted during the annotation process, these unannotated cells have a significant feature similarity with the annotated ones. Leveraging this characteristic, we propose a novel calibrated loss named Semantic Energy Loss (SEL). Specifically, our SEL automatically adjusts the loss to be lower for unannotated regions with similar semantic to the labeled ones, while penalizing regions with lager semantic difference. Besides, to prevent all regions from having similar semantics during training, we propose Stretched Feature Loss (SFL) that widen the semantic distance. We evaluate our method on two different IHC datasets and achieve significant performance improvements in both sparse and exhaustive annotation scenarios. Furthermore, we also validate that our method holds significant potential for detecting multiple types of cells. Our code is available at here. Xianglong Du, Yuxin Kang, Hong Lv, Lei Cui 0004, Hansheng Li, Yaqiong Xing, Jun Feng 0003, Lin Yang 0002 |
BIBM | 10 |
| 2023 | Speech Emotion Recognition Based on Low-Level Auto-Extracted Time-Frequency FeaturesabstractDeep-learning based methods that aim to extract effective high-level features have steadily improved the performance on the speech emotion recognition. However, low-level features that contain important emotion-related information have not gained much attention. In this paper, we propose a novel low-level feature extraction method based on the Time-Frequency Attention (TFA) module and Time-Frequency Weighting (TFW) module. First, the TFA module is designed to learn notable regions in the detail-rich low-level feature maps produced by the scale-specific convolutional layers. Then, the TFW module is proposed to extract discriminative features from the time and frequency dimensions respectively. Finally, the speech emotion recognition task is completed by the subsequent multi-branch network. Experimental results on the IEMOCAP and RAVDESS datasets demonstrate the importance of low-level features, and show that the proposed method outperforms other state-of-the-art approaches. Jingzhao Hu, Jun Feng 0003 |
ICASSP | 3 |
| 2023 | Speech Emotion Recognition via Heterogeneous Feature LearningabstractSpeech emotion recognition (SER) based on multi-view learning has made some progress on speaker-independent scenarios. How-ever, the existing SER methods always rely on excessive feature views and ignore the importance of heterogeneous feature learning. In this paper, we propose a novel multi-level attention method to effectively learn the heterogeneous information from the hand-crafted feature (MFCC) and the feature (W2V2) extracted from the pre-trained model. Specifically, we first design an Attention based Multi-scale Low-level Feature (A-MLF) extractor to extract scale-specific emotion-related regions from MFCC. Then, the Multi-Unit Attention (MUA) module is used to simultaneously learn discriminative features in three different dimensions. Finally, a two-stage feature fusion strategy is used for joint representation space learning. We demonstrate our method on two speaker-independent validation strategies and interpret the SOTA performance by visualizing the feature distribution. Dongya Wu, Dekui Wang, Jun Feng 0003 |
ICASSP | 4 |
| 2023 | Speech Emotion Recognition Via Two-Stream Pooling Attention With Discriminative Channel WeightingabstractMulti-view Speech Emotion Recognition (SER) based on the pre-trained model has achieved success in speaker-independent scenarios. However, the existing SER methods rely on excessive feature views and have complicated feature fusion strategies. In this paper, we propose a novel method to learn effective emotion-related information from two feature views. First, we present a Discriminative Channel Weighting (DCW) module to weight the channel dimension of the features produced by a set of multi-scale convolution layers. This module allows for discriminative weighting of complex channel dimensions. Second, a concise Two-stream Pooling Attention (TsPA) strategy is proposed to generate two groups of fusion features based on different channel-level embeddings with different emphasis. Finally, the SER task is completed by three consecutive fully connected layers. The effectiveness of the proposed method has been demonstrated on two speaker-independent validation strategies, outperforming other state-of-the-art approaches. Dekui Wang, Dongya Wu, Jun Feng 0003 |
ICASSP | 4 |
| 2023 | Local Feature Enhanced Adversarial Network for the Blind Image Quality AssessmentabstractAs a hot research topic in the field of computer vision, blind image quality assessment (BIQA) can provide high-quality images for end-users and promote the development of other fields of computer vision. Although the existing BIQA based on convolution neural networks has made significant progress in synthetic distortion evaluation, it still cannot be well extended to authentic distortion and algorithm-related distortion. Therefore, this paper proposes a BIQA adversarial network with local feature enhancement to deal with this challenge. First, the ResNeSt50 network with local feature enhancement is used to extract the features of images, which effectively combines the overall semantic information and the local features of the images. Then, the mapping of distorted images to their quality scores is learned by the adversarial network. Extensive experiments demonstrate that the proposed method performs best on three categories of distorted scenario databases (nine databases) compared with state-of-the-art BIQA methods. Xiaomei Shi, Shouhai Xia, Ruxue Zhang, Jun Feng 0003 |
ICASSP | 5 |
| 2023 | A Sentiment and Syntactic-Aware Graph Convolutional Network for Aspect-Level Sentiment ClassificationabstractAspect-level sentiment classification (ASC) is a significant problem in fine-grained sentiment analysis, which automatically predicts the sentiment polarity of a given aspect in a sentence. Dependency tree-based graph convolutional networks have been widely studied for their ability to effectively capture the dependencies of aspect words with other words. However, constructing more accurate syntactic trees by introducing external knowledge has limited improvement on ungrammatical informal texts and has led to over-parameterization of the model. To alleviate this problem, we propose a sentiment and syntactic-aware graph convolutional network (SaS-GCN) that combines syntactic and sentiment relations. We use an attention mechanism and the Sparsemax activation function to construct a sparse sentiment-dependent graph. Compared with existing methods that use LSTM or CNN to obtain semantics from text directly, this graph, combined with a GCN, contains more semantic features. Moreover, we redesign the network structure of GCN, calling it EN-GCN, to make it sensitive to node dimensional features and hence to have a strong feature mining ability. The experimental results indicate that our model outperforms state-of-the-art methods. In particular, when evaluated on the Rest15 and Rest16 datasets, the F1 scores of the proposed lightweight model are 4.15% and 3.77% better than BERT respectively. Qiang Lu 0006, Richard F. E. Sutcliffe, Jun Feng 0003 |
ICASSP | 5 |
| 2023 | Improving Chinese Spelling Correction by RankingabstractChinese Spelling Check (CSC) aims to detect and correct Chinese spelling errors. Most Chinese spelling errors are the misuse of semantically, phonetically or graphically similar characters. Previous state-of-the-art works on the CSC task pursue transitions from misspelled sentences to correct sentences directly. However, it is difficult to force the current CSC methods to find the correct answer at one run. Thus, we propose a simple and effective method for CSC task by making fully use of the trained model to generate multiple candidate sentences and simply ranking to select the best, in which no additional training and parameters are required. The experimental results show that our approach outperforms previous methods and achieves the state-of-the-art performances. Jun Feng 0003, Wenbiao Yin, Lin Shang 0001 |
IJCNN | 1 |
| 2023 | Utilization of Spatio-Temporal and Social Information for POI Group RecommendationabstractPOI group recommendation offers a list of locations to a group of users based on their visiting preferences, which is crucial for Location-Based Social Networks(LBSNs) to improve user experience quality and group satisfaction. Current studies either regard the POI as a general item for group recommendation, or do not make full use of the geospatial information of POIs and users’ social friends’ information to generate the group visiting preference representations. In this paper, a neural network-based POI group recommendation method named STSPGR is proposed. STSPGR first learns the user embedding vector that fuses temporal, spatial, categorical, and social information to represent each user’s visiting preference. Then it utilizes the attention network to dynamically learn the impact degrees of each member in the group decision-making process to aggregate the group visiting preference embedding. Finally, the POI recommender decodes the group’s embedding to the preference scores over all POIs to make the recommendation. Experiments are conducted on three real-world datasets, which show that the proposed STSPGR has better recommendation accuracy than other POI group recommendation methods. Pengyu Niu, Boting Qu, Jun Feng 0003 |
MDM | 3 |
| 2023 | Segment Membranes and Nuclei from Histopathological Images via Nuclei Point-Level Supervision
Hansheng Li, Xiaoshuang Shi, Yuxin Kang, Qirong Bu, Hong Lv, Mingzhen Lin, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
MICCAI (6) | 11 |
| 2023 | SPR-Net: Structural Points Based Registration for Coronary Arteries Across Systolic and Diastolic Phases
Xiao Zhang 0028, Feihong Liu, Yuning Gu, Xiaosong Xiong, Caiwen Jiang, Jun Feng 0003, Dinggang Shen |
MICCAI (7) | 6 |
| 2023 | Rib segmentation algorithm for X-ray image based on unpaired sample augmentation and multi-scale network
Hongyu Wang 0007, Songtao Ding, Zhanyi Gao, Jun Feng 0003, Shaohua Wan 0001 |
Neural Comput. Appl. | 5 |
| 2023 | Wse-MF: A weighting-based student exercise matrix factorization modelabstractStudents who have been taught new ideas need to develop their skills by carrying out further work in their own time. This often consists of a series of exercises which must be completed. While students can choose exercises themselves from online sources, they will learn more quickly and easily if the exercises are specifically tailored to their needs. A good teacher will always aim to do this, but with the large groups of students who typically take advantage of open online courses, it may not be possible. Exercise prediction, working with large-scale matrix data, is a better way to address this challenge, and a key stage within such prediction is to calculate the probability that a student will answer a given question correctly. Therefore, this paper presents a novel approach called Weighting-based Student Exercise Matrix Factorization (Wse-MF) which combines student learning ability and exercise difficulty as prior weights. In order to learn how to complete the matrix, we apply an iterative optimization method that makes the approach practical for large-scale educational deployment. Compared with eight models in cognitive diagnosis and matrix factorization, our research results suggest that Wse-MF significantly outperforms the state-of-the-art on a range of real-world datasets in both prediction quality and time complexity. Moreover, we find that there is an optimal value of the latent factor K (the inner dimension of the factorization) for each dataset, which is related to the relationship between skills and exercises in that dataset. Similarly, the optimal value of hyperparameter c0 is linked to the ratio between exercises and students. Taken as a whole, we demonstrate improvements to matrix factorization within the context of educational data. Richard F. E. Sutcliffe, Zhizezhang Gao, Wenying Kang, Jun Feng 0003 |
Pattern Recognit. | 6 |
| 2023 | A Fast FPGA Connection Router Using Prerouting-Based Parallel Local Routing AlgorithmabstractRouting is one of the most time-consuming steps in the field-programmable gate array (FPGA) design process. Even if unceasing efforts have been made to accelerate FPGA routing, the existing work seldom pays attention to the underlying FPGA connection router. In this article, we present a fast FPGA connection router called PRoute which implements a novel prerouting-based parallel local routing algorithm. Basically, PRoute precomputes the potential routing solutions for various connection patterns on FPGAs, which can be directly used in the later practical routing. On the whole, PRoute is composed of A-star maze expansion and parallel local search. In the first part, PRoute gradually expands the maze wavefront toward the lowest-cost node to search for the target sink. For a wire-type node, PRoute invokes a fast parallel local search instead taking advantage of the prerouting results, and hence the time-expensive maze expansion can be reduced. Particularly, it allows PRoute to call one another between A-star maze expansion and parallel local search. This enables the runtime efficiency of PRoute while ensuring its global search ability. In addition, we put forward an engineering improvement to further speed up PRoute by avoiding the exploration of block output pins. To our best knowledge, this work is the first to apply the idea of prerouting for FPGAs. Experimental results show that PRoute achieves speedups of$1.8\times $,$2.4\times $,$3.2\times $,$4.1\times $, and$5.1\times $with 1, 4, 8, 16, and 32 threads relative to the baseline versatile place to route’s connection router, respectively, without degrading the quality of results. Dekui Wang, Jun Feng 0003, Wei Zhou 0012, Xingxing Hao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | FCRoute: A Fast FPGA Connection Router Using Soft Routing-Space Pruning AlgorithmabstractRouting is one of the most time-consuming stages in the field-programmable gate array (FPGA) design flow. Even if various attempts have been made to reduce route time, the existing work rarely focuses on improving the underlying A*-based FPGA connection router. In this article, we present a fast FPGA connection router called FCRoute based on a novel soft routing-space pruning algorithm. Within FCRoute, a routing resource priority mechanism is applied to classify the routing resource nodes into high-priority nodes and low-priority ones. On the whole, FCRoute is composed of fast maze search and backtracking process. During the fast maze search, we explore only the high-priority nodes in the routing space. In this way, a great deal of unnecessary work can be avoided. When the fast maze search fails to find the target sink, it allows the backtracking process to explore the low-priority nodes promising to be on the best path, after which a new fast maze search is called. By avoiding the exploration of the majority of low-priority nodes, FCRoute maintains runtime efficiency while ensuring global search ability. In addition, we further accelerate FCRoute with an engineering enhancement which simplifies the cost computations of nodes. Runtime and quality of results are compared with the state-of-the-art connection router in VPR 8. Experimental results show that on average FCRoute explores less than half the number of routing resource nodes, and therefore reduces runtime by 38% while enabling the quality of results. When combined with the enhancement, FCRoute achieves an average 45% reduction on runtime without sacrificing the quality of results. Dekui Wang, Jun Feng 0003, Wei Zhou 0012, Xingxing Hao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Invariant Content Synergistic Learning for Domain Generalization on Medical Image SegmentationabstractAlthough deep convolution neural networks (DC-NNs) can achieve remarkable success on medical image segmentation, their performance might significantly deteriorate when confronting testing data with the new distribution. Recent studies suggest that one major cause of this issue is the strong inductive bias of DCNNs, which towards image styles (e.g., superficial texture) that are sensitive to change, instead of the invariant content (e.g., object shapes). Inspired by this, we propose a novel method, named Invariant Content Synergistic Learning (ICSL), to improve the generalization ability of DCNNs on unseen data by controlling the inductive bias. Specifically, ICSL first mixes the style of training instances to perturb the training distribution, so that more diverse domains or styles would be made available for training DCNNs. Then, based on the perturbed distribution, we carefully design a dual-branches invariant content synergistic learning strategy to prevent style-biased predictions and maintain the invariant content. Extensive experimental results demonstrate the superior performance of the proposed method over state-of-the-art domain generalization methods on two typical medical segmentation tasks. Yuxin Kang, Hansheng Li, Xiaoshuang Shi, Feihong Liu, Qingguo Yan, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 9 |
| 2022 | A Random Feature Augmentation for Domain Generalization in Medical Image SegmentationabstractDeep convolutional neural networks (DCNNs) significantly improve the performance of medical image segmentation. Nevertheless, medical images frequently experience distribution discrepancies, which fails to maintain their robustness when applying trained models to unseen clinical data. To address this problem, domain generalization methods were proposed to enhance the generalization ability of DCNNs. Feature space-based data augmentation methods have proven their effectiveness to improve domain generalization. However, existing methods still mainly rely on certain prior knowledge or assumption, which has limitations in enriching the diversity of source domain data. In this paper, we propose a random feature augmentation (RFA) method to diversify source domain data at the feature level without prior knowledge. Specifically, we explore the effectiveness of random convolution at the feature level for the first time and prove experimentallyt hat itc an adequately preserve domain-invariant information while perturbing domainspecific information. Furthermore, tocapture the same domain-invariant information from the augmented features of RFA, we present a domain-invariant consistent learning strategy to enable DCNNs to learn a more generalized representation. Our proposed method achieves state-of-the-art performance on two medical image segmentation tasks, including optic cup/disc segmentation on fundus images and prostate segmentation on MRI images. Yuxin Kang, Hansheng Li, Jiayu Luo, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 6 |
| 2022 | ASRL: An Adaptive GPS Sampling Method Using Deep Reinforcement LearningabstractWith the development of positioning technology, massive GPS trajectory data are obtained to provide location-based services. However, most GPS trajectory data are obtained by a fixed sampling rate, which may lead to tremendous data redundancy, causing huge communication overhead, storing and computing issues, and high battery consumptions of mobile devices. In this paper, we propose an Adaptive Sampling method by Reinforcement Learning called ASRL. ASRL adjusts the sampling rate based on object moving status, aiming at reducing the size of the GPS trajectory without sacrificing the tracking accuracy. ASRL follows an actor-critic reinforcement learning framework and learns a GPS sampling policy network. The proper reward function in ASRL is generated by utilizing the Inverse Reinforcement Learning (IRL), which learns from the map matching results of historical trajectories and estimates the importance of each object moving status feature from demonstrations. The proposed ASRL method is evaluated by using three real GPS trajectory datasets. The result shows that the ASRL method can reduce more than 95% of GPS points while keeping the reasonable trajectory accuracy. Boting Qu, Mengjiao Zhao, Jun Feng 0003 |
MDM | 3 |
| 2022 | Progressive Deep Segmentation of Coronary Artery via Hierarchical Topology Learning
Xiao Zhang 0028, Jingyang Zhang, Lei Ma 0006, Peng Xue 0005, Dijia Wu, Yiqiang Zhan, Jun Feng 0003, Dinggang Shen |
MICCAI (5) | 8 |
| 2022 | Time-Frequency Attention for Speech Emotion Recognition with Squeeze-and-Excitation Blocks
Chen Wang 0093, Jiayue Chen, Jun Feng 0003 |
MMM (1) | 4 |
| 2022 | BBW: a batch balance wrapper for training deep neural networks on extremely imbalanced datasets with few minority samplesabstractAbstract In recent years, Deep Neural Networks (DNNs) have achieved excellent performance on many tasks, but it is very difficult to train good models from imbalanced datasets. Creating balanced batches either by majority data down-sampling or by minority data up-sampling can solve the problem in certain cases. However, it may lead to learning process instability and overfitting. In this paper, we propose the Batch Balance Wrapper (BBW), a novel framework which can adapt a general DNN to be well trained from extremely imbalanced datasets with few minority samples. In BBW, two extra network layers are added to the start of a DNN. The layers prevent overfitting of minority samples and improve the expressiveness of the sample distribution of minority samples. Furthermore, Batch Balance (BB), a class-based sampling algorithm, is proposed to make sure the samples in each batch are always balanced during the learning process. We test BBW on three well-known extremely imbalanced datasets with few minority samples. The maximum imbalance ratio reaches 1167:1 with only 16 positive samples. Compared with existing approaches, BBW achieves better classification performance. In addition, BBW-wrapped DNNs are 16.39 times faster, relative to unwrapped DNNs. Moreover, BBW does not require data preprocessing or additional hyper-parameter tuning, operations that may require additional processing time. The experiments prove that BBW can be applied to common applications of extremely imbalanced data with few minority samples, such as the classification of EEG signals, medical images and so on. Jingzhao Hu, Hao Zhang 0202, Richard F. E. Sutcliffe, Jun Feng 0003 |
Appl. Intell. | 5 |
| 2022 | MKPM: Multi keyword-pair matching for natural language sentencesabstractAbstract Sentence matching is widely used in various natural language tasks, such as natural language inference, paraphrase identification and question answering. For these tasks, we need to understand the logical and semantic relationship between two sentences. Most current methods use all information within a sentence to build a model and hence determine its relationship to another sentence. However, the information contained in some sentences may cause redundancy or introduce noise, impeding the performance of the model. Therefore, we propose a sentence matching method based on multi keyword-pair matching (MKPM), which uses keyword pairs in two sentences to represent the semantic relationship between them, avoiding the interference of redundancy and noise. Specifically, we first propose a sentence-pair-based attention mechanismsp-attentionto select the most important word pair from the two sentences as a keyword pair, and then propose a Bi-task architecture to model the semantic information of these keyword pairs. The Bi-task architecture is as follows: 1. In order to understand the semantic relationship at the word level between two sentences, we design a word-pair task (WP-Task), which uses these keyword pairs to complete sentence matching independently. 2. We design a sentence-pair task (SP-Task) to understand the sentence level semantic relationship between the two sentences by sentence denoising. Through the integration of the two tasks, our model can understand sentences more accurately from the two granularities of word and sentence. Experimental results show that our model can achieve state-of-the-art performance in several tasks. Our source code is publicly available1. Yi Gao 0003, Jun Feng 0003, Richard F. E. Sutcliffe |
Appl. Intell. | 5 |
| 2022 | General discriminative optimization for point set registration
Yan Zhao 0042, Wen Tang 0004, Jun Feng 0003, Tao Ruan Wan, Long Xi 0001 |
Comput. Graph. | 3 |
| 2022 | TransferSense: towards environment independent and one-shot wifi sensing
Qirong Bu, Xingxia Ming, Jingzhao Hu, Jun Feng 0003 |
Pers. Ubiquitous Comput. | 5 |
| 2022 | Deep transfer learning for gesture recognition with WiFi signals
Qirong Bu, Xingxia Ming, Jun Feng 0003 |
Pers. Ubiquitous Comput. | 5 |
| 2022 | Speech Emotion Recognition via Multi-Level Attention NetworkabstractAiming to improve the performance of human speech emotion recognition (SER), the existing work has made great progress based on the popular mel-scale frequency cepstral coefficient (MFCC). However, the existing work rarely pays attention to the low-level emotion related features in MFCC, such as the underlying interactive relations. In this letter, we propose a novel multi-level attention network (MLAnet), which contains a multi-scale low-level feature (MLF) extractor and a multi-unit attention (MUA) module. Within the MLF extractor, we minimize the task-irrelevant information which harms the performance of SER by applying the attention mechanism. Since the features extracted by the MLF extractor contain rich domain-specific emotion information, we further present a MUA module to simultaneously weight the features in terms of time, frequency and channel dimensions. In this way, the discriminative emotion features in different dimensions can be extracted by corresponding weighting blocks. Experimental results on two benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art approaches. Dekui Wang, Dongya Wu, Jun Feng 0003 |
IEEE Signal Process. Lett. | 5 |
| 2022 | A Novel Encoding and Decoding Calibration Guiding Pathway for Pathological Image AnalysisabstractDiagnostic pathology is the foundation and gold standard for identifying carcinomas, and the accurate quantification of pathological images can provide objective clues for pathologists to make more convincing diagnosis. Recently, the encoder-decoder architectures (EDAs) of convolutional neural networks (CNNs) are widely used in the analysis of pathological images. Despite the rapid innovation of EDAs, we have conducted extensive experiments based on a variety of commonly used EDAs, and found them cannot handle the interference of complex background in pathological images, making the architectures unable to focus on the regions of interest (RoIs), thus making the quantitative results unreliable. Therefore, we proposed a pathway named GLobal Bank (GLB) to guide the encoder and the decoder to extract more features of RoIs rather than the complex background. Sufficient experiments have proved that the architecture remoulded by GLB can achieve significant performance improvement, and the quantitative results are more accurate. Hansheng Li, Yuxin Kang, Chunbao Wang 0002, Feihong Liu, Wenli Hui, Qirong Bo, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2022 | Dynamic Key-Value Memory Networks With Rich Features for Knowledge TracingabstractKnowledge tracing is an important research topic in student modeling. The aim is to model a student's knowledge state by mining a large number of exercise records. The dynamic key-value memory network (DKVMN) proposed for processing knowledge tracing tasks is considered to be superior to other methods. However, through our research, we have noticed that the DKVMN model ignores both the students' behavior features collected by the intelligent tutoring system (ITS) and their learning abilities, which, together, can be used to help model a student's knowledge state. We believe that a student's learning ability always changes over time. Therefore, this article proposes a new exercise record representation method, which integrates the features of students' behavior with those of the learning ability, thereby improving the performance of knowledge tracing. Our experiments show that the proposed method can improve the prediction results of DKVMN. Richard F. E. Sutcliffe, Jun Feng 0003 |
IEEE Trans. Cybern. | 6 |
| 2022 | How Many Vehicles Do We Need? Fleet Sizing for Shared Autonomous Vehicles With RidesharingabstractFleet sizing is critical for shared autonomous vehicle(SAV) fleet management to reduce maintenance costs and traffic congestions. As the passenger demands are dynamic, how to make the demand-aware dynamic ridesharing and calculate the fleet size is an important and challenging problem. In this paper, we propose a minimum fleet sizing method called Fleet Sizing for demand-aware Dynamic Ridesharing (FSDR) to accurately determine the fleet size for ridesharing enabled SAV system. Specifically, the travel demands are first predicted by an ensemble method that takes account of temporal correlations between regions. Then a concept called demand utility is proposed to measure the travel demands when planning vehicle paths, and the ride-matching dispatches vehicles to high travel demand regions by maximizing the demand utility along the vehicle path. Based on the ride-matching result, the minimum fleet size is calculated based on a trip graph by the Hopcroft-Karp algorithm. FSDR is evaluated on two real GPS taxi trajectory datasets from Wuhan, China, and San Francisco, USA. The result validates that the proposed demand-aware ridesharing in FSDR can significantly reduce the fleet size compared to the existing ride-matching methods. Moreover, the result shows that by allowing ridesharing, the vehicle fleet size can be reduced by 30%. Boting Qu, Linran Mao, Zhenzhou Xu, Jun Feng 0003, Xin Wang 0004 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | A Dynamic Ridesplitting Method With Potential Pick-Up Probability Based on GPS TrajectoriesabstractRidesplitting is a convenient and budget-friendly for-hire transportation service to arrange one-time shared rides on-the-fly. One crucial component for a ridesplitting system is the effective and efficient rider allocation method to match drivers to riders. Due to the uncertainty of ride requests, the difficulty in locating new riders is one of the problems in rider allocations. In this paper, a dynamic ridesplitting method based on the potential pick-up probability named DRPP is proposed. Given drivers and riders, DRPP aims to allocate the riders to maximize the drivers’ potential pick-up probability, subject to the riders’ time constraints and drivers’ capacity constraint. In DRPP, a grid network is first constructed to predict each grid’s pick-up probability and the traveling time between grids from historical GPS trajectories. To allocate multiple riders, an iterated local search method called ILSAS is proposed to find the solution with overall maximized potential pick-up probability for the drivers. Moreover, we propose the data structure TKdS-tree to improve the rider allocation efficiency. DRPP is evaluated on two real trajectory datasets. The experiment shows that DRPP performed better than other methods in service rate, share rate, and rider waiting time. Boting Qu, Xinyu Ren, Jun Feng 0003, Xin Wang 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Boosting Boundary Representation for Gland Instance SegmentationabstractAccurate and automated gland instance segmentation on histology images can assist pathologists to analyze the malignancy degree of adenocarcinoma. Recently, deep-learning-based segmentation networks have been significantly developed to achieve this goal. However, the gland instances are generally proximate to each other and have indiscernible boundaries (i.e., homogeneous intensity values). Most of the existed networks do not define discriminative boundaries representation as context information, resulting in segmenting proximate instances incorrectly. In this paper, to improve the segmentation accuracy between proximate instances, we propose a Boundary Definition Module to boost boundaries feature representation by the guidance of the intra-and-extra glandular features. Moreover, we propose to use the Gumbel-Softmax distribution estimator to clarify the final prediction of boundaries further. Finally, we embed the Boundary Definition Module and Gumbel-Softmax distribution estimator into the gland instance network(FullNet) for performance verification. Experiments on the 2015 MICCAI Gland Segmentation Challenge dataset demonstrate that our proposed method achieves state-of-the-art performance. Yuxin Kang, Hansheng Li, Zhuoyue Wu, Feihong Liu, Dongqing Hu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 7 |
| 2021 | Robust Pathological Detector Training Method on Sparsely Annotated Datasets via Spatial CuesabstractComputer-aided diagnosis of pathological images usually requires detection and examination of all positive cells and lesions to make an accurate diagnosis. Therefore, there is an unprecedented demand for effective and reliable methods of training pathological detectors than ever. To train a reliable detector, the training dataset is required to fully annotate all positive instances, such a requirement is challenge and laborious, and is not guaranteed in most cases. However, sparse annotations will limit the training performance of detectors. Here, we propose a novel module named Collaborative Correction Sibling (CCS), which is embedded into the original object detection network to enhance the training performance on sparse annotations in a pioneering way. Specifically, instance-level annotations in the image space can be calibrated by positive instances’ spatial features provided by CCS. Extensive experiments have been conducted on both cellular-and-lesion-level detection tasks, compared with the state of the art methods, our CCS demonstrates the training effectiveness on pathological images. Hansheng Li, Yuxin Kang, Lingyu Hu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 6 |
| 2021 | Fashion Landmark Detection via Deep Residual Spatial Attention NetworkabstractFashion landmark detection is challenging due to the large spatial variances of the landmarks and the scale variations of the clothing images. Therefore, the fashion landmark detection model requires two abilities for accurately locating the coordinates of the landmarks. One is to autonomously focus on the task-related features of clothing images to adapt to the diverse spatial distribution of the fashion landmarks. The other is to extract features that contain multi-scale context information to deal with the scale variations. For these purposes, first, we propose the Direction-aware Spatial Attention Module(DASM), which embeds the direction-aware information into spatial attention to capture global contexts and help the network enhance features. And we integrate our DSAM with a bottleneck to building a Spatial Attention ResBlock(SARB). Based on the SARB, we establish a residual-style network for the feature extraction. Then, two feature fusion operations are performed for encoding multi-scale contexts, which are the feature refinement in a top-to-down way and the multi-scale feature aggregation. We name the proposed model the Deep Residual Spatial Attention Network. We demonstrate the effectiveness of our proposed method by the experimental results on two benchmark datasets, which show the proposed fashion landmark detection network outperforms the state-of-the-art methods. Jun Feng 0003, Qirong Bu |
ICTAI | 2 |
| 2021 | EEG-Based Emotion Recognition Fusing Spacial-Frequency Domain Features and Data-Driven Spectrogram-Like Features
Chen Wang 0093, Jingzhao Hu, Qiaomei Jia, Jiayue Chen, Kun Yang 0001, Jun Feng 0003 |
ISBRA | 7 |
| 2021 | Semantic-enhanced sequential modeling for personality trait recognition from texts
Xia Xue, Jun Feng 0003 |
Appl. Intell. | 2 |
| 2021 | GRU-based capsule network with an improved loss for personnel performance prediction
Xia Xue, Yi Gao 0003, Jun Feng 0003 |
Appl. Intell. | 6 |
| 2021 | Salient object detection via light-weight multi-path cascaded networks
Qirong Bu, Jun Feng 0003 |
Neurocomputing | 5 |
| 2021 | ScalingNet: Extracting features from raw EEG data for emotion recognitionabstractConvolutional Neural Networks (CNNs) have achieved remarkable performance breakthroughs in a variety of tasks. Recently, CNN-based methods that are fed with hand-extracted EEG features have steadily improved their performance on the emotion recognition task. In this paper, we propose a novel convolutional layer, called the Scaling Layer, which can adaptively extract effective data-driven spectrogram-like features from raw EEG signals. Furthermore, it exploits convolutional kernels scaled from one data-driven pattern to exposed a frequency-like dimension to address the shortcomings of prior methods requiring hand-extracted features or their approximations. ScalingNet, the proposed neural network architecture based on the Scaling Layer, has achieved state-of-the-art results across the established DEAP and AMIGOS benchmark datasets. Jingzhao Hu, Chen Wang 0093, Qiaomei Jia, Qirong Bu, Richard F. E. Sutcliffe, Jun Feng 0003 |
Neurocomputing | 6 |
| 2021 | Reweighted Discriminative Optimization for least-squares problems with point cloud registrationabstractOptimization plays a pivotal role in computer graphics and vision. Learning-based optimization algorithms have emerged as a powerful optimization technique for solving problems with robustness and accuracy because it learns gradients from data without calculating the Jacobian and Hessian matrices. The key aspect of the algorithms is the least-squares method, which formulates a general parametrized model of unconstrained optimizations and makes a residual vector approach to zeros to approximate a solution. The method may suffer from undesirable local optima for many applications, especially for point cloud registration, where each element of transformation vectors has a different impact on registration. In this paper, Reweighted Discriminative Optimization (RDO) method is proposed. By assigning different weights to components of the parameter vector, RDO explores the impact of each component and the asymmetrical contributions of the components on fitting results. The weights of parameter vectors are adjusted according to the characteristics of the mean square error of fitting results over the parameter vector space at per iteration. Theoretical analysis for the convergence of RDO is provided, and the benefits of RDO are demonstrated with tasks of 3D point cloud registrations and multi-views stitching. The experimental results show that RDO outperforms state-of-the-art registration methods in terms of accuracy and robustness to perturbations and achieves further improvement than non-weighting learning-based optimization. Yan Zhao 0042, Wen Tang 0004, Jun Feng 0003, Tao Ruan Wan, Long Xi 0001 |
Neurocomputing | 3 |
| 2021 | Gaussianization of Diffusion MRI Data Using Spatially Adaptive Filtering
Feihong Liu, Jun Feng 0003, Geng Chen 0001, Dinggang Shen, Pew-Thian Yap |
Medical Image Anal. | 2 |
| 2021 | Classification of EEG Signals for Epileptic Seizures Using Feature Dimension Reduction Algorithm based on LPP
Jun Feng 0003, Jingzhao Hu |
Multim. Tools Appl. | 3 |
| 2021 | DCE-MRI interpolation using learned transformations for breast lesions classification
Hongyu Wang 0007, Jun Feng 0003, Xiaoying Pan, Bao-ying Chen |
Multim. Tools Appl. | 3 |
| 2021 | Word Representation Learning Based on Bidirectional GRUs With Drop Loss for Sentiment ClassificationabstractSentiment classification is a fundamental task in many natural language processing applications. Neural networks have achieved great successes on the sentiment classification task in recent years, since recurrent neural networks and long-short-term memory networks have the ability to deal with sequences of different lengths and to capture contextual semantic information. However, the effectiveness of these methods is limited when used to extract contextual information from relatively long texts. Therefore, in our model, we apply bidirectional gated recurrent units to capture contextual information as far as possible when learning word representations, which may effectively reduce the noise compared to other methods. We also propose a novel loss function namely drop loss (DL) which makes the model focus on the hard examples - examples which are easily classified incorrectly - in order to improve the accuracy of the model. We experiment on four commonly used datasets, and the results show that the proposed method has a good performance on four datasets, and needs fewer parameters compared with recent benchmarks, such as CoVe, ULMFiT, embeddings from language models, and bidirectional encoder representations from transformers. Furthermore, we demonstrate that the classification performance of existing shallow network models can be significantly improved by using DL. In particular, the accuracy of the CNN+LSTM model improves 9% on the IMDB-10 dataset. Yi Gao 0003, Richard F. E. Sutcliffe, Shou Xi Guo, Xin Wang 0104, Jun Feng 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2020 | Feature Enhancement And Fusion For Image-Based Particle Matter Estimation With F-MSE LossabstractAir pollution is a major hazard to environment and human health. Particle matter with a diameter less than 2.5 micrometers (PM25) is a very harmful air pollutant that can penetrate deeply into lungs through human respiratory system. In this paper, we propose an efficient and reliable method to estimate PM25concentration using outdoor images. Firstly, a prior attention block based on gradient features is used to enhance the boundary area between the sky region and the object in a feature map. After that, an embedding layer is applied to encode weather information and fuse it with image features. Finally, a deep neural network model with a novel loss function, F-MSE, is constructed to combine the prediction error of each model level during the training process and to further improve the effectiveness of the presented method. The proposed method was evaluated on a PM2.5dataset with 1,514 images and the experimental results demonstrate that our method outperformed other existing methods. Qirong Bo, Jun Feng 0003, Jingzhao Hu, Yuxin Kang |
ICIP | 4 |
| 2020 | Personalized tourism route recommendation based on user's active interests*abstractTourism is both an important industry and popular leisure activity undertaken by millions around the world. How to effectively mine the user’s travel mode and visit preferences based on the user’s historical travel data is a challenge. The tourism resources of different visiting areas, such as the popularity of POI (Point of Interest), influence the user’s interests dynamically. Therefore, a user’s interest preferences during traveling may differ between geographical region. In this paper, we introduce a personalized travel route recommendation framework, named PTDR, based on region dependent personal interest. PTDR consists of two parts, which are POI recommendation and itinerary generation. We analyzed the user’s history interest from check-in behavior in detail and constructed a convolutional neural network to extract the potential features of the target visiting area. Then the user’s active interest is learned from the user’s history interest and the potential features of the target area. Finally, we optimize the itinerary for a user based on the orienteering problem, which takes into account the user’s travel restrictions, such as time limits, starting attraction restrictions, and destination attraction restrictions. We evaluated the proposed algorithm on four cities of Flickr datasets and compared them to existing travel recommendations, including accuracy, recall, and F1. Experiments verify the effectiveness of the proposed method. Zhizhou Duan, Yuan Gao 0045, Jun Feng 0003 |
MDM | 3 |
| 2020 | A Novel Loss Calibration Strategy for Object Detection Networks Training on Sparsely Annotated Pathological Datasets
Hansheng Li, Yuxin Kang, Xiaoshuang Shi, Mengdi Yan, Zixu Tong, Qirong Bu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
MICCAI (5) | 9 |
| 2020 | Completely Blind Image Quality Assessment with Visual Saliency Modulated Multi-feature Collaboration
Wenjing Hou, Jun Feng 0003 |
PRCV (1) | 5 |
| 2020 | Blind Image Quality Assessment with Visual Sensitivity Enhanced Dual-Channel Deep Convolutional Neural NetworkabstractRecent years, various blind image quality assessment (BIQA) methods based on deep neural network have been proposed and achieved excellent performance. Most existing deep BIQA methods learn a regression model from distorted images with corresponding human subjective scores with end-to-end neural networks. However, such schemes ignore the characteristics of human visual system (HVS) since human beings are the ultimate receivers of the images. This paper proposed a dual-channel deep neural architecture for BIQA, which incorporated the visual sensitivity with taken the psychophysical characteristics of human visual system (HVS) into consideration. Furthermore, a new loss function is employed, which penalizes the deep network when the order of prediction scores is different from the ground truth order. The experimental results on two benchmark IQA databases show that the proposed method outperforms the state-of-the-arts. Wenjing Hou, Jun Feng 0003 |
QoMEX | 4 |
| 2020 | A deep learning-based framework for lung cancer survival analysis with biomarker interpretationabstractBACKGROUND: Lung cancer is the leading cause of cancer-related deaths in both men and women in the United States, and it has a much lower five-year survival rate than many other cancers. Accurate survival analysis is urgently needed for better disease diagnosis and treatment management. RESULTS: In this work, we propose a survival analysis system that takes advantage of recently emerging deep learning techniques. The proposed system consists of three major components. 1) The first component is an end-to-end cellular feature learning module using a deep neural network with global average pooling. The learned cellular representations encode high-level biologically relevant information without requiring individual cell segmentation, which is aggregated into patient-level feature vectors by using a locality-constrained linear coding (LLC)-based bag of words (BoW) encoding algorithm. 2) The second component is a Cox proportional hazards model with an elastic net penalty for robust feature selection and survival analysis. 3) The third commponent is a biomarker interpretation module that can help localize the image regions that contribute to the survival model's decision. Extensive experiments show that the proposed survival model has excellent predictive power for a public (i.e., The Cancer Genome Atlas) lung cancer dataset in terms of two commonly used metrics: log-rank test (p-value) of the Kaplan-Meier estimate and concordance index (c-index). CONCLUSIONS: In this work, we have proposed a segmentation-free survival analysis system that takes advantage of the recently emerging deep learning framework and well-studied survival analysis methods such as the Cox proportional hazards model. In addition, we provide an approach to visualize the discovered biomarkers, which can serve as concrete evidence supporting the survival model's decision. Lei Cui 0004, Hansheng Li, Wenli Hui, Lin Yang 0002, Yuxin Kang, Qirong Bo, Jun Feng 0003 |
BMC Bioinform. | 8 |
| 2020 | Semantic trajectory segmentation based on change-point detection and ontologyabstractTrajectory segmentation is a fundamental issue in GPS trajectory analytics. The task of dividing a raw trajectory into reasonable sub-trajectories and annotating them based on moving subject’s intentions and application domains remains a challenge. This is due to the highly dynamic nature of individuals’ patterns of movement and the complex relationships between such patterns and surrounding points of interest. In this paper, we present a framework called SEMANTIC-SEG for automatic semantic segmentation of trajectories from GPS readings. For the decomposition component of SEMANTIC-SEG, a moving pattern change detection (MPCD) algorithm is proposed to divide the raw trajectory into segments that are homogeneous in their movement conditions. A generic ontology and a spatiotemporal probability model for segmentation are then introduced to implement a bottom-up ontology-based reasoning for semantic enrichment. The experimental results on three real-world datasets show that MPCD can more effectively identify the semantically significant change-points in a pattern of movement than four existing baseline methods. Moreover, experiments are conducted to demonstrate how the proposed SEMANTIC-SEG framework can be applied. Yuan Gao 0045, Longfei Huang, Jun Feng 0003, Xin Wang 0004 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2020 | Multi-depth dilated network for fashion landmark detection with batch-level online hard keypoint mining
Qirong Bu, Jun Feng 0003 |
Image Vis. Comput. | 4 |
| 2019 | Global Bank: A Guided Pathway of Encoding and Decoding for Pathological Image AnalysisabstractThe encoder-decoder architecture of convolutional neural networks (CNNs) is widely used in computer vision tasks and various analyses of medical images. However, extracting semantic features from regions of interest (RoIs) in pathological images remains a challenging task because RoIs of different morphologies and scales are embedded in a blurred background. Additionally, it is well known that the classic encoder-decoder architecture is vulnerable to interference from a blurred background and is thus not entirely suitable for precise analysis of pathological images. In this paper, we propose a pathway named global bank (GLB) to guide the encoder and decoder to focus more on the RoIs by providing the decoder with additional effective features of the RoIs. We extend the U-Net and feature pyramid network (FPN) with GLB and evaluate the resulting models on gland segmentation and cancer embolus detection tasks, respectively. Extensive experiments demonstrate that our proposal can significantly improve the performance of the encoder-decoder architecture. The U-Net with GLB achieves the best semantic segmentation performance on the 2015 MICCAI Gland Challenge dataset. Additionally, the FPN with GLB achieves improvements of 2% in average precision and 3.4% in recall on the embolus detection task. Hansheng Li, Jun Feng 0003, Baosheng Kang, Yuxin Kang, Feihong Liu, Wenli Hui, Qirong Bo, Chunbao Wang 0002, Lin Yang 0002, Lei Cui 0004 |
BIBM | 2 |
| 2019 | $S^{3}$ Net: Trained on a Small Sample Segmentation Network for Biomedical Image AnalysisabstractFully convolutional networks (FCNs) are powerful methods to extract hierarchies of features that have achieved remarkable success in various biomedical image analysis tasks. However, the successful training of FCN requires more than hundreds of pixel-level annotated training samples, which poses a challenge for biomedical image processing tasks. In this paper, we present S3Net, a network that makes more efficient use of available annotated samples on biomedical image segmentation. S3Net is essentially a deeply-supervised encoder-decoder network where the decoder has been redesigned to efficiently restore multi-level encoded feature resolution in a single step. We have conducted extensive experiments on the 2015 MICCAI Gland Challenge dataset. Compared with other methods, S3Net achieves a 0.781 dice-score using 11 images for training, higher than U-Net++ 18.6% and U-Net 12.9%, which verifies the performance of S3Net trained on a small sample. Further, with training on 85 images, S3Net achieves a 0.910 dice-score by using resnet50 as the encoder, which is higher than state-of-the-art semantic segmentation results by 4%. Mengdi Yan, Hansheng Li, Baosheng Kang, Jun Feng 0003, Yuxin Kang, Lin Yang 0002, Lei Cui 0004 |
BIBM | 4 |
| 2019 | Salient Object Detection via Light-Weight Multi-path Refinement Networks
Jun Feng 0003, Qirong Bu |
PRCV (1) | 2 |
| 2019 | High throughput automatic muscle image segmentation using parallel frameworkabstractBACKGROUND: Fast and accurate automatic segmentation of skeletal muscle cell image is crucial for the diagnosis of muscle related diseases, which extremely reduces the labor-intensive manual annotation. Recently, several methods have been presented for automatic muscle cell segmentation. However, most methods exhibit high model complexity and time cost, and they are not adaptive to large-scale images such as whole-slide scanned specimens. METHODS: In this paper, we propose a novel distributed computing approach, which adopts both data and model parallel, for fast muscle cell segmentation. With a master-worker parallelism manner, the image data in the master is distributed onto multiple workers based on the Spark cloud computing platform. On each worker node, we first detect cell contours using a structured random forest (SRF) contour detector with fast parallel prediction and generate region candidates using a superpixel technique. Next, we propose a novel hierarchical tree based region selection algorithm for cell segmentation based on the conditional random field (CRF) algorithm. We divide the region selection algorithm into multiple sub-problems, which can be further parallelized using multi-core programming. RESULTS: We test the performance of the proposed method on a large-scale haematoxylin and eosin (H &E) stained skeletal muscle image dataset. Compared with the standalone implementation, the proposed method achieves more than 10 times speed improvement on very large-scale muscle images containing hundreds to thousands of cells. Meanwhile, our proposed method produces high-quality segmentation results compared with several state-of-the-art methods. CONCLUSIONS: This paper presents a parallel muscle image segmentation method with both data and model parallelism on multiple machines. The parallel strategy exhibits high compatibility to our muscle segmentation framework. The proposed method achieves high-throughput effective cell segmentation on large-scale muscle images. Lei Cui 0004, Jun Feng 0003, Zizhao Zhang 0002, Lin Yang 0002 |
BMC Bioinform. | 2 |
| 2018 | Semi-supervised Deep Linear Discriminant Analysis for Histopathology Image Classification
Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 2 |
| 2018 | Skull Gender Identification Based on Skull Contour and Convolutional Neural Network
Xiaoning Liu 0001, FangFang Qiao, Jun Feng 0003, Shanghao Zhao, Wen Yang 0003 |
BIBM | 4 |
| 2018 | Deep Convolution Neural Networks for Drug-Drug Interaction Extraction
Jun Feng 0003, Xiaodong Du |
BIBM | 2 |
| 2018 | Particle Pollution Estimation from Images Using Convolutional Neural Network and Weather FeaturesabstractAirborne particulate matter with a diameter less than 2.5 micrometers (PM2.5) is one of the most harmful air pollutants, because PM2.5 can be inhaled into human body and cause serious health problems by transmitting hazardous chemicals deeply into lung and bloodstream. A reliable, easily accessible, and low-cost PM2.5 monitoring system can greatly help people raise public awareness of PM2.5 and reduce health hazards of air pollution. In this paper, we combine image and weather information to estimate PM2.5 indices of outdoor images using deep learning and support vector regression (SVR) techniques. The proposed method first uses a convolutional neural network (CNN) to predict the PM2.5 index based on image information, and then the PM2.5 predicted by CNN and two weather features, humidity and wind speed, are combined to yield final estimated PM2.5 index using a created SVR model. We assessed our method using two datasets collected from Shanghai City and Beijing City in China and experimental results demonstrated the effectiveness of the proposed method for PM2.5 estimation. Qirong Bo, Nabin Rijal, Yilin Xie, Jun Feng 0003 |
ICIP | 5 |
| 2018 | A Shallow ResNet with Layer Enhancement for Image-Based Particle Pollution Estimation
Jun Feng 0003, Qirong Bo |
PRCV (2) | 2 |
| 2018 | Breast mass classification via deeply integrating the contextual information from multi-view data
Hongyu Wang 0007, Jun Feng 0003, Zizhao Zhang 0002, Hai Su, Lei Cui 0004 |
Pattern Recognit. | 2 |
| 2017 | Normalized Euclidean Super-Pixels for Medical Image Segmentation
Feihong Liu, Jun Feng 0003, Wenhuo Su, Zhaohui Lv, Fang Xiao |
ICIC (3) | 2 |
| 2017 | Classifying biomedical knowledge in PubMed using multi-label vector machines with weaker optimization constraints
Jun Feng 0003, Su-Shing Chen, Feijuan He |
Neural Comput. Appl. | 3 |
| 2016 | Image dehazing base on two-peak channel priorabstractHaze is one of the major factors that degrade outdoor images. Removing haze from an image is a challenge problem. In this paper, a two-peak channel prior model is proposed for general image dehazing. Firstly, the estimation of medium transmission function is derived and analyzed comprehensively. Secondly, a new calculation method estimating atmospheric light is proposed for more robust dehazing with a new compensation parameter. The experimental results illustrate that the proposed method is able to achieve more satisfied dehazing results than two state-of-the-art methods. Xiaoxu Han, Hongwei Feng, Qirong Bu, Jun Feng 0003, Xiaoning Liu 0001 |
ICIP | 4 |
| 2014 | Classifying Lung Cancer Knowledge in PubMed According to GO Terms Using Extreme Learning MachineabstractFor a well-established digital library (e.g., PubMed), searching in terms of a newly established ontology (e.g., Gene Ontology (GO)) is an extremely difficult task. Making such a digital library adaptive to any new ontology or to reorganize knowledge automatically is our main objective. The decomposition of the knowledge base into classes is a first step toward our main objective. In this paper, we will demonstrate an automated linking scheme for PubMed citations with GO terms using an improved version of extreme learning machine (ELM) type algorithms. ELM is an emergent technology, which has shown excellent performance in large data classification problems, with fast learning speeds. Xuebin Xu, Jun Feng 0003, Su-Shing Chen |
Int. J. Intell. Syst. | 4 |
| 2012 | Using surface variability characteristics for segmentation of deformable 3D objects with application to piecewise statistical deformable model
Horace Ho-Shing Ip, Bei Hua, Jun Feng 0003 |
Vis. Comput. | 4 |
| 2009 | Segmenting deformable soft-body meshes based on statistical variation information for piecewise Active Shape ModelabstractThis paper proposes an algorithm for segmenting deforming soft-body meshes based on statistical variation information extracted from the deforming meshes. The variation information is extracted by performing a global principal component analysis (PCA) on the set of meshes. eigen-variation similarity (EVS) and eigen-variation magnitude (EVM) are then defined for the vertices and triangle faces of the meshes based on the extracted variation information. A multiple-source region growing algorithm is presented for segmenting a mesh that favors grouping faces with similar variations into a same component. We apply the proposed mesh segmentation algorithm to the construction of piecewise active shape model (ASM) and use such piecewise ASM to reconstruct unseen meshes. Experimental results show that our algorithm outperforms several state-of-the-art methods in terms of reconstruction accuracy. Horace Ho-Shing Ip, Jun Feng 0003, Bei Hua |
CAD/Graphics | 3 |
| 2009 | A multi-resolution statistical deformable model (MISTO) for soft-tissue organ reconstruction
Jun Feng 0003, Horace Ho-Shing Ip |
Pattern Recognit. | 1 |
| 2009 | Mr-SDM: a novel statistical deformable model for object deformation
Qizhen He, Horace Ho-Shing Ip, Jun Feng 0003, Xianbin Cao 0001 |
Vis. Comput. | 3 |
| 2008 | Clustered Microcalcification detection based on a Multiple Kernel Support Vector Machine with Grouped Features (GF-SVM)abstractClustered microcalcification is an important signal for breast cancer in the early stages. In this paper, we propose a multiple kernel SVM with group features (GF-SVM) to tackle problems associated with heterogeneous features of clustered microcalcification and normal breast tissues in suspicious regions. Specifically, different types of features such as being gradient, geometric and textural are grouped and modeled by different kernels, respectively. The prior knowledge from different resources is then combined into the framework of the multiple kernel SVM based classification scheme. Experimental results demonstrate that our classification scheme reduces the false positive rate significantly while maintaining the true positive rate. Tian-Tian Chang, Jun Feng 0003, Horace Ho-Shing Ip |
ICPR | 2 |
| 2008 | Robust point correspondence matching and similarity measuring for 3D models by relative angle-context distributions
Jun Feng 0003, Horace Ho-Shing Ip, Lap Yi Lai, Alf D. Linney |
Image Vis. Comput. | 1 |
| 2006 | MISTO: A Multi-Resolution Deformable Model for Segmentation of Soft-Tissue OrgansabstractWe propose a multi-resolution integrated model for the segmentation of soft-tissue organs called MISTO. The model is constructed hierarchically to represent the most significant deformations from the training set as well as to generate representative deformation modes of the organ shapes. The clutter surrounding of the surface points are formulated in terms of an external functional which is also learnt automatically from the training samples. By combining a set of powerful shape models and context constraints, the segmentation process can be carried out very effectively. To avoid the local minimum during model optimization, the deformation strategies are designed such that the portions of the surface for which we have more reliable prior knowledge on their possible deformations are deformed first, followed by deformation on the less informed portions. The experimental and validation results verify that our proposed approaches can be robustly applied to highly deformable anatomies such as soft-tissue organs. Jun Feng 0003, Horace Ho-Shing Ip |
ICIP | 1 |
| 2005 | Iterative 3D Point-Set Registration Based on Hierarchical Vertex Signature (HVS)
Jun Feng 0003, Horace Ho-Shing Ip |
MICCAI (2) | 1 |
| 2004 | A 3D Geometric Deformable Model for Tubular Structure SegmentationabstractIn this paper, we present a relational-tubular (ReTu) deformable model for segmenting a complex and the entire tubular network structure with branches in close proximity of each other. Specifically, we incorporate a priori knowledge of the target anatomy structure as well as the spatial relationship between branches to reduce possible segmentation errors due to the effects of a variety of imaging artifacts and noise. To get more robust description of the data properties than a simple 3D edge map, a new data energy functional is proposed based on testing the volumetric density within the model cross-sections. The deformation process is formulated as a two-stage procedure: tubular medial axis deformation and tubular surface deformation. The efficiency of this approach is demonstrated by our experiments which show that satisfactory quantifications of the entire zebrafish vasculature recorded from the fluorescence confocal microscope. The experiments also demonstrate the robustness of our deformable model in the presence of complex issue structure that adhered to the vessel branches. Jun Feng 0003, Horace Ho-Shing Ip, Shuk Han Cheng |
MMM | 1 |
| 2001 | Affine-Invariant Sketch-Based Retrieval of ImagesabstractThe advent of the digital library and multimedia database require robust techniques for multimedia content searches. Content-based retrieval techniques have been developed to overcome some of the limitations associated with conventional keyword-based browsing or searching of visual data. We present an affine invariant shape-based retrieval technique for image retrieval which is efficient and robust. More importantly, the technique supports a query presented in the form of hand-drawn sketches and has the potential of supporting affine invariant partial shape retrieval. Horace Ho-Shing Ip, Angus K. Y. Cheng, William Y. F. Wong, Jun Feng 0003 |
Computer Graphics International | 4 |