EDBT 2026 Demo / reviewers in the wild / expert
Ming Feng
dblp:69/4238
· DBLP profile ↗
25ranked-venue papers
6as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery
Lumin Chen, Zhiying Wu, Tianye Lei, Xuexue Bai, Ming Feng, Yuxi Wang 0001, Gaofeng Meng, Zhen Lei 0001, Hongbin Liu 0001 |
MICCAI (9) | 5 |
| 2025 | Reconstructing 3D Hand-Instrument Interaction from a Single 2D Image in Medical Scenes
Xiangyu Zhu 0001, Jinlin Wu, Ming Feng, Zelin Zang, Hongbin Liu 0001, Zhen Lei 0001 |
MICCAI (10) | 4 |
| 2025 | Self-Supervised Multi-Scale Multi-Modal Graph Pool Transformer for Sellar Region Tumor DiagnosisabstractThe sellar region tumor is a brain tumor that only exists in the brain sellar, which affects the central nervous system. The early diagnosis of the sellar region tumor subtypes helps clinicians better understand the best treatment and recovery of patients. Magnetic resonance imaging (MRI) has proven to be an effective tool for the early detection of sellar region tumors. However, the existing sellar region tumor diagnosis still remains challenging due to the small amount of dataset and data imbalance. To overcome these challenges, we propose a novel self-supervised multi-scale multi-modal graph pool Transformer (MMGPT) network that can enhance the multi-modal fusion of small and imbalanced MRI data of sellar region tumors. MMGPT can strengthen feature interaction between multi-modal images, which makes our model more robust. A contrastive learning equipped auto-encoder (CAE) via self-supervised learning (SSL) is adopted to learn more detailed information between different samples. The proposed CAE transfers the pre-trained knowledge to the downstream tasks. Finally, a hybrid loss is equipped to relieve the performance degradation caused by data imbalance. The experimental results show that the proposed method outperforms state-of-the-art methods and obtains higher accuracy and AUC in the classification of sellar region tumors. Bai Ying Lei, Gege Cai, Yun Zhu 0006, Tianfu Wang 0001, Cheng Zhao 0003, Xinzhi Hu, Huijun Zhu, Ming Feng, Renzhi Wang 0002 |
IEEE J. Biomed. Health Informatics | 11 |
| 2024 | Enhancing Cloud-Native Security Through eBPF TechnologyabstractIn the cloud industry, eBPF (extended Berkeley Packet Filter) security technology is one of the most popular and influential technologies in the Linux kernel in recent years. With the rapid development of networking and cloud-native technologies, eBPF is extensively applied in network and security, performance analysis, container and cloud-native environments, operations and troubleshooting, as well as observability of application systems. This paper focuses on the practical application of eBPF technology in cloud environments for ensuring the secure operation, event-driven security handling capability, and performance hotspots observation of transaction systems. The integration of eBPF technology significantly enhances efficiency and reduces costs across the development, testing, and operational phases of systems, providing effective means for security optimization, assisting in testing, and facilitating fault localization and troubleshooting during secure production operations. Moreover, eBPF plays a crucial role in handling security events in cloud environments. Ming Feng |
CSCloud | 1 |
| 2024 | Co-Axial Slender Tubular robot (CAST): Towards Robotized Operation for Transorbital Neurosurgery with Minimal InvasivenessabstractTransorbital Neuro Surgery (TNS) offers a novel treatment towards the lesion inside skull pursuing minimal invasiveness. Most conventional TNS tools are rigid and straight, limiting the dexterity and accessibility in passing a small port. Bendable and steerable surgical tools provides an alternative for this issue. In this work, we proposed a dual-segment slender surgical robot arm for TNS, which is a Co-Axial Slender Tubular robot (CAST), and modelled it using novel approaches. Another contribution is tendon-mortise shaped slits along the axial direction, enhancing the overall stiffness. The bending of CAST is actuated by pushing/pulling distance, and the maximum diameter is only 1.7mm with high dexterity after mounting on a rigid robot arm. Experiments demonstrates that the proposed the slit design doubles the stiffness properties compared to traditional rectangle slit designs. The path-following task shows that the position error was maximally 3mm in open-looped control. Test on a skull model demonstrates that the whole system could successfully perform electrocoagulation procedure inside the depth of skull in a robotized manner effectively. Shuai Wang 0024, Qingxiang Zhao, Jian Chen 0036, Mingcong Chen, Guanglin Cao, Runfeng Zhu, Danny Tat-Ming Chan, Ming Feng, Hongbin Liu 0001 |
ICRA | 9 |
| 2024 | BioPRO: Context-Infused Prompt Learning for Biomedical Entity LinkingabstractRecent research tends to address the biomedical entity linking problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing thevarietychallenge of the heterogeneous naming of biomedical concepts. Yet, theambiguitychallenge that the same word under different contexts can be used to refer to distinct concepts is usually ignored. To address this challenge, we propose BioPRO, a two-stage entity linking algorithm to enhance the biomedical entity representations based on context-infused prompt learning. The first stage includes a coarse-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity's surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a fine-grained encoder based on prompt-tuning that sufficiently stimulates knowledge in contextual information of mentions and entities. Furthermore, the trained fine-grained encoder can be utilized to generate deep representations of bio-entities and boost candidate retrieval in the first stage. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art (SOTA) techniques on 4 biomedical corpora. We also observe by cases that the proposed context-infused prompt-tuning strategy is effective in solving both thevarietyandambiguitychallenges in the linking task. Tiantian Zhu 0002, Yang Qin 0001, Ming Feng, Qingcai Chen, Baotian Hu, Yang Xiang 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Raw Ultrasound-Based Phonetic Segments Classification Via Mask ModelingabstractUltrasound tongue imaging is widely used in clinical linguistics and phonetics. Recently, deep neural networks, especially convolutional neural networks, have been widely used in the interpretation and analysis of ultrasound tongue images (UTI). Despite achieving satisfactory performance, deep models rely on a large amount of manually labeled data, which is often difficult to obtain in practical settings. To address this issue, this paper focuses on how to utilize a large amount of unlabeled UTI data to improve the performance of UTI classification task. Specifically, we explore self-supervised learning with masking modeling strategy. By predicting the masked part, our pre-trained model enables the neural network to infer contextual information. Then, we fine-tune the pre-trained model with a small amount of labeled data. Compared with the previous competing algorithms, our method can improve the classification accuracy by an average of 13.33% in four different scenarios. Kang You, Bo Liu 0014, Kele Xu, Yunsheng Xiong, Qisheng Xu, Ming Feng, Tamás Gábor Csapó, Boqing Zhu |
ICASSP | 6 |
| 2023 | IFS-SED: Incremental Few-Shot Sound Event Detection Using Explicit Learning and CalibrationabstractSound event detection (SED) refers to recognizing the sound events in a continuous audio signal, which has drawn increasing interest during recent decades. The applications of SED seem to be evident in many fields, ranging from surveillance to monitoring applications. Despite the sustainable efforts that have been made, most of the previous attempts are performed on the closed-set, as only fixed and known sound event classes can be employed during the training. In this paper, we present our incremental few-shot SED framework under the open-set settings, as a practical machine listening system should be able to address unknown sound events. Specifically, an explicit learning and calibration-based multi-stage learning framework is utilized to address the challenges of catastrophic forgetting, and aim to achieve a better trade-off between stability and plasticity. To compress the model efficiently, the model prune and self-distillation paradigm are combined used for the model compression, thus our system can be deployed for the resource-limited devices. Our framework can also provide an uncertain estimation for the inference. Lastly, an interactive interface is presented to demonstrate the functions of our system. Ming Feng, Kele Xu, Hengxing Cai |
ACM Multimedia | 1 |
| 2023 | 3D Shuffle-Mixer: An Efficient Context-Aware Vision Learner of Transformer-MLP Paradigm for Dense Prediction in Medical VolumeabstractDense prediction in medical volume provides enriched guidance for clinical analysis. CNN backbones have met bottleneck due to lack of long-range dependencies and global context modeling power. Recent works proposed to combine vision transformer with CNN, due to its strong global capture ability and learning capability. However, most works are limited to simply applying pure transformer with several fatal flaws (i.e., lack of inductive bias, heavy computation and little consideration for 3D data). Therefore, designing an elegant and efficient vision transformer learner for dense prediction in medical volume is promising and challenging. In this paper, we propose a novel 3D Shuffle-Mixer network of a new Local Vision Transformer-MLP paradigm for medical dense prediction. In our network, a local vision transformer block is utilized to shuffle and learn spatial context from full-view slices of rearranged volume, a residual axial-MLP is designed to mix and capture remaining volume context in a slice-aware manner, and a MLP view aggregator is employed to project the learned full-view rich context to the volume feature in a view-aware manner. Moreover, an Adaptive Scaled Enhanced Shortcut is proposed for local vision transformer to enhance feature along spatial and channel dimensions adaptively, and a CrossMerge is proposed to skip-connect the multi-scale feature appropriately in the pyramid architecture. Extensive experiments demonstrate the proposed model outperforms other state-of-the-art medical dense prediction methods. Jianye Pang, Cheng Jiang 0001, Jianbo Chang, Ming Feng, Renzhi Wang 0002, Jianhua Yao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2022 | EasySED: Trusted Sound Event Detection with Self-Distillation
Qingsong Zhou, Kele Xu, Ming Feng |
AAAI | 3 |
| 2022 | Ideal Midsagittal Plane Detection Using Deep Hough Plane Network for Brain Surgical Planning
Chenchen Qin, Wenxue Zhou, Jianbo Chang, Dasheng Wu, Yixun Liu, Ming Feng, Renzhi Wang 0002, Wenming Yang, Jianhua Yao 0001 |
MICCAI (8) | 7 |
| 2022 | Seeing Speech: Magnetic Resonance Imaging-Based Vocal Tract Deformation Visualization Using Cross-Modal TransformerabstractAs an essential component to advance speech science, understanding of speech production can be greatly helpful to improve our understanding of motor control, dynamical systems of humans during natural speech. Different medical imaging modalities have been leveraged to visualize the dynamic process, in which Magnetic resonance imaging (MRI) provides a valuable tool for evaluating static postures. In this demo, we present our solution to visualize the vocal tract deformation, leveraging the correlation between the MRI and the acoustical signals. We first formulate the problem as a cross-modal prediction task and a novel cross-modal Transformer network is proposed. Thus, we can infer the deformation of the vocal tract by only utilizing the acoustical signals. Then, we present an interactive framework, which can be used to visualize the deformation utilizing the aforementioned network. We hope our solution can also be helpful in pronunciation training for children with sound speech disorders and second language learning. Kele Xu, Ming Feng, Weiquan Huang |
ACM Multimedia | 2 |
| 2022 | Masked Modeling-based Audio Representation for ACM Multimedia 2022 Computational Paralinguistics ChallengEabstractIn this paper, we present our solution for ACM Multimedia 2022 Computational Paralinguistics Challenge. Our method employs the self-supervised learning paradigm, as it achieves promising results in computer vision and audio signal processing. Specifically, we firstly explore modifying the Swin Transformer architecture to learn general representation for the audio signals, accompanied with random masking on the log-mel spectrogram. The main goal of the pretext task is to predict the masked parts, by combining the advantages of the Swin-Transformer and masked modeling. For the downstream tasks, we utilize the labelled datasets to fine-tune the pre-trained model. Compared with the competitive baselines, our approach can provide significant performance improvements without ensembling. Kang You, Kele Xu, Boqing Zhu, Ming Feng, Bo Liu 0014, Bo Ding 0001 |
ACM Multimedia | 4 |
| 2022 | A bilevel whale optimization algorithm for risk management scheduling of information technology projects considering outsourcing
Fuqiang Lu, Tongren Yan, Hualing Bi, Ming Feng, Suxin Wang, Min Huang 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Colony search optimization algorithm using global optimization
Heng Wen, Suxin Wang, Fuqiang Lu, Ming Feng, Lei Zhen Wang, Jun Kai Xiong, Ma Cong Si |
J. Supercomput. | 4 |
| 2022 | Automatic Brain Midline Surface Delineation on 3D CT Images With Intracranial HemorrhageabstractBrain midline delineation plays an important role in guiding intracranial hemorrhage surgery, which still remains a challenging task since hemorrhage shifts the normal brain configuration. Most previous studies detected brain midline on 2D plane and did not handle hemorrhage cases well. We propose a novel and efficient hemisphere-segmentation framework (HSF) for 3D brain midline surface delineation. Specifically, we formulate the brain midline delineation as a 3D hemisphere segmentation task, and employ an edge detector and a smooth regularization loss to generate the midline surface. We also introduce a distance-weighted map to keep the attention on the midline. Furthermore, we adopt rectification learning to handle various head poses. Finally, considering the complex situation of ventricle break-in for hemorrhages in bilateral intraventricular (B-IVH) cases, we identify those cases via a classification model and design a midline correction strategy to locally adjust the midline. To our best knowledge, it is the first study focusing on delineating the brain midline surface on 3D CT images of hemorrhage patients and handling the situation of ventricle break-in. Extensive validation on our large in-house datasets (519 patients) and the public CQ500 dataset (491 patients), demonstrates that our method outperforms state-of-the-art methods on brain midline delineation. Dasheng Wu, Haoming Li 0012, Jianbo Chang, Chenchen Qin, Yixun Liu, Bingsheng Huang, Ming Feng, Renzhi Wang 0002, Jianhua Yao 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Improving Ultrasound Tongue Contour Extraction Using U-Net and Shape Consistency-Based RegularizerabstractB-mode ultrasound tongue imaging is widely used to visualize the tongue motion, due to its appearing properties. Extracting the tongue surface contour in the B-mode ultrasound image is still a challenge, while it is a prerequisite for further quantitative analysis. Recently, deep learning-based approach has been adopted in this task. However, the standard deep models fail to address faint contour when the ultrasound wave goes parallel to the tongue surface. To address the faint or missing contours in the sequence, we explore the shape consistency-based regularizer, which can take sequential information into account. By incorporating the regularizer, the deep model not only can extract frame-specific contours, but also can enforce the similarity between the contours extracted from adjacent frames. Extensive experiments are conducted both on the synthetic and real ultrasound tongue imaging dataset and the results demonstrate the effectiveness of proposed method. To better promote the research in this field, we have released our code at1. Ming Feng, Kele Xu, Huaimin Wang 0001, Bo Ding 0001 |
ICASSP | 1 |
| 2021 | Batch Weighted Nuclear-Norm Minimization for Medical Image Sequence Segmentation
Kele Xu, Zijian Gao, Jilong Wang 0007, Ming Feng |
ISBRA | 5 |
| 2021 | 3D Brain Midline Delineation for Hematoma Patients
Chenchen Qin, Haoming Li 0012, Yixun Liu, Hong Shang, Hanqi Pei, Jianbo Chang, Ming Feng, Renzhi Wang 0002, Jianhua Yao 0001 |
MICCAI (5) | 9 |
| 2021 | MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification ChallengeabstractDetecting various types of cells in and around the tumor matrix holds a special significance in characterizing the tumor micro-environment for cancer prognostication and research. Automating the tasks of detecting, segmenting, and classifying nuclei can free up the pathologists' time for higher value tasks and reduce errors due to fatigue and subjectivity. To encourage the computer vision research community to develop and test algorithms for these tasks, we prepared a large and diverse dataset of nucleus boundary annotations and class labels. The dataset has over 46,000 nuclei from 37 hospitals, 71 patients, four organs, and four nucleus types. We also organized a challenge around this dataset as a satellite event at the International Symposium on Biomedical Imaging (ISBI) in April 2020. The challenge saw a wide participation from across the world, and the top methods were able to match inter-human concordance for the challenge metric. In this paper, we summarize the dataset and the key findings of the challenge, including the commonalities and differences between the methods developed by various participants. We have released the MoNuSAC2020 dataset to the public. Ruchika Verma, Neeraj Kumar 0002, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E Ahmed Raza, Nasir M. Rajpoot, Xiyi Wu, Huai Chen, Lisheng Wang, Hyun Jung, G. Thomas Brown, Shuolin Liu, Seyed Alireza Fatemi Jahromi, Aliasghar Khani, Ehsan Montahaei, Mahdieh Soleymani Baghshah, Hamid Behroozi, Pavel Semkin, Alexandr Rassadin, Prasad Dutande, Romil Lodaya, Ujjwal Baid, Bhakti Baheti, Sanjay N. Talbar, Amirreza Mahbod, Rupert Ecker, Isabella Ellinger, Bin Dong 0006, Zhengyu Xu, Yuehan Yao, Ming Feng, Kele Xu, Hasib Zunair, A. Ben Hamza, Steven M. Smiley, Tang-Kai Yin, Qi-Rui Fang, Shikhar Srivastava 0001, Dwarikanath Mahapatra, Lubomira Trnavska, Hanyun Zhang, Priya Lakshmi Narayanan, Justin Law, Yinyin Yuan, Abhiroop Tejomay, Aditya Mitkari, Dinesh Koka, Vikas Ramachandra, Lata Kini, Amit Sethi |
IEEE Trans. Medical Imaging | 39 |
| 2020 | A Quantitative Comparison of Different Machine Learning Approaches for Human Spermatozoa Quality Prediction Using Multimodal DatasetsabstractDespite remarkable advances in medical data analysis fields, they are severely restrained from the limited property of the employed single modality, usually medical imaging data. However, other modalities (such as patient-related information) should also be taken into account in the process of clinical decision. How to fully employ the multi-modal dataset is still under-explored. In this paper, we make a quantitative comparison of different machine learning approaches for the human spermatozoa quality prediction task, leveraging multiple modalities dataset. To empirically investigate the advantages and disadvantages of different machine learning approaches, we perform extensive experiments. Leveraging different features, we achieve state-of-the-art performance on most of the tasks. The obtained results show that simple models can provide better performance, which emphasizes the importance of avoiding overfitting. For the sake of reproducibility, we have released our code to facilitate the research community. Ming Feng, Kele Xu |
ACM Multimedia | 1 |
| 2020 | Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions for Beauty Product RetrievalabstractThe application of beauty and personal-care product retrieval seems to be evident in our daily life, and it has attracted increasing research interests during the last decade. However, the retrieval task is suffered from different image variations and complicated backgrounds. Recent works have demonstrated that Generalized-attention Regional Maximal Activation of Convolutions (GRMAC) descriptor can provide state-of-the-art performance for the retrieval task. However, GRMAC descriptor is restrained from the essentially limited property of the employed feature from a single layer. Features from a single layer are not robust enough for scale variations, shape deformation, and heavy occlusion. In this paper, we propose a novel descriptors, named Multi-Scale Generalized Attention-Based Regional Maximum Activation of Convolutions (MS-GRMAC). This method introduces multi-scale generalized attention mechanism to reduce the influence of scale variations, thus, can boost the performance of the retrieval task. To empirically investigate the effectiveness of the proposed approach, we conduct extensive experiments on the dataset containing more than half-million personal-care products (Perfect-500K) and obtain satisfactory results without ensemble. Kele Xu, Yuzhong Liu, Ming Feng, Jianqiao Zhao, Huaimin Wang 0001, Hengxing Cai |
ACM Multimedia | 3 |
| 2019 | MSNET-Blockchain: A New Framework for Securing Mobile Satellite Communication NetworkabstractIn this paper, the security problem for mobile satellite communication networks (MSNET) has been investigated. With the rapidly growth of communication needs, mobile satellite systems represent a significant solution to provide high-quality communication services to mobile users in under-populated regions, in emergency areas, on planes, trains and ships. However, lacking an effective framework to secure mobile satellite communication networks seriously limited the practicality of satellite services. Therefore, a new security framework have been developed in this paper to address the security challenges in mobile satellite communication network. Firstly, the mobile satellite communication networks have been formulated as delay-tolerance network (DTN). Then, the blockchain technique has been adopted and used in two aspects, i.e. 1) integrating with DTN structure to secure the data communication, 2) combing with the practical satellite constellation management algorithm to defend the unexpected cyber attacks physically. Through integrating emerging blockchain techniques with both communication and physical aspects, the developed framework cannot only effectively detect the cyber attacks, but also better defend the mobile satellite communication networks through communication and satellite management aspects. Eventually, the numerical simulation and experimental tests have been provided to demonstrate the effectiveness of developed MSNET-Blockchain framework. Ming Feng, Hao Xu 0002 |
SECON | 1 |
| 2012 | Study on digital coded technology in active radar calibrator of SARabstractThe fundamental method of the active coded radiometric calibration technique is analyzed. The theory of reducing the influence of background clutter signal to the precision of radiometric calibration by active coded radiometric calibration technique is explained. Digital coded technique is applied to avoid phase-shifting error of the phase-shifter used in the analog coded method. An experiment is carried out, and the experiment results proved that the active coded radiometric calibration technique can restrain the influence of background clutter signal effectively and improve the precision of radiometric calibration. Hong Jun, Ming Feng |
IGARSS | 3 |
| 2009 | Back to the future: a non-automated method for constructing transfer models
Ming Feng, Joseph E. Beck |
EDM | 1 |