EDBT 2026 Demo / reviewers in the wild / expert
Mingfei Gao
dblp:67/6825
· DBLP profile ↗
29ranked-venue papers
10as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 9 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 7 since 2021Computer networks · 2Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuningabstractWe present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development. Haotian Zhang 0005, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang 0002, Yanghao Li, Sam Dodge, Keen You, Aleksei Timofeev, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You |
ICLR | 2 |
| 2025 | UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and GenerationabstractWe introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen’s image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to future research. Code is available at https://github.com/apple/ml-unigen. Mingfei Gao, Jiasen Lu, Zuxuan Wu, Yinfei Yang, Afshin Dehghan |
NeurIPS | 2 |
| 2024 | 4M-21: An Any-to-Any Vision Model for Tens of Tasks and ModalitiesabstractCurrent multimodal and multitask foundation models, like 4M or UnifiedIO, show promising results. However, their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually small) number of modalities and tasks they are trained on. In this paper, we develop a single any-to-any model trained on tens of highly diverse modalities and by performing co-training on large-scale multimodal datasets and text corpora. This includes training on images and text along with several semantic and geometric modalities, feature maps from recent state of the art models like DINOv2 and ImageBind, pseudo labels of specialist models like SAM and 4DHumans, and a range of new modalities that allow for novel ways to interact with the model and steer the generation, for example, image metadata or color palettes.
A crucial step in this process is performing discrete tokenization on various modalities, whether they are image-like, neural network feature maps, vectors, structured data like instance segmentation or human poses, or data that can be represented as text.
Through this, we show the possibility of training one model to solve at least 3x more tasks/modalities than existing models and doing so without a loss in performance. In addition, this enables more fine-grained and controllable multimodal generation capabilities and allows studying the distillation of models trained on diverse data and objectives into one unified model.
We scale the training to a three billion parameter and different datasets. The multimodal models and training code are open sourced at https://4m.epfl.ch/. Roman Bachmann 0001, Oguzhan Fatih Kar, David Mizrahi, Ali Garjani, Mingfei Gao, David Griffiths, Afshin Dehghan, Amir Zamir |
NeurIPS | 5 |
| 2023 | Mask-Free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask AnnotationsabstractExisting instance segmentation models learn task-specific information using manual mask annotations from base (training) categories. These mask annotations require tremendous human effort, limiting the scalability to annotate novel (new) categories. To alleviate this problem, Open-Vocabulary (OV) methods leverage large-scale image-caption pairs and vision-language models to learn novel categories. In summary, an OV method learns task-specific information using strong supervision from base annotations and novel category information using weak supervision from image-captions pairs. This difference between strong and weak supervision leads to overfitting on base categories, resulting in poor generalization towards novel categories. In this work, we overcome this issue by learning both base and novel categories from pseudo-mask annotations generated by the vision-language model in a weakly supervised manner using our proposed Mask-free OVIS pipeline. Our method automatically generates pseudo-mask annotations by leveraging the localization ability of a pre-trained vision-language model for objects present in image-caption pairs. The generated pseudo-mask annotations are then used to supervise an instance segmentation model, freeing the entire pipeline from any labour-expensive instance-level annotations and overfitting. Our extensive experiments show that our method trained with just pseudo-masks significantly improves the mAP scores on the MS-COCO dataset and OpenImages dataset compared to the recent state-of-the-art methods trained with manual masks. Codes and models are provided in https://vibashan.github.io/ovis-web/. Vibashan VS, Ning Yu 0006, Chen Xing, Can Qin, Mingfei Gao, Juan Carlos Niebles, Vishal M. Patel, Ran Xu 0001 |
CVPR | 5 |
| 2023 | ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D UnderstandingabstractThe recognition capabilities of current state-of-the-art 3D models are limited by datasets with a small number of annotated data and a pre-defined set of categories. In its 2D counterpart, recent advances have shown that similar problems can be significantly alleviated by employing knowledge from other modalities, such as language. Inspired by this, leveraging multimodal information for 3D modality could be promising to improve 3D understanding under the restricted data regime, but this line of research is not well studied. Therefore, we introduce ULIP to learn a unified representation of image, text, and 3D point cloud by pre-training with object triplets from the three modalities. To overcome the shortage of training triplets, ULIP leverages a pre-trained vision-language model that has already learned a common visual and textual space by training with massive image-text pairs. Then, ULIP learns a 3D representation space aligned with the common image-text space, using a small number of automatically synthesized triplets. ULIP is agnostic to 3D backbone networks and can easily be integrated into any 3D architecture. Experiments show that ULIP effectively improves the performance of multiple recent 3D backbones by simply pre-training them on ShapeNet55 using our framework, achieving state-of-the-art performance in both standard 3D classification and zero-shot 3D classification on ModelNet40 and ScanObjectNN. ULIP also improves the performance of PointMLP by around 3% in 3D classification on ScanObjectNN, and outperforms PointCLIP by 28.8% on top-1 accuracy for zero-shot 3D classification on ModelNet40. Our code and pre-trained models will be released. Le Xue, Mingfei Gao, Chen Xing, Roberto Martin Martin, Jiajun Wu 0001, Caiming Xiong, Ran Xu 0001, Juan Carlos Niebles, Silvio Savarese |
CVPR | 2 |
| 2023 | Robustness Evaluation of Transformer-Based Form Field Extractors via Form Attacks
Le Xue, Mingfei Gao, Zeyuan Chen 0001, Caiming Xiong, Ran Xu 0001 |
ICDAR (2) | 2 |
| 2023 | 4M: Massively Multimodal Masked ModelingabstractCurrent machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of capabilities, hinting at a possibility for similarly versatile models in computer vision.
In this paper, we take a step in this direction and propose a multimodal training scheme called 4M. It consists of training a single unified Transformer encoder-decoder using a masked modeling objective across a wide range of input/output modalities – including text, images, geometric, and semantic modalities, as well as neural network feature maps. 4M achieves scalability by unifying the representation space of all modalities through mapping them into discrete tokens and performing multimodal masked modeling on a small randomized subset of tokens.
4M leads to models that exhibit several key capabilities: (1) they can perform a diverse set of vision tasks out of the box, (2) they excel when fine-tuned for unseen downstream tasks or new input modalities, and (3) they can function as a generative model that can be conditioned on arbitrary modalities, enabling a wide variety of expressive multimodal editing capabilities with remarkable flexibility.
Through experimental analyses, we demonstrate the potential of 4M for training versatile and scalable foundation models for vision tasks, setting the stage for further exploration in multimodal learning for vision and other domains. David Mizrahi, Roman Bachmann 0001, Oguzhan Fatih Kar, Teresa Yeo, Mingfei Gao, Afshin Dehghan, Amir Zamir |
NeurIPS | 5 |
| 2022 | TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation
Jun Wang 0090, Mingfei Gao, Yuqian Hu, Ramprasaath R. Selvaraju, Chetan Ramaiah, Ran Xu 0001, Joseph F. JáJá, Larry Davis 0001 |
BMVC | 2 |
| 2022 | DocQueryNet: Value Retrieval with Arbitrary Queries for Form-like DocumentsabstractWe propose, DocQueryNet, a value retrieval method with arbitrary queries for form-like documents to reduce human effort of processing forms. Unlike previous methods that only address a fixed set of field items, our method predicts target value for an arbitrary query based on the understanding of the layout and semantics of a form. To further boost model performance, we propose a simple document language modeling (SimpleDLM) strategy to improve document understanding on large-scale model pre-training. Experimental results show that DocQueryNet outperforms previous designs significantly and the SimpleDLM further improves our performance on value retrieval by around 17% F1 score compared with the state-of-the-art pre-training method. Code is available here, https://github.com/salesforce/QVR-SimpleDLM. Mingfei Gao, Le Xue, Chetan Ramaiah, Chen Xing, Ran Xu 0001, Caiming Xiong |
COLING | 1 |
| 2022 | Open Vocabulary Object Detection with Pseudo Bounding-Box Labels
Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li 0001, Ran Xu 0001, Wenhao Liu 0003, Caiming Xiong |
ECCV (10) | 1 |
| 2022 | Burn After Reading: Online Adaptation for Cross-domain Streaming Data
Luyu Yang, Mingfei Gao, Zeyuan Chen 0001, Ran Xu 0001, Abhinav Shrivastava, Chetan Ramaiah |
ECCV (33) | 2 |
| 2021 | WOAD: Weakly Supervised Online Action Detection in Untrimmed VideosabstractOnline action detection in untrimmed videos aims to identify an action as it happens, which makes it very important for real-time applications. Previous methods rely on tedious annotations of temporal action boundaries for training, which hinders the scalability of online action detection systems. We propose WOAD, a weakly supervised framework that can be trained using only video-class labels. WOAD contains two jointly-trained modules, i.e., temporal proposal generator (TPG) and online action recognizer (OAR). Supervised by video-class labels, TPG works offline and targets at accurately mining pseudo frame-level labels for OAR. With the supervisory signals from TPG, OAR learns to conduct action detection in an online fashion. Experimental results on THUMOS’14, ActivityNet1.2 and ActivityNet1.3 show that our weakly-supervised method largely outperforms weakly-supervised baselines and achieves comparable performance to the previous strongly-supervised methods. Beyond that, WOAD is flexible to leverage strong supervision when it is available. When strongly supervised, our method obtains the state-of-the-art results in the tasks of both online per-frame action recognition and online detection of action start. Mingfei Gao, Yingbo Zhou 0002, Ran Xu 0001, Richard Socher, Caiming Xiong |
CVPR | 1 |
| 2021 | Deep Co-Training with Task Decomposition for Semi-Supervised Domain AdaptationabstractSemi-supervised domain adaptation (SSDA) aims to adapt models trained from a labeled source domain to a different but related target domain, from which unlabeled data and a small set of labeled data are provided. Current methods that treat source and target supervision without distinction overlook their inherent discrepancy, resulting in a source-dominated model that has not effectively use the target supervision. In this paper, we argue that the labeled target data needs to be distinguished for effective SSDA, and propose to explicitly decompose the SSDA task into two sub-tasks: a semi-supervised learning (SSL) task in the target domain and an unsupervised domain adaptation (UDA) task across domains. By doing so, the two sub-tasks can better leverage the corresponding supervision and thus yield very different classifiers. To integrate the strengths of the two classifiers, we apply the well established co-training framework, in which the two classifiers exchange their high confident predictions to iteratively "teach each other" so that both classifiers can excel in the target domain. We call our approach Deep Co-training with Task decomposition (DeCoTa). DeCoTa requires no adversarial training and is easy to implement. Moreover, DeCoTa is well founded on the theoretical condition of when co-training would succeed. As a result, DeCoTa achieves state-of-the-art results on several SSDA datasets, outperforming the prior art by a notable 4% margin on DomainNet. Code is available at https://github.com/LoyoYang/DeCoTa. Luyu Yang, Yan Wang 0051, Mingfei Gao, Abhinav Shrivastava, Kilian Q. Weinberger, Wei-Lun Chao, Ser-Nam Lim |
ICCV | 3 |
| 2020 | Consistency-Based Semi-supervised Active Learning: Towards Minimizing Labeling Cost
Mingfei Gao, Sercan Ö. Arik, Larry Davis 0001, Tomas Pfister |
ECCV (10) | 1 |
| 2020 | InfoFocus: 3D Object Detection for Autonomous Driving with Dynamic Information Modeling
Jun Wang 0090, Shiyi Lan, Mingfei Gao, Larry Davis 0001 |
ECCV (10) | 3 |
| 2019 | WSLLN: Weakly Supervised Natural Language Localization NetworksabstractMingfei Gao, Larry Davis, Richard Socher, Caiming Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingfei Gao, Larry Davis 0001, Richard Socher, Caiming Xiong |
EMNLP/IJCNLP (1) | 1 |
| 2019 | StartNet: Online Detection of Action Start in Untrimmed VideosabstractWe propose StartNet to address Online Detection of Action Start (ODAS) where action starts and their associated categories are detected in untrimmed, streaming videos. Previous methods aim to localize action starts by learning feature representations that can directly separate the start point from its preceding background. It is challenging due to the subtle appearance difference near the action starts and the lack of training data. Instead, StartNet decomposes ODAS into two stages: action classification (using ClsNet) and start point localization (using LocNet). ClsNet focuses on per-frame labeling and predicts action score distributions online. Based on the predicted action scores of the past and current frames, LocNet conducts class-agnostic start detection by optimizing long-term localization rewards using policy gradient methods. The proposed framework is validated on two large-scale datasets, THUMOS'14 and ActivityNet. The experimental results show that StartNet significantly outperforms the state-of-the-art by 15%-30% p-mAP under the offset tolerance of 1-10 seconds on THUMOS'14, and achieves comparable performance on ActivityNet with 10 times smaller time offset. Mingfei Gao, Larry Davis 0001, Richard Socher, Caiming Xiong |
ICCV | 1 |
| 2019 | Temporal Recurrent Networks for Online Action DetectionabstractMost work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions as soon as each video frame arrives, based only on current and historical observations. In this paper, we propose a novel framework, the Temporal Recurrent Network (TRN), to model greater temporal context of each frame by simultaneously performing online action detection and anticipation of the immediate future. At each moment in time, our approach makes use of both accumulated historical evidence and predicted future information to better recognize the action that is currently occurring, and integrates both of these into a unified end-to-end architecture. We evaluate our approach on two popular online action detection datasets, HDD and TVSeries, as well as another widely used dataset, THUMOS'14. The results show that TRN significantly outperforms the state-of-the-art. Mingfei Gao, Yi-Ting Chen 0001, Larry Davis 0001, David Crandall |
ICCV | 2 |
| 2019 | Goal-oriented Object Importance Estimation in On-road Driving VideosabstractWe formulate a new problem as Object Importance Estimation (OIE) in on-road driving videos, where the road users are considered as important objects if they have influence on the control decision of the ego-vehicle's driver. The importance of a road user depends on both its visual dynamics, e.g., appearance, motion and location, in the driving scene and the driving goal, e.g., the planned path, of the ego vehicle. We propose a novel framework that incorporates both visual model and goal representation to conduct OIE. To evaluate our framework, we collect an on-road driving dataset at traffic intersections in the real world and conduct human-labeled annotation of the important objects. Experimental results show that our goal-oriented method outperforms baselines and has much more improvement on the left-turn and right-turn scenarios. Furthermore, we explore the possibility of using object importance for driving control prediction and demonstrate that binary brake prediction can be improved with the information of object importance. Mingfei Gao, Ashish Tawari, Sujitha Martin |
ICRA | 1 |
| 2018 | Dynamic Zoom-In Network for Fast Object Detection in Large ImagesabstractWe introduce a generic framework that reduces the computational cost of object detection while retaining accuracy for scenarios where objects with varied sizes appear in high resolution images. Detection progresses in a coarse-to-fine manner, first on a down-sampled version of the image and then on a sequence of higher resolution regions identified as likely to improve the detection accuracy. Built upon reinforcement learning, our approach consists of a model (R-net) that uses coarse detection results to predict the potential accuracy gain for analyzing a region at a higher resolution and another model (Q-net) that sequentially selects regions to zoom in. Experiments on the Caltech Pedestrians dataset show that our approach reduces the number of processed pixels by over 50% without a drop in detection accuracy. The merits of our approach become more significant on a high resolution test set collected from YFCC100M dataset, where our approach maintains high detection performance while reducing the number of processed pixels by about 70% and the detection time by over 50%. Mingfei Gao, Ruichi Yu, Ang Li 0001, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 1 |
| 2018 | NISP: Pruning Networks Using Neuron Importance Score PropagationabstractTo reduce the significant redundancy in deep Convolutional Neural Networks (CNNs), most existing methods prune neurons by only considering the statistics of an individual layer or two consecutive layers (e.g., prune one layer to minimize the reconstruction error of the next layer), ignoring the effect of error propagation in deep networks. In contrast, we argue that for a pruned network to retain its predictive power, it is essential to prune neurons in the entire neuron network jointly based on a unified goal: minimizing the reconstruction error of important responses in the "final response layer" (FRL), which is the second-to-last layer before classification. Specifically, we apply feature ranking techniques to measure the importance of each neuron in the FRL, formulate network pruning as a binary integer optimization problem, and derive a closed-form solution to it for pruning neurons in earlier layers. Based on our theoretical analysis, we propose the Neuron Importance Score Propagation (NISP) algorithm to propagate the importance scores of final responses to every neuron in the network. The CNN is pruned by removing neurons with least importance, and it is then fine-tuned to recover its predictive power. NISP is evaluated on several datasets with multiple CNN models and demonstrated to achieve significant acceleration and compression with negligible accuracy loss. Ruichi Yu, Ang Li 0001, Chun-Fu Chen 0001, Jui-Hsin Lai, Vlad I. Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, Larry Davis 0001 |
CVPR | 7 |
| 2018 | C-WSL: Count-Guided Weakly Supervised Localization
Mingfei Gao, Ang Li 0001, Ruichi Yu, Vlad I. Morariu, Larry Davis 0001 |
ECCV (1) | 1 |
| 2015 | Learning discriminative occlusion feature for depth ordering inference on monocular imageabstractIn this paper, a novel depth ordering inference approach is presented. Our main insight is to integrate the discriminative feature selection, occlusion feature learning and same-layer (S-L) relationship judgement into a uniform sparsity based classification objective, which cannot only supply the precise segmentation for the occlusion edge, but also reduce the solution space for the depth ordering inference efficiently. In addition, a novel triple descriptor is adopted to judge the foreground relationship, which is more discriminative than conversional local cues and can further reduce the solution space. The inference is executed by finding a valid path on a directed graph model. We validate our approach on the Cornell depth-order dataset and the NYU 2 dataset, and the convincing experimental results demonstrate the effectiveness of our approach. Anlong Ming, Baofeng Xun, Jia Ni, Mingfei Gao, Yu Zhou 0016 |
ICIP | 4 |
| 2015 | Resource management in device-to-device underlaying cellular networkabstractThis paper addresses resource management problem in a heterogeneous network consisting of cellular users and multiple D2D pairs. The D2D users share spectrum with cellular uplink under the QoS constraint of cellular system and pay for the interference they cause. We propose a two-step resource management scheme to optimize D2D user's transmitting power and the spectrum efficiency of the network. Firstly, the interference pricing and power allocation problem between a D2D pair and their allocated cellular uplink is formulated as a Stackelberg game. Then the Stackelberg equilibrium, i.e., the optimal interference price and transmitting power that maximize utilities of two players (the base station and the D2D transmitter), is obtained in close form. In the second step, with the close form expressions as important parameters, the situation is extended to multiple D2D pairs and a Hungarian algorithm based method is utilized to assign spectrum band to each D2D pair with the purpose of maximizing system spectrum efficiency. Simulation results demonstrate the advantageous performance of our scheme in network capacity and spectrum efficiency. Yuchi Zhang, Mingfei Gao, Qixun Zhang, Huidi Li, Zhiyong Feng 0001 |
WCNC | 3 |
| 2014 | Improved Energy Detector for Full Duplex SensingabstractUnder current popularity of full duplex radio, this paper explores the adaption of energy detector (ED) to full duplex scenario. To improve sensing performance,firstly a correlation based channel estimation algorithm is formulated to cancel out self-interference. We then notice traditional ED has a long delay in finding sensing error. To tackle this problem, sliding window ED is proposed, in which samples utilized for different sensing instances are allowed to overlap. Moreover, it is also observed that sensing performance degrades seriously when primary user sate changes in the sensing period, which happens frequently in full duplex sensing. Therefore, weighted ED is proposed, in which samples at the end of the sensing period are assigned bigger weights. Simulation results show that our strategies can better adapt energy detector to full duplex sensing with marginal loss in accuracy. Xiao Yan 0002, Mingfei Gao, Jian Yang 0021, Yuchi Zhang, Zhiyong Feng 0001, Yifan Zhang 0003 |
VTC Fall | 3 |
| 2014 | Full-Duplex Spectrum Sensing Scheme Based on Phase DifferenceabstractWith recent breakthroughs in full-duplex radio, there is a growing trend to combine it with cognitive radio. Thus the topic of spectrum sensing in full-duplex scenario is addressed in this paper. Firstly, the concept of full-duplex spectrum sensing and its advantage over traditional half-duplex spectrum sensing are analyzed. Then a correlation based least square algorithm is formulated to cancel out self-interference, in which no preamble is required. We notice that the phase distribution of noise differs greatly from that of noise-perturbed signals. Therefore, a novel spectrum sensing scheme using phase difference as test statistics is introduced. Its theoretical performance is also analyzed by approximating the test statistics as Gaussian distribution. The proposed detector is simple and immune to noise uncertainty due to the independence of its threshold on noise power. Simulation results show that robust sensing performance can be achieved in full-duplex scenario using correlation based least square and the phase based sensing scheme. Jian Yang 0021, Ying Zhu 0005, Mingfei Gao, Yifan Zhang 0003, Zhiyong Feng 0001, Yuchi Zhang |
VTC Fall | 4 |
| 2014 | A novel spectrum sensing scheme based on phase differenceabstractSpectrum sensing is one of the most challenging tasks in cognitive radio. Unfortunately traditional schemes fail to balance between accuracy and complexity, which are the key indicators for the performance of spectrum sensing. In this paper, a new spectrum sensing scheme based on phase difference is proposed. Through analyzing the distributions of phase difference between adjacent samples of noise and noise-perturbed primary signal, we notice that the mean of phase difference varies from noise when primary signal is present. On this basis, a novel sensing scheme using the accumulation of phase difference as test statistics is formulated. Then the analytical performance of our scheme is derived and its complexity is analyzed. Our proposed scheme is simple, accurate and immune to noise uncertainty. Simulation results show that our scheme outperforms conventional energy detection and can achieve a detection probability of 99% at −7dB signal-to-noise ratio using 500 data samples. Jian Yang 0021, Xiao Yan 0002, Mingfei Gao, Hao Lian, Han Zhang 0006, Zhiyong Feng 0001, Yifan Zhang 0003 |
WCNC | 3 |
| 2013 | Channel Correlation Assisted Fast Spectrum SensingabstractTo dynamically and efficiently utilize the vacant spectrum resources, cognitive radio is proposed as a potential solution, where spectrum sensing is one of the indispensable techniques. As one of the remaining issues in spectrum sensing, how to achieve the fast and energy efficient spectrum sensing in face of a wide band spectrum with an acceptable sensing accuracy is still a big challenge. Therefore, to reduce the spectrum sensing cost, a fast spectrum sensing scheme is proposed in this paper. Firstly, the channels of the same service are classified into highly correlated groups by a modified version of the greedy algorithm for set covering problem. In each group, only one representative channel (RC) is detected and the current busy-idle states of other channels, namely, estimated channels (EC), can be inferred according to their historical states and the current state of RC by the proposed joint Markov and channel correlation algorithm. Based on the real-time measurement results on GSM service and TV service in Beijing, the proposed scheme is proved and verified to be efficient on the premise of low estimated error. Mingfei Gao, Xiao Yan 0002, Ying Zhu 0005, Qixun Zhang, Zhiyong Feng 0001, Baoling Liu |
VTC Fall | 1 |
| 2013 | Sensing Performance of Improved Cyclostationary Detector with Multiple Correlated Antennas over Nakagami Fading ChannelabstractIn this paper, we analyze the sensing performance of multi-cycle cyclostationary (MC) detection- based spectrum sensing in a secondary user (SU) possessing multiple correlated antennas when the channel from the primary user (PU) to the SU suffers from Nakagami fading. We first propose an improved MC detector, aiming at reducing the computational complexity of conventional MC detector by simplifying its test statistic. Compared with the conventional one, our proposed detector achieves low-computational complexity and high- accuracy on sensing performance. Based on the proposed detector, square-law combining (SLC) diversity of multiple antennas technique is introduced to improve the detection capability over the Nakagami fading channel. Subsequently, the effect of multiple correlated antennas is investigated. A special correlated antenna case of a linear array of 2 and 4 arbitrarily correlation is then treated. The corresponding closed-form average detection probability is derived by using the moment generation function (MGF) approach. Finally, results show the efficiency and reliability of the proposed detector and the degradation on sensing performance with correlated antennas over the Nakagami fading channel. Ying Zhu 0005, Zhiyong Feng 0001, Mingfei Gao, Qixun Zhang |
VTC Fall | 3 |