Yike Ma

dblp:94/10052 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 14 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 IQGS: Instance Query-based Gaussian Segmentation
abstract
In recent years, Gaussian scene representations have achieved a series of promising results in 3D reconstruction. Compared to the previous 3DGS paradigm, the latest reconstruction approach 2DGS can achieve more accurate geometric representation using fewer Gaussian points. Accordingly, developing a panoramic segmentation algorithm suitable for 2DGS-reconstructed scenes is of significant importance. However, existing segmentation methods are primarily designed for 3DGS. They either fail to account for all objects in complex segmentation scenes or suffer from significant performance degradation when applied to 2D Gaussian scenes. Moreover, these methods consistently exhibit poor cross-dataset generalization. To address these issues, we propose IQGS, a segmentation framework applicable to 2DGS representations. Specifically, IQGS employs per-instance query and relaxed object-level supervision instead of strict pixel-level ID supervision, effectively mitigating the segmentation performance degradation that occurs when applied to 2DGS. At the same time, by learning features independent of specific object ID assignments, IQGS enhances its ability to generalize across diverse datasets. Our method achieves impressive panoramic segmentation results across multiple datasets, with an average mIoU of 66.6%, surpassing the state-of-the-art method Gaussian Grouping, which achieves 57.17%.
Yichao Gao, Xinyuan Liu 0003, Yike Ma
AAAI3
2026 Semantic-decoupled spatial partition guided point-supervised oriented object detection
Xinyuan Liu 0003, Yike Ma, Chenggang Yan 0001
Pattern Recognit.4
2025 Exact: Exploring Space-Time Perceptive Clues for Weakly Supervised Satellite Image Time Series Semantic Segmentation
abstract
Automated crop mapping through Satellite Image Time Series (SITS) has emerged as a crucial avenue for agricultural monitoring and management. However, due to the low resolution and unclear parcel boundaries, annotating pixel-level masks is exceptionally complex and time-consuming in SITS. This paper embraces the weakly supervised paradigm (i.e., only image-level categories available) to liberate the crop mapping task from the exhaustive annotation burden. The unique characteristics of SITS give rise to several challenges in weakly supervised learning: (1) noise perturbation from spatially neighboring regions, and (2) erroneous semantic bias from anomalous temporal periods. To address the above difficulties, we propose a novel method, termed exploring space-time perceptive clues (Exact). First, we introduce a set of spatial clues to explicitly capture the representative patterns of different crops from the most class-relative regions. Besides, we leverage the temporal-to-class interaction of the model to emphasize the contributions of pivotal clips, thereby enhancing the model perception for crop regions. Building upon the space-time perceptive clues, we derive the clue-based CAMs to effectively supervise the SITS segmentation network. Our method demonstrates impressive performance on various SITS benchmarks. Remarkably, the segmentation network trained on Exact-generated masks achieves 95% of its fully supervised performance, showing the bright promise of weakly supervised paradigm in crop mapping scenario. Our code will be publicly available here.
Jiayu Xiao, Tianxiang Xiao, Yike Ma
CVPR5
2025 TopoPoint: Enhance Topology Reasoning via Endpoint Detection in Autonomous Driving
abstract
Topology reasoning, which unifies perception and structured reasoning, plays a vital role in understanding intersections for autonomous driving. However, its performance heavily relies on the accuracy of lane detection, particularly at connected lane endpoints. Existing methods often suffer from lane endpoints deviation, leading to incorrect topology construction. To address this issue, we propose TopoPoint, a novel framework that explicitly detects lane endpoints and jointly reasons over endpoints and lanes for robust topology reasoning. During training, we independently initialize point and lane query, and proposed Point-Lane Merge Self-Attention to enhance global context sharing through incorporating geometric distances between points and lanes as an attention mask . We further design Point-Lane Graph Convolutional Network to enable mutual feature aggregation between point and lane query. During inference, we introduce Point-Lane Geometry Matching algorithm that computes distances between detected points and lanes to refine lane endpoints, effectively mitigating endpoint deviation. Extensive experiments on the OpenLane-V2 benchmark demonstrate that TopoPoint achieves state-of-the-art performance in topology reasoning (48.8 on OLS). Additionally, we propose DET$_p$ to evaluate endpoint detection, under which our method significantly outperforms existing approaches (52.6 v.s. 45.2 on DET$_p$). The code is released at https://github.com/Franpin/TopoPoint.
Yanping Fu, Xinyuan Liu 0003, Tianyu Li 0004, Yike Ma
NeurIPS4
2025 DASC-SPT: Towards Self-Supervised Panoramic Semantic Segmentation
abstract
Self-Supervised Semantic Segmentation, aiming to leverage masses of unlabeled data for boosting semantic segmentation, has been rapidly emerging as an active task in recent years. However, existing self-supervised semantic segmentation approaches mainly focus on planar images, leaving multiple distorted objects encountered in panoramic images unexplored due to the formidable challenge of handling heterogeneous degrees of distortions across different locations. In this paper, we propose a novel Self-Supervised Panoramic Semantic Segmentation model, termed DASC-SPT, built upon the mainstream contrastive learning framework. Towards distortions in panoramic images, we present two structures to better learn from distorted features by applying planar images. For the input images of self-supervision, we design a Spherical Projection Transformation (SPT) strategy that involves randomly projecting planar images onto various locations of the sphere to introduce the distortions. For pixel-wise distorted features, we construct a Deformation-aware Sampling Consistency (DASC) framework to further utilize the shared content and discrepancies caused by different distortions of paired views, where the deformation-aware consistency can be quantified on pixel-wise features. Both of the two components facilitate the model to adapt to distortions and boost panoramic semantic segmentation. Extensive comprehensive experiments on three panoramic datasets demonstrate the effectiveness and superiority of DASC-SPT approach.
Tianlong Tan, Bin Chen 0021, Hongliang Cao, Chenggang Yan 0001, Yike Ma
WACV5
2025 A Robotic AI Algorithm for Fusing Generative Large Models in Agriculture Internet of Things
abstract
Robot perception difficulties in complex environments seriously affect robot operational efficiency. The application of generative big model (GBM) technology on robots through the Agriculture Internet of Things (AIoT) can solve the difficulty of low-operational efficiency. Therefore, this article proposes an automatic decision-making localization algorithm for robots fused with GBMs in AIoT. First, in the AIoT, the complex agricultural scene information is acquired in real-time by the Realsense D435i equipped on the robot, which is accessed in the GBM through a proprietary network to realize real-time intelligent sensing and control between the environmental information and the robot. Then, the innovative reinforcement learning method based on human solid feedback (S-RLHF) and the automatic generation method of weakly supervised fine-tuning data (WS-FT DAGM) are designed in the generative large model. At the same time, by combining the characteristics of the robot operation, two generative significant model recommendation methods are designed, which solves the problem of the difficulty of the target perception in the complex agricultural scene. Finally, by integrating the AIoT and generative large models, the critical method of real-time analysis of crop shading characteristics by the large model is innovatively proposed to solve the problem of low efficiency in robot operation. In the robot test experiments, the operation efficiency using the fusion of AIoT and generative large model reaches more than 92%, significantly improving the operation efficiency compared to the small model (CNN) method without AIoT and generative large model in the traditional agricultural robot.
Guangyu Hou, Runxin Niu, Haihua Chen 0003, Yancong Wang, Zhenqiang Zhu, Yike Ma, Tongbin Li
IEEE Internet Things J.7
2024 Rethinking Boundary Discontinuity Problem for Oriented Object Detection
abstract
Oriented object detection has been developed rapidly in the past few years, where rotation equivariance is crucial for detectors to predict rotated boxes. It is expected that the prediction can maintain the corresponding rotation when objects rotate, but severe mutation in angular prediction is sometimes observed when objects rotate near the boundary angle, which is well-known boundary discontinuity problem. The problem has been long believed to be caused by the sharp loss increase at the angular boundary, and widely used joint-optim IoU-like methods deal with this problem by loss-smoothing. However, we experimentally find that even state-of-the-art IoU-like methods actually fail to solve the problem. On further analysis, we find that the key to solution lies in encoding mode of the smoothing function rather than in joint or independent optimization. In existing IoU-like methods, the model essentially attempts to fit the angular relationship between box and object, where the break point at angular boundary makes the predictions highly unstable. To deal with this issue, we propose a dual-optimization paradigm for angles. We decouple reversibility and joint-optim from single smoothing function into two distinct entities, which for the first time achieves the objectives of both correcting angular boundary and blending angle with other parameters. Extensive experiments on multiple datasets show that boundary discontinuity problem is well-addressed. More-over, typical IoU-like methods are improved to the same level without obvious performance gap. The code is available at https://github.com/hangxu-cv/cvpr24acm.
Xinyuan Liu 0003, Yike Ma, Zunjie Zhu, Chenggang Yan 0001
CVPR4
2024 MISA: MIning Saliency-Aware Semantic Prior for Box Supervised Instance Segmentation
Jiayu Xiao, Yike Ma
IJCAI4
2024 TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes
abstract
As an emerging task that integrates perception and reasoning, topology reasoning in autonomous driving scenes has recently garnered widespread attention. However, existing work often emphasizes "perception over reasoning": they typically boost reasoning performance by enhancing the perception of lanes and directly adopt vanilla MLPs to learn lane topology from lane query. This paradigm overlooks the geometric features intrinsic to the lanes themselves and are prone to being influenced by inherent endpoint shifts in lane detection. To tackle this issue, we propose an interpretable method for lane topology reasoning based on lane geometric distance and lane query similarity, named TopoLogic. This method mitigates the impact of endpoint shifts in geometric space, and introduces explicit similarity calculation in semantic space as a complement. By integrating results from both spaces, our methods provides more comprehensive information for lane topology. Ultimately, our approach significantly outperforms the existing state-of-the-art methods on the mainstream benchmark OpenLane-V2 (23.9 v.s. 10.9 in TOP$_{ll}$ and 44.1 v.s. 39.8 in OLS on subsetA). Additionally, our proposed geometric distance topology reasoning method can be incorporated into well-trained models without re-training, significantly enhancing the performance of lane topology reasoning. The code is released at https://github.com/Franpin/TopoLogic.
Yanping Fu, Wenbin Liao, Xinyuan Liu 0003, Yike Ma
NeurIPS5
2023 Gaussian Label Distribution Learning for Spherical Image Object Detection
abstract
Spherical image object detection emerges in many applications from virtual reality to robotics and automatic driving, while many existing detectors use$l_{n}$-norms loss for regression of spherical bounding boxes. There are two intrinsic flaws for$l_{n}$-norms loss, i.e., independent optimization of parameters and inconsistency between metric (dominated by IoU) and loss. These problems are common in planar image detection but more significant in spherical image detection. Solution for these problems has been extensively discussed in planar image detection by using IoU loss and related variants. However, these solutions cannot be migrated to spherical image object detection due to the undifferentiable of the Spherical IoU (SphIoU). In this paper, we design a simple but effective regression loss based on Gaussian Label Distribution Learning (GLDL) for spherical image object detection. Besides, we observe that the scale of the object in a spherical image varies greatly. The huge differences among objects from different categories make the sample selection strategy based on SphIoU challenging. Therefore, we propose GLDL-ATSS as a better training sample selection strategy for objects of the spherical image, which can alleviate the drawback of IoU threshold-based strategy of scale-sample imbalance. Extensive results on various two datasets with different baseline detectors show the effectiveness of our approach.
Xinyuan Liu 0003, Qiang Zhao 0005, Yike Ma, Chenggang Yan 0001
CVPR4
2023 Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods
abstract
Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i.e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical approximate IoU methods have been proposed recently. We find that the key of these approximate methods is to map spherical boxes to planar boxes. However, there exists two problems in these methods: (1) they do not eliminate the influence of panoramic image distortion; (2) they break the original pose between bounding boxes. They lead to the low accuracy of these methods. Taking the two problems into account, we propose a new sphere-plane boxes transform, called Sph2Pob. Based on the Sph2Pob, we propose (1) an differentiable IoU, Sph2Pob-IoU, for spherical boxes with low time-cost and high accuracy and (2) an agent Loss, Sph2Pob-Loss, for spherical detection with high flexibility and expansibility. Extensive experiments verify the effectiveness and generality of our approaches, and Sph2Pob-IoU and Sph2Pob-Loss together boost the performance of spherical detectors. The source code is available at https://github.com/AntXinyuan/sph2pob.
Xinyuan Liu 0003, Bin Chen 0021, Qiang Zhao 0005, Yike Ma, Chenggang Yan 0001
IJCAI5
2023 Multi-Robot Task Allocation in Agriculture Scenarios Based on the Improved NSGA-II Algorithm
abstract
Agricultural multi-robot task allocation is an important direction of intelligent agricultural robots. In this paper, to ensure the shortest total movement distance of robots and balance the workload, we transform the task allocation of agricultural multi-robot into a multi objective multiple traveling salesman problem (MO-MTSP). A processing framework is also designed to adjust the allocation plan in real-time by monitoring the status of robots and tasks. In this processing framework, an improved Non-dominated genetic algorithm (INSGA-II) is utilized to achieve better performance of task allocation. The improvement includes: 1) An objective function is designed to ensure a balanced task workload distribution for each robot considering the actual conditions of the farmland; 2) A neighborhood search algorithm is used to accelerate the convergence of the algorithm by adjusting the initial population plan; 3) An adaptive adjustment mechanism for crossover and mutation probabilities is proposed to enhance the flexibility of the task allocation plan searching. Finally, the proposed algorithm is evaluated on both public datasets and actual agricultural plot datasets. Experimental results show that the improved algorithm can achieve better allocation results, demonstrating the practical applicability of the algorithm in task allocation of multiple robots employed in agriculture scenarios.
Zaiwang Lu, Long Long, Yike Ma, LeiLi LeiLi, Jintao Li 0001
VTC Fall4
2022 Unbiased IoU for Spherical Image Object Detection
abstract
As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while it is challenging for spherical ones. Some existing methods utilize planar bounding boxes to represent spherical objects. However, they are biased due to the distortions of spherical objects. Others use spherical rectangles as unbiased representations, but they adopt excessive approximate algorithms when computing the IoU. In this paper, we propose an unbiased IoU as a novel evaluation criterion for spherical image object detection, which is based on the unbiased representations and utilize unbiased analytical method for IoU calculation. This is the first time that the absolutely accurate IoU calculation is applied to the evaluation criterion, thus object detection algorithms can be correctly evaluated for spherical images. With the unbiased representation and calculation, we also present Spherical CenterNet, an anchor free object detection algorithm for spherical images. The experiments show that our unbiased IoU gives accurate results and the proposed Spherical CenterNet achieves better performance on one real-world and two synthetic spherical object detection datasets than existing methods.
Bin Chen 0021, Yike Ma, Bailan Feng, Chenggang Yan 0001, Qiang Zhao 0005
AAAI4
2022 PANDORA: A Panoramic Detection Dataset for Object with Orientation
Qiang Zhao 0005, Yike Ma, Bailan Feng, Chenggang Yan 0001
ECCV (8)3
2022 Global Boundary Refinement for Semantic Segmentation via Optimal Transport
Shuaibin Zhang, Hao Liu 0068, Yike Ma, Qiang Zhao 0005
PRICAI (3)4
2021 Bipartite Matching for Crowd Counting with Point Supervision
abstract
For crowd counting task, it has been demonstrated that imposing Gaussians to point annotations hurts generalization performance. Several methods attempt to utilize point annotations as supervision directly. And they have made significant improvement compared with density-map based methods. However, these point based methods ignore the inevitable annotation noises and still suffer from low robustness to noisy annotations. To address the problem, we propose a bipartite matching based method for crowd counting with only point supervision (BM-Count). In BM-Count, we select a subset of most similar pixels from the predicted density map to match annotated pixels via bipartite matching. Then loss functions can be defined based on the matching pairs to alleviate the bad effect caused by those annotated dots with incorrect positions. Under the noisy annotations, our method reduces MAE and RMSE by 9% and 11.2% respectively. Moreover, we propose a novel ranking distribution learning framework to address the imbalanced distribution problem of head counts, which encodes the head counts as classification distribution in the ranking domain and refines the estimated count map in the continuous domain. Extensive experiments on four datasets show that our method achieves state-of-the-art performance and performs better crowd localization.
Hao Liu 0068, Qiang Zhao 0005, Yike Ma
IJCAI3
2021 Dense Scale Network for Crowd Counting
abstract
Crowd counting has been widely studied by computer vision community in recent years. Due to the large scale variation, it remains to be a challenging task. Previous methods adopt either multi-column CNN or single-column CNN with multiple branches to deal with this problem. However, restricted by the number of columns or branches, these methods can only capture a few different scales and have limited capability. In this paper, we propose a simple but effective network called DSNet for crowd counting, which can be easily trained in an end-to-end fashion. The key component of our network is the dense dilated convolution block, in which each dilation layer is densely connected with the others to preserve information from continuously varied scales. The dilation rates in dilation layers are carefully selected to prevent the block from gridding artifacts. To further enlarge the range of scales covered by the network, we cascade three blocks and link them with dense residual connections. We also introduce a novel multi-scale density level consistency loss for performance improvement. To evaluate our method, we compare it with state-of-the-art algorithms on five crowd counting datasets (ShanghaiTech, UCF-QNRF, UCF_CC_50, UCSD and WorldExpo'10). Experimental results demonstrate that DSNet can achieve the best overall performance and make significant improvements.
Hao Liu 0068, Yike Ma, Xi Zhang 0008, Qiang Zhao 0005
ICMR3
2021 Image stitching via deep homography estimation
Qiang Zhao 0005, Yike Ma, Chunfeng Yao, Bailan Feng
Neurocomputing2
2020 Dilated Convolutional Neural Networks for Panoramic Image Saliency Prediction
abstract
Saliency prediction is an important way to understand human's behavior and has a wide range of applications. Although lots of algorithms have been designed to predict saliency for planar images, there are few works for 360° images. In this paper, we propose an encoder-decoder network for panoramic image saliency prediction. Dilated convolutional layers are deployed in the network, which can extract more representative features and improve the accuracy of saliency prediction. To deal with the image distortions in 360° images, our network takes cube map format as input and processes six faces of cube map simultaneously. Respecting the saliency distribution of ground truth, we also propose a new data augmentation method to train the network, which is validated to be helpful for performance improvement. Extensive experiments show that our method gives new state-of-the-art results on 360° image saliency prediction.
Youqiang Zhang, Yike Ma, Qiang Zhao 0005
ICASSP3
2020 Light Field Reconstruction Using Dynamically Generated Filters
Xiuxiu Jing, Yike Ma, Qiang Zhao 0005, Ke Lyu
MMM (1)2
2020 Panoramic Light Field From Hand-Held Video and Its Sampling for Real-Time Rendering
abstract
By providing angular and spatial information of light rays, light field images are widely used in many applications. To capture large field of view light field, existing approaches either stitch small field of view light fields, whose apertures are also small, or leverage specialized equipments, which are not accessible to ordinary users. In this paper, we present a method to extract fully 360° field of view panoramic light fields from densely sampled hand-held videos. Unlike previous works, our method handles the large exposure variation and motion blur problems, which are common in the panoramic capture and hand-held videos. To provide real-time rendering performance, our method applies light field sampling to the extracted panoramic light field. Based on the unstructured representation, our method can sample the light field without explicit parameterization. We formulate the sampling problem into a set multicover problem and solve it globally using integer linear programming. Compared with the greedy based method, our approach can use much less video frames to represent the panoramic light field, but still can get better rendering results.
Qiang Zhao 0005, Jing Lv, Yike Ma, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Near-infrared Image Guided Neural Networks for Color Image Denoising
abstract
Noisy color image and guided near-infrared (NIR) image can be jointly employed to eliminate noise and enhance details. Existing methods mostly rely on explicit designed filters and hand-crafted objective function optimization. These methods usually introduce erroneous structures from guidance signal. Besides, they are time-consuming and not suitable for real time applications. In this paper, we come up with a learning based method. The noisy color image and NIR image are fused, then fed into a fully convolutional neural network. The network learns a directly map from degraded image to restored sharp image. Our architecture can effectively eliminate image noise and transfer detail structure from guided image. Our trained network accepts any resolution of input image and runs in constant time. We evaluate the presented approach on both synthetic and real images. Results show that our approach outperforms the state-of-art methods.
Yike Ma, Junbo Guo, Qiang Zhao 0005, Yongdong Zhang 0001
ICASSP3
2019 Freely Explore the Scene with 360°Field of View
abstract
By providing 360° field of view, spherical panoramas are widely used in virtual reality (VR) systems and street view services. However, due to bandwidth or storage limitations, existing systems only provide sparsely captured panoramas and have limited interaction modes. Although there are methods that can synthesize novel views based on captured panoramas, the generated novel views all lie on the lines connecting existing views. Therefore these methods do not support free viewpoint navigation. In this paper, we propose a new panoramic image based rendering method. Our method takes pre-captured images as input and can synthesize panoramas at novel views that are far from input camera positions. Thus, it supports to freely explore the scene with 360° field of view.
Yike Ma, Juan Cao 0001, Qiang Zhao 0005, Yongdong Zhang 0001
VR3
2019 Scene-adaptive coded aperture imaging
Yike Ma, Ke Gao 0012, Yongdong Zhang 0001
Multim. Tools Appl.3
2018 Wide Range Depth Estimation from Binocular Light Field Camera
Yike Ma, Guoqing Jin, Qiang Zhao 0005
BMVC3
2018 Semantic Preserving Hash Coding Through VAE-GAN
abstract
This paper proposes a novel framework for fast image retrieval. The proposed framework combines variational autoencoder with generative adversarial network to generate content preserving images for learning-based hashing. By accepting real image and systhesized image in a pairwise form, a semantic perserving binary mapping model is learned using pairwise ranking loss under an adversarial generative process. Extensive experiments on several benchmark datasets demonstrate that the proposed method shows substantial improvement over the state-of-the-art hashing methods.
Guoqing Jin, Dongming Zhang 0004, Junbo Guo, Yike Ma, Yongdong Zhang 0001
ICIP5
2018 Panoramic Light Field Video Acquisition
abstract
Due to the limited field of view (FOV), the light field cameras can only capture a portion of the scene. Although there are methods designed to increase the FOV of light field images, these approaches stitch multiple sequentially captured light field images into a panoramic one and are restricted to static scenes. In this paper, we present a panoramic light field video acquisition system for dynamic scenes. Our system consists of 48 light field cameras and can capture panoramic light field videos at 6.3 fps. Two challenging problems, i.e. large volume data transmission, storage and camera synchronization, have been solved in our system. By using the light field stitching algorithm, the generated panoramic light field videos have 9×9 angular resolution, 6283×1200 spatial resolution and 360°×70° FOV, enabling us to do video refocusing, video focus tracking and other features not supported by light field static images.
Jing Lv, Qiang Zhao 0005, Yike Ma, Yongdong Zhang 0001
ICME5
2018 Distortion-aware CNNs for Spherical Images
abstract
Convolutional neural networks are widely used in computer vision applications. Although they have achieved great success, these networks can not be applied to 360 spherical images directly due to varying distortion effect. In this paper, we present distortion-aware convolutional network for spherical images. For each pixel, our network samples a non-regular grid based on its distortion level, and convolves the sampled grid using square kernels shared by all pixels. The network successively approximates large image patches from different tangent planes of viewing sphere with small local sampling grids, thus improves the computational efficiency. Our method also deals with the boundary problem, which is an inherent issue for spherical images. To evaluate our method, we apply our network in spherical image classification problems based on transformed MNIST and CIFAR-10 datasets. Compared with the baseline method, our method can get much better performance. We also analyze the variants of our network.
Qiang Zhao 0005, Yike Ma, Guoqing Jin, Yongdong Zhang 0001
IJCAI4
2018 Light Field Foreground Matting Based on Defocus and Correspondence
Jianshe Zhou, Tuya Naren, Yike Ma, Jie Liu 0022
MMM (1)4
2018 Variable aperture panoramic imaging
Jing Lv, Qiang Zhao 0005, Yike Ma, Yongdong Zhang 0001
Multim. Tools Appl.4
2018 Spherical Superpixel Segmentation
abstract
These days, superpixel algorithms are widely used in computer vision and multimedia applications. However, existing algorithms are designed for planar images, which are less suited to deal with wide angle images. In this paper, we present a superpixel segmentation method for 360° spherical images. Unlike previous methods, our approach explicitly considers the geometry for spherical images and makes clustering to spherical image pixels. It starts with the seeds defined by Hammersley points sampled on the sphere, then iterates between assignment step and update step, which are both based on the distance metric respecting spherical geometry. We evaluate our method on the transformed Berkeley segmentation dataset and panorama segmentation dataset collected by ourselves. Experimental results show that our method can gain better performance in terms of adherence to image boundaries and superpixel structural regularity. Furthermore, superpixels generated by our method can reserve the coherence across image boundaries and all have closed contours.
Qiang Zhao 0005, Yike Ma, Jiawan Zhang, Yongdong Zhang 0001
IEEE Trans. Multim.3
2015 Lenselet image compression scheme based on subaperture images streaming
abstract
Plenoptic cameras capture the light field in a scene with a single shot and produce lenselet images. From a lenselet image, light field can be reconstructed, with which we can render images with different viewpoints and focal length. Because of large volume data, high efficient image compression scheme for storage and transmission is urgent. Containing 4D light field information, lenselet images have much more redundant information than traditional 2D images. In this paper, we propose a subaperture images streaming scheme to compress lenselet images, in which rotation scan mapping is adopted to further improve compression efficiency. The experiment results show our approach can efficient compress the redundancy in lenselet images and outperform traditional image compression method.
Jun Zhang 0007, Yike Ma, Yongdong Zhang 0001
ICIP3
2015 Automatic foreground segmentation using light field images
abstract
Foreground segmentation is a fundamental method in computer vision. Traditional foreground segmentation algorithms are sensitive to blurry degree of background, smooth foreground regions and camouflage foreground. To deal with these problems, we use light field images as input by exploiting its focusness cue. In this paper, we propose an automatic foreground segmentation algorithm for light field images. Firstly we calculate focusness measures in refocused stack and oversegment all focus images. Then the graph cut framework is introduced to generate foreground result. Experiments show that our method has a superior performance against counterparts and is more robust in scenes like ambiguous foreground.
Yike Ma, Yongdong Zhang 0001
VCIP3
2014 Inverse optimal design of spacecraft rendezvous problem with disturbances
abstract
This paper investigates the stabilization problem of spacecraft rendezvous with target spacecraft in an arbitrary elliptical orbit. A linearized dynamic model, obtained from the Hill-Clohessy-Wiltshire (HCW) equations, is used to describe the relative motion of two spacecrafts. An inverse optimal method is introduced to deal with the stabilization problem in presence of external disturbances. With Lyapunov analysis, A group of inverse optimal control laws is presented, which guarantees the input-to-state stabilization of the whole system, and at the same time, is optimal with respect to a performance index incorporating a penalty on the states, the disturbance acceleration, and the control effort. Simulation results are presented to elucidate the effectiveness of the control strategy.
Yike Ma, Haibo Ji
ICARCV1
2013 Highly parallel mode decision method for HEVC
abstract
High Efficiency Video Coding (HEVC) standard achieves double compression efficiency compared to H.264/AVC with the adoption of more flexible coding structure and advanced coding tools. On the other hand, the coding mode space is too large and it's very time consuming for an HEVC encoder to search for the best coding mode. With the development of multi-core or many-core computing architecture, parallelizing HEVC encoding on such platforms is an efficient approach to fulfill the high computational requirement. In this paper, we exploit the potential parallelism in HEVC mode decision (MD) process and propose a highly parallel MD method which works in a motion estimation region (MER). Specifically, we analyze and remove data dependencies that hinder parallel MD, including motion estimation (ME) dependencies and entropy coding dependencies, and then the MD computation for different blocks within the same MER can be computed concurrently. Experimental results show that our proposed parallel MD method gets an overall speed up of more than 14x with negligible quality loss (1.79% bit rate increasing), compared with the non-parallel baseline.
Jun Zhang 0007, Yike Ma, Yongdong Zhang 0001
PCS3
2012 Efficient Parallel Framework for H.264/AVC Deblocking Filter on Many-Core Platform
abstract
The H.264/AVC deblocking filter is becoming the performance bottleneck of H.264/AVC parallelization on many-core platform. Efficient parallelization of the deblocking filter on a many-core platform is challenging, because the deblocking filter has complicated data dependencies, which provide insufficient parallelism for so many cores. Furthermore, parallelization may have significant synchronization and load imbalance overhead. At present, research on the parallelizing deblocking filter on a many-core platform is rare and focuses on data-level parallelization. In this paper, we propose a three-step framework considering task-level segmentation and data-level parallelization to efficiently parallelize the deblocking filter. First, we review the entire deblocking filter process in 4 × 4 block edge-level and divide it into two parts: 1) boundary strength computation (BSC) and 2) edge discrimination and filtering (EDF), which increases the parallelism. Then, we apply the Markov empirical transition probability matrix and Huffman tree (METPMHT) to the BSC, which alleviate the load imbalance problem. Finally, we use an independent pixel connected area parallelization (IPCAP) for the EDF, which increases the parallelism and reduces the synchronization. In experiments, we apply our parallel method to the deblocking filter of the H.264/AVC reference software JM15.1 on the Tile64 platform without any Tile64 platform-based optimizations. Compared to the well-known 2D-wavefront method, the proposed method achieves on average 14.85, 17.83, and 10.60 times speed-up for QCIF, CIF, and HD videos using 62 cores, respectively.
Yongdong Zhang 0001, Chenggang Yan 0001, Yike Ma
IEEE Trans. Multim.4
2011 Optimization of stateful hardware acceleration in hybrid architectures
abstract
In many computing domains, hardware accelerators can improve throughput and lower power consumption, instead of executing functionally equivalent software on the general-purpose micro-processors cores. While hardware accelerators often are stateless, network processing exemplifies the need for stateful hardware acceleration. The packet oriented streaming nature of current networks enables data processing as soon as packets arrive rather than when the data of the whole network flow is available. Due to the concurrence of many flows, an accelerator must maintain and switch contexts between many states of the various accelerated streams embodied in the flows, which increases overhead associated with acceleration. We propose and evaluate dynamic reordering of requests of different accelerated streams in a hybrid on-chip/memory based request queue in order to reduce the associated overhead.
Xiaotao Chang, Yike Ma, Hubertus Franke, Kun Wang 0005, Rui Hou 0001, Hao Yu 0008, Terry Nelms
DATE2
2011 Parallel deblocking filter for H.264/AVC implemented on Tile64 platform
abstract
For the purpose of accelerating deblocking filter, which accounts for a significant percentage of H.264/AVC decoding time, some researchers use multi-core platforms to achieve the required performance. We study the problem under the context of many-core systems. Parallelization of deblocking filter on many-core platform is challenging not only because deblocking filter has complicated data dependencies which provides insufficient parallelism for so many cores but also because parallelization may have significant synchronization overhead. We present a new method to exploit the implicit parallelism and reduce the synchronization overhead. We apply our implementation to the deblocking filter of the H.264/AVC reference software JM15.1 on Tile64 platform. The proposed method achieves up to 817%, 604% and 532% speedup for CIF, SD and HD videos compared to the well-known wavefront method using 62 cores, respectively.
Chenggang Yan 0001, Yongdong Zhang 0001, Yike Ma, Licheng Chen, Lingjun Fan, Yasong Zheng
ICME4