VLDB 2026 Research / reviewers in the wild / expert
Gang Wang 0023
dblp:71/4292-23
· DBLP profile ↗
27ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0001-5283-7126ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A camera-light detection and ranging sensor online extrinsic calibration network based on mamba-like linear attention mechanism for unstructured off-road environments
Ren Xiao, Huayan Pu, Gang Wang 0023, Mingliang Zhou 0001, Jun Luo 0006 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Simultaneous Optimization of Hand-Eye and Robot-World Parameters Exploiting Accuracy Discrepancies in Robotic and Vision SensorabstractAccurate hand-eye and robot-world calibration remains crucial for vision-guided robotic applications. Existing methods predominantly focus on deriving closed-form or nonlinear optimization solutions for hand-eye equationAX=XBor hand-eye-robot-world calibration equationAX=YBunder various rotation parameterizations, while the critical role of input data refinement on calibration accuracy has been largely overlooked. This paper addresses this gap by proposing a pose graph optimization model that strategically enhances the precision of key pose chains in robotic vision systems, particularly under significant discrepancies between robotic positioning and visual measurement accuracies. Building upon the optimized pose chains associated with the hand-eye and robot-world parameters, a pair of dual equations forXandYis then constructed and solved via our proposed Kronecker product-based least squares algorithm, allowing simultaneous and accurate estimation on rotation matrices and translation components ofXandY. Simulation and real-world experimental results have validated the superiority of the proposed pose graph optimization model, under significant robot-vision precision disparities. Moreover, even when the error magnitudes between the robot and visual measurement system are comparable, the proposed method achieves superior accuracy for both the hand-eye and robot-world parameters compared to nine state-of-the-art approaches. Huayan Pu, Gang Wang 0023, Dengyu Xiao, Song Shao, Jun Luo 0006 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Bio-inspired Shape Self-Assembly in Large-Scale Swarm Robots Under Information Asymmetry *abstractThis study investigates the problem of large-scale swarm robots shape self-assembly problem under conditions of information asymmetry. Existing methods assume complete sharing of global information; however, this assumption has significant limitations in terms of resource consumption and swarm emergence. On the other hand, strategies that rely entirely on local information struggling to achieve the self-assembly of complex shapes, especially disconnected shapes. To address these challenges, this study proposes a novel bio-inspired distributed self-assembly strategy specifically designed for information asymmetric swarm. The strategy draws inspiration from the task specialization mechanism between scout ants and worker ants in social insects, guiding individuals to efficiently complete shape-assembly under information asymmetry through local perception, neighborhood interactions, and dynamic rule adjustments. Experimental results show that the proposed strategy successfully achieves the self-assembly of various shapes, including simple shapes (e.g., circle, rectangle), complex shapes (e.g., human, flower, and letter "A"), and disconnected shapes (e.g., letter "IO"). This demonstrates the strategy’s adaptability to shape complexity. Furthermore, experiments with varying swarm sizes validate the strategy’s robustness and scalability across different scales. During the experiments, we unexpectedly observed emergent behaviors within the swarm, further confirming that the proposed strategy not only significantly enhances task flexibility but also strengthens swarm emergence. These results indicate that the proposed method provides an efficient, scalable, and innovative solution for swarm robots self-assembly under information asymmetry. Dengyu Xiao, Gang Wang 0023, Huayan Pu, Jun Luo 0006 |
IROS | 3 |
| 2025 | FIPNet: Self-supervised low-light image enhancement combining feature and illumination priors
Qijie Zou, Xin Ning 0001, Gang Wang 0023 |
Neurocomputing | 5 |
| 2025 | RESEARCH NOTES: Multiclass Classification for Self-Admitted Technical Debt via Large Pre-Trained Language ModelabstractTechnical debt refers to suboptimal solutions adopted for short-term goals. Self-admitted technical debt (SATD) is the debt that is explicitly marked through comments or documentation, making it traceable. Multi-classification of SATD helps developers understand different debt types and improve efficiency. This paper proposes a SATD multi-classification method based on Fine-Tuning the GPT-3.5-turbo model for SATD prediction. This study uses a public dataset containing 10 projects with code comments. We classify design debt, requirement debt, and defect debt and evaluate our method’s performance. The experimental results show that compared to the best baseline model, our method achieves average improvements of 11.41%, 1.72% and 3.72% in MacroF, MacroP and MacroR metrics, respectively, in the MTO scenario. In the OTO scenario, improvements are 2.33%, 3.70% and 2.18%, respectively. These results indicate that our method has a strong generalization ability in SATD multi-classification and offers a new approach to managing technical debt. Yiyang Du, Xingguang Yang, Zhenyu Shu, Zijie Huang 0001, Gang Wang 0023, Libo Xu |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2025 | AR-MANet: A Low-Quality Image Restoration Method Based on Multi-Feature FusionabstractExposure problems often affect image quality more significantly than noise and blurring, seriously affecting their effectiveness in computer vision applications. To address this core problem, this paper proposes an innovative deep-learning method. The method uses pyramidal multi-scale decomposition to repair the image layer by layer, handling both overexposure and underexposure problems. Firstly, we introduce the Strengthen-Operate-Subtract (SOS) feature enhancement module, which refines and enhances the image by building previously estimated images to make the exposure correction effect more pronounced. Additionally, we propose using the Multi-scale Aggregation Back-projection (MAB) module to improve the U-Net model. Through the error feedback mechanism, the features of non-adjacent layers are compensated and integrated effectively, and the loss of detailed information is reduced. We introduce a dense block of matte residuals at the maximum scale to expand the network's receptive field and effectively capture local information. These innovations significantly enhance image exposure correction, reducing detail loss, regional contrast saturation, and color imbalance. Extensive evaluations confirm the effectiveness of our model, which demonstrates superior performance across all datasets, particularly in PSNR, achieving an average value of 26.271. Yugui Zhang, Hong Qin 0007, Gang Wang 0023 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Dual Guidance Enabled Fuzzy Inference for Enhanced Fine-Grained RecognitionabstractIn the field of fine-grained visual recognition (FGVR), the ability to resolve minute and often subtle differences between highly similar object categories is paramount. The advent of vision transformers (ViTs) has marked a significant advancement in this domain, primarily due to their capacity to model the intricate interdependencies among object parts represented as image patches. However, their inherent single-scale processing limitation hampers their effectiveness in FGVR tasks. Furthermore, the challenge of uncertainty inherent in FGVR tasks remains unresolved, necessitating the development of methods that bolster the robustness of these models, particularly across varying scales of visual features. We introduce a new plug-in module that can be seamlessly integrated into ViT, called dual guidance enabled fuzzy inference (DGEFI), which combines fuzzy inference with dual guidance mechanisms. Dual guidance includes scale-aware guidance and probability guidance. The former strengthens the model's focus on salient scales, and the latter refines the distinction between similar categories by optimizing intraclass compactness and interclass separability. Fuzzy inference enables the model to adaptively tweak the influence of distinct scales in the final decision-making phase, thereby enhancing the overall accuracy of recognition tasks. We demonstrate the versatility and efficacy of our DGEFI module by integrating it into several leading ViT backbones, including ViT, Swin, Mvitv2, and EVA-02. Empirical results exhibit exceptional performance gains, with the integration of DGEFI into EVA-02 remarkable accuracy improvements, reaching 93.6% on the CUB-200-2011 dataset and 94.5% on the NA-Birds dataset, respectively, improving over the state-of-the-art method 0.5% and 1.5%. Qiupu Chen, Feng He 0008, Gang Wang 0023, Xiao Bai 0001, Long Cheng 0003, Xin Ning 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | Brain-Inspired Fast- and Slow-Update Prompt Tuning for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) aims to learn new classes incrementally with a limited number of samples per class. Foundation models combined with prompt tuning showcase robust generalization and zero-shot learning (ZSL) capabilities, endowing them with potential advantages in transfer capabilities for FSCIL. However, existing prompt tuning methods excel in optimizing for stationary datasets, diverging from the inherent sequential nature in the FSCIL paradigm. To address this issue, taking inspiration from the "fast and slow mechanism" of the complementary learning systems (CLSs) in the brain, we present fast- and slow-update prompt tuning FSCIL (FSPT-FSCIL), a brain-inspired prompt tuning method for transferring foundation models to the FSCIL task. We categorize the prompts into two groups: fast-update prompts and slow-update prompts, which are interactively trained through meta-learning. Fast-update prompts aim to learn new knowledge within a limited number of iterations, while slow-update prompts serve as meta-knowledge and aim to strike a balance between rapid learning and avoiding catastrophic forgetting. Through experiments on multiple benchmark tests, we demonstrate the effectiveness and superiority of FSPT-FSCIL. The code is available at https://github.com/qihangran/FSPT-FSCIL. Hang Ran, Xingyu Gao 0001, Lusi Li, Weijun Li 0002, Songsong Tian, Gang Wang 0023, Hailong Shi, Xin Ning 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | A Framework for Real-time Generation of Multi-directional Traversability Maps in Unstructured EnvironmentsabstractIn complex unstructured environments, accurate terrain traversability analysis is a fundamental requirement for the successful execution of any movements of ground robots, especially given that terrain traversability often exhibits anisotropy. However, the difficulty in obtaining multi-directional terrain labels hinders the emergence of end-to-end multi-directional traversability network. This paper introduces a framework for real-time multi-directional traversability maps (MTraMap) generation tailored for unstructured environments. It involves pre-training a uni-directional traversability classifier, termed UniTraT, through self-supervised learning using ground robot travel simulation. Furthermore, it employs Uni-directional to Multi-directional Traversability Distillation (UMTraDistill) to distill a multi-directional traversability network, termed MultiTCNN, which is capable of directly generating MTraMap. We evaluated both networks on our traversability dataset, achieving an 89% accuracy in terrain traversability classification with the UniTraT. Compared to UniTraT, the accuracy of the MultiTCNN distilled via UMTraDistill only decreases by 1.8%, and it can process 10 m × 10 m elevation map at a speed of 74 fps. Field robotics experiments were also conducted and showed that MultiTCNN can generate MTraMap of the surrounding 20 m × 20 m environment at a rate of 9.39 fps, with a slight reduction of 0.61 fps compared to the lidar data publishing rate, and the generated MTraMap can clearly delineate the multi-directional traversability of the surrounding environments. Tao Huang 0010, Gang Wang 0023, Tao Zhu 0003, Huayan Pu, Jun Luo 0006 |
ICRA | 2 |
| 2024 | BE-SLAM: BEV-Enhanced Dynamic Semantic SLAM with Static Object ReconstructionabstractThe quality of a robot’s environmental perception determines whether it can achieve more intelligent applications, such as semantic interaction with humans. SLAM, on the other hand, is one of the crucial capabilities for a robot to perceive its environment. However, when only a monocular image is provided, dynamic objects in the environment significantly impact the accuracy of map construction by the robot, leading to erroneous perception results. To address this issue, we propose a Visual SLAM framework based on BEV perception results, named BE-SLAM. With this framework, we can handle dynamic objects, occlusions, and incompletely observed objects. It can construct a stable static map by strengthening trust in static objects. Considering that object-level semantic maps can enhance a robot’s perception abilities, we also reconstruct static objects in the map and use them to optimize the pose. Through experiments on existing publicly available dataset, we compare BE-SLAM with several existing methods that have shown good performance. The experimental results demonstrate that BE-SLAM performs exceptionally well on high-dynamic sequences and achieves comparable results on static or low-dynamic sequences. Jun Luo 0003, Gang Wang 0023, Tao Huang 0010, Dengyu Xiao, Huayan Pu, Jun Luo 0006 |
IROS | 2 |
| 2024 | Leveraging Contrastive Language-Image Pre-Training and Bidirectional Cross-attention for Multimodal Keyword Spotting
Dong Liu 0037, Qirong Mao, Lijian Gao, Gang Wang 0023 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Fast 3D Object Measurement Based on Point Cloud ModelingabstractAutomated object measurement is becoming increasingly important due to its ability to reduce manual costs, increase production efficiency, and minimize errors in various fields. In this paper, we present a novel approach to three-dimensional (3D) object measurement based on point cloud modeling. Our method introduces a fast point cloud modeling computation framework consisting of five stages: coordinate centralization, rotation and translation, noise filtering, plane projection, and geometric computation. Furthermore, we propose a fast convex hull optimization algorithm to reduce the high complexity problem of traditional convex hull calculation. Our extensive experiments demonstrate that our approach outperforms existing methods in terms of measurement error rate and time savings, with a maximum time saving of 31.03% under certain error conditions. Gang Wang 0023, Mingliang Zhou 0001, Bin Fang 0001, Yugui Zhang, Shouqin Guan, Bin Ruan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2023 | Reconstruction of Smooth Skin Surface Based on Arbitrary Distributed Sparse Point CloudsabstractAdaptively reconstruction of the skin surface covering the skeletons of aircrafts, ships, high-speed trains, buildings with steel frames, etc., based on on-site measured point clouds is an important issue in both industry and construction. However, the skeleton has the characteristics of long and narrow shape and variable structure, resulting in a narrow and sparse distribution of measured point cloud on its surface with variable shape in different areas, which brings great challenges to the reconstruction of skin surface. In this article, a method for accurately reconstructing the skin surface covering the skeleton structures based on the arbitrary distributed sparse on-site measured points of the skeleton surface is proposed. At first, a robust B-spline curve fitting method based on the least square principle is proposed to construct the boundary curves of the point cloud. Then, a Coons-B-spline surface fitting method based on the generated boundary curves is proposed to generate an initial skin surface. Next, the initial skin surface is considered as a curved thin plate with stiffness, and a method of surface deformation considering the target points of deformation and the tensile and shear stiffness of the surface is proposed to obtain the surface with high accuracy and good smoothness. To show the feasibility of the proposed method, simulations and experiments are carried out. It is proved that the proposed method can achieve better accuracy while maintaining smoothness. Gang Wang 0023, Jun Luo 0003, Lisheng Mou, Huayan Pu, Jun Luo 0006 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Vehicular Abandoned Object Detection Based on VANET and Edge AI in Road ScenesabstractRapid processing of abandoned objects is one of the most important tasks in road maintenance. Abandoned object detection heavily relies on traditional object detection approaches at a fixed location. However, detection accuracy and range are still far from satisfactory. This study proposes an abandoned object detection approach based on vehicular ad-hoc networks (VANETs) and edge artificial intelligence (AI) in road scenes. We propose a vehicular detection architecture for abandoned objects to achieve task-based AI technology for large-scale road maintenance in mobile computing circumstances. To improve detection accuracy and reduce repeated detection rates in mobile computing, we propose a detection algorithm that combines a deep learning network and a deduplication module for high-frequency detection. Finally, we propose a location estimation approach for abandoned objects based on the World Geodetic System 1984 (WGS84) coordinate system and an affine projection model to accurately compute the positions of abandoned objects. Experimental results show that our proposed algorithm achieves an average accuracy of 99.57% and 53.11% on the two datasets, respectively. Additionally, our whole system achieves real-time detection and high-precision localization performance on real roads. Gang Wang 0023, Mingliang Zhou 0001, Xuekai Wei, Guang Yang 0006 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | An Accelerated and Flexible SIFT Parallel-Computing Approach Based on the General Multi-Core PlatformabstractVisual retrieval has been a significant technology in the computer vision task. Visual feature descriptors are the key to the visual retrieval. The famous local feature descriptor is called the Scale Invariant Feature Transform (SIFT), which can keep invariant mapping for the scale, rotate and simulate images. To utilize effectively the SIFT feature descriptor for visual matching on different hardware platforms, this paper proposes an accelerated SIFT algorithm based on the SIFT feature computing principle of the general multi-core platform. First, our multi-core task allocation method introduces the WFM theory into task assignment for each core to improve the core computing resource utilization for high-efficient parallel computing. Then, to improve the efficiency of picture matching, we introduce global geometric constraints condition to optimal picture matching for the multi-core parallelization approach. Experimental results show that the proposed approach can save on average 87.31% on the Intel X86 platform, compared to the single-core time. Also, our approach can save on average 33.79% on the Raspberry Pi platform, compared to the single-core time. Gang Wang 0023, Mingliang Zhou 0001, Bin Fang 0001, Haichao Huang, Zhenyu Shu, Xueshu Chen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2022 | An active contour model based on adaptively variable exponent combining Legendre polynomial for image segmentation
Jiajie Zhu 0001, Bin Fang 0001, Mingliang Zhou 0001, Futing Luo, Weizhi Xian, Gang Wang 0023 |
Multim. Tools Appl. | 6 |
| 2022 | Trajectory Planning and Optimization for Robotic Machining Based On Measured Point CloudabstractIndustrial robots are characterized by good flexibility and a large working space, and offer a new approach for the machining of large and complex parts with small machining allowances (extra material allowed for subsequent machining). Parts of this type (such as aircraft skin parts, wind turbine blades, etc.) are easily deformed due to their large scale and low stiffness. Therefore, these parts cannot be directly machined according to the designed model. A feasible method is to plan a robotic machining path by using the point clouds of parts after clamping from onsite measurement which contains inherent defects of measurement such as noise points and abrupt points. In this article, a novel method is proposed to plan and optimize a robotic machining path that meets the requirements of smoothness, dexterity, and stiffness based on the point cloud from onsite measurement. The dual nonuniform rational B-spline curves of the machining path points and tool axis points are generated at first. Next, an objective function of smoothness optimization is established to filter out the local mutation of the path by considering the constraints of both the deformation energy and the deviation. Then, the objective function of robot postures optimization is established to optimize dexterity and Cartesian stiffness of a robot during the machining process. To show the feasibility of the proposed method, simulation and experiments are carried out. It is proved that the proposed method can generate a smooth machining trajectory. The stability of joint rotation and the rigidity and dexterity of the robot are improved during the machining process. Gang Wang 0023, Wenlong Li 0001, Cheng Jiang 0007, Dahu Zhu, Wei Xu 0027, Huan Zhao 0001, Han Ding 0001 |
IEEE Trans. Robotics | 1 |
| 2021 | Deep Sensor Fusion Based on Frustum Point Single Shot Multibox Detector for 3D Object DetectionabstractWe present a deep sensor fusion method based on frustum point single shot multibox detector (PointSSD) for autonomous driving scenarios. The proposed method solves the problem of precision degradation in frustum PointNets (F-PointNet) caused by relying heavily on 2D detection and making insufficient use of RGB information. The method mainly consists of two subnetworks: pyramid segmentation network (PSNet) and PointSSD. The proposed PSNet uses a novel architecture capable of performing semantic segmentation on RGB information to generate high quality image semantic information. Using these image semantic information, point cloud semantic information is obtained through projection and is then fused with raw 3D spatial features by deep fusion. The fusion results are processed by PointSSD, which is proposed for classification and bounding box regression. Evaluated on the KITTI dataset, our method is superior to other methods in 3D classification and 3D localization. In addition, our method guarantees robustness to 2D false detections. Ye Zhang 0008, Shaohua Zhai, Hao Chen 0014, Shaoqi Shi, Gang Wang 0023 |
ICIP | 6 |
| 2021 | A Fast Perceptual Surveillance Video Coding (PSVC) Based on Background Model-Driven JND EstimationabstractPerceptual video coding (PVC) optimization has been an important video coding technique, which can be consistent with the perception characteristics of the human visual system (HVS). Currently, PVC schemes incorporating the just noticeable distortion (JND) model can obtain better performance gain in all PVC schemes. To further accelerate the JND computation for real-time video coding applications (e.g. surveillance video coding and conference video coding), this paper proposes a fast perceptual surveillance video coding (PSVC) scheme based on background model-driven JND estimation method. First, to utilize the surveillance scene characteristics, the computation complexity of JND estimation can be significantly decreased by reusing the content complexity of background regions. Then we apply the perceptive video coding scheme into the background modeling-based surveillance video codec. The proposed scheme adopts background modeling frame as background anchor. Experimental results show that the proposed scheme can yield remarkable time saving of 42.33% maximum and on average 34.76% with approximate bitrate reductions and similar subjective quality, compared to HEVC and other state-of-the-art schemes. Gang Wang 0023, Mingliang Zhou 0001, Haiheng Cao, Bin Fang 0001, Shiting Wen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2021 | Surveillance Video Coding for Traffic Scene based on Vehicle Knowledge and Shared Library by Cloud-Edge Computing in Cyber-Physical-Social SystemsabstractWith rapid development of intelligent video surveillance systems based on cloud computing devices and edge computing devices in Cyber-Physical-Social Systems, massive surveillance video data has brings enormous challenge for video storage and transmission. However, existing surveillance video coding approaches hardly utilize intelligent video analysis results for improving video coding. This paper proposed a surveillance video coding scheme for traffic scene based on vehicle knowledge and shared library by cloud-edge computing in Cyber-Physical-Social Systems. Firstly, in order to provide the object library for synchronous application at the encode and decode side offline, a generation method of shared long-term foreground reference object library is proposed by using the existing large-scale monitoring vehicle object datasets. Then, to meet the requirement of low complexity and high-performance coding, a virtual foreground reference picture generation method with coding-oriented object retrieval is proposed. Experimental results show that the proposed scheme can obtain the satisfactory effect of the virtual foreground reference picture. Also, it can yield remarkable bit rate reductions, compared to HEVC. Gang Wang 0023, Mingliang Zhou 0001, Bin Fang 0001, Shiting Wen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2021 | Simultaneous Calibration of Multicoordinates for a Dual-Robot System by Solving the AXB = YCZ ProblemabstractMultirobot systems have shown great potential in dealing with complicated tasks that are impossible for a single robot to achieve. One essential problem encountered in cooperatively working of the multirobot systems is the unknown initial transformation relationships from hand to eye, base to base, and flange to tool. In this article, the problem of multicoordinates calibration for a dual-robot system is formulated to a matrix equation AXB = YCZ. A novel approach for simultaneously solving the unknowns in equation AXB = YCZ is proposed, which is composed of a closed form method based on the Kronecker product and an iterative method which converts the calculation of a nonlinear problem to an optimization problem of a strictly convex function. The closed form method is used to quickly obtain an initial estimation for the iterative method to improve the efficiency and accuracy of iteration. In addition, a series of conditions on the solvability of the problem are proposed to guide the operators to select appropriate robot attitudes during the calibration process. To show the feasibility and superiority of the proposed iterative method, two other calibration methods are chosen to be compared to the proposed method through simulation and practical experiments. The comparison results verify the superiority of the proposed method in accuracy, efficiency, and stability. Gang Wang 0023, Wenlong Li 0001, Cheng Jiang 0007, Dahu Zhu, He Xie, Xingjian Liu, Han Ding 0001 |
IEEE Trans. Robotics | 1 |
| 2019 | Texture-Classification Accelerated CNN Scheme for Fast Intra CU Partition in HEVCabstractHigh Efficiency Video Coding (HEVC) achieves significant coding performance over H.264. However, the performance gain is achieved at the cost of substantially higher encoding complexity, in which the coding tree unit (CTU) partition is one of the most time-consuming parts due to the rate-distortion optimization-based ergodic search of all possible quad-tree partitions. To address this problem, this paper proposes a texture-classification accelerated convolutional neural network (CNN)-based fast intra CU partition scheme to reduce the encoding complexity for intra-coding in HEVC, by taking into consideration of the heterogeneous texture characteristics into the CNN-based classification. First, a threshold-based texture classification model is developed to identify the heterogeneous and homogeneous CTUs, through jointly consideration of the CU depth, quantization parameter and texture complexity. Second, three different CNN structures are designed and trained to predict the CU partition mode for each CU layer in the heterogeneous CTUs. Finally, extensive experimental results show that the proposed scheme can reduce intra-mode encoding time by 62.13% with negligible BD-rate loss of 2.01%, consistently outperforming two state-of-the-art CNN-based schemes in terms of both coding performance and complexity reduction. Yongfei Zhang, Gang Wang 0023, Mai Xu, C.-C. Jay Kuo |
DCC | 2 |
| 2019 | Multi-View Frustum Pointnet for Object Detection in Autonomous DrivingabstractLIDAR point cloud and RGB images are often used for object detection in autonomous driving scenarios. This paper develops a multi-view version of Frustum PointNet (F-PointNet), to be called MVFP to reduce the rate of missed detection in F-PointNet by adding auxiliary bird's eye view (BEV) detection part. In processing MVFP, initial object detection results are obtained from F-PointNet by combining the RGB image and raw LIDAR point cloud. Simultaneously, raw LIDAR point cloud is encoded into BEV feature maps, from which 2D bounding boxes are predicted. In missed detection judgement, the intersection over union (IoU) is used as a criteria for the matching of preliminary object detection results from F-PointNet and BEV maps prediction results. 2D boxes belonging to missed-detected objects from BEV maps are projected to the pipeline of F-PointNet until all the objects in BEV maps find a matching detection result in the set of F-PointNet detection results. To evaluate the performance of MVFP, 3D object detection experiments are conducted on KITTI benchmark. The experiment results demonstrate that MVFP outperforms the original F-PointNet by 5% and 4% higher recall on the hard mode of pedestrian and cyclist. Hao Chen 0014, Ye Zhang 0008, Gang Wang 0023 |
ICIP | 4 |
| 2019 | Highly Paralleled Low-Cost Embedded HEVC Video Encoder on TI KeyStone Multicore DSPabstractAlthough HEVC, the emerging video coding standard, has doubled the coding performance of its predecessor H.264/AVC, its significantly increased computational complexity imposes great obstacles for HEVC encoders to be employed in real-time applications with embedded processors, such as digital signal processors (DSPs). In this paper, a TI Keystone multicore TMS320C6678 DSP-based highly paralleled low-cost fast HEVC encoding solution is well designed and implemented. First, the overall structure of HEVC encoder with CTU-level parallelism is re-designed to well support the encoding parallelism, with full consideration of the hardware characteristics. Second, a low-delay and low-memory multicore data transmission mechanism is proposed to reduce the latency of data access between internal L2 memory and external DDR3. Third, the encoding bottlenecks, i.e., the most time-consuming encoding modules, are identified and optimized for acceleration with TI powerful C6000 SIMD instructions. Experimental results show that our proposed HEVC encoder on TI TMS320C6678 DSPs can significantly improve the real-time capacity with tolerable performance loss, 0.93 dB performance loss under on average 465.50 times speedup as compared to CPU-based HM reference software, more specifically, which makes it desirable in power-constrained real-time video applications. Rui Fan 0002, Yongfei Zhang, Gang Wang 0023, Zhe Li 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | A fast and HEVC-compatible perceptual video coding scheme using a transform-domain Multi-Channel JND model
Gang Wang 0023, Yongfei Zhang, Bo Li 0006, Rui Fan 0002, Mingliang Zhou 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Background Modeling and Referencing for Moving Cameras-Captured Surveillance Video Coding in HEVCabstractSurveillance video coding is crucial for improving compression efficiency in intelligent video surveillance systems and applications. Plenty of work has been done, which can be roughly divided into two categories: the former mainly focuses on low-complexity background modeling to obtain the clear background, while the latter focuses on an appropriate coding strategy to generate the high-quality background reference picture for effective background prediction. However, almost all existing works focus only on stationary camera scenes, while moving cameras-captured surveillance video coding is left untouched and is still an open problem. In this paper, a background modeling and referencing scheme for moving cameras-captured surveillance video coding in high-efficiency video coding (HEVC) is proposed. First, this paper proposes a low-complexity motion background modeling algorithm for surveillance video coding using the running average based on a global-motion-compensation method. To obtain the global motion vector, we propose a global motion detection method based on character blocks by establishing a low-rank singular value decomposition model for clustering and estimating motion vectors of background character blocks in the cameras movement circumstance. Second, we propose a background referencing coding strategy, in which the motion background coding tree units (MBCTUs) would be selected by anchoring the input video frame on the modeling background frame and coded with the optimized quantization parameter. Then, the reconstructed MBCTU will be used to update the previous coding tree unit in the global compensation location of the background reference picture. Extensive experimental results show that the proposed scheme can achieve significant bit savings of up to 26.6% and, on average, 6.7% with similar subjective quality and negligible encoding complexity, compared to HM12.0. Besides, the proposed scheme consistently outperforms two state-of-the-art surveillance video coding schemes with remarkable bitrate savings. Gang Wang 0023, Bo Li 0006, Yongfei Zhang, Jinhui Yang |
IEEE Trans. Multim. | 1 |
| 2017 | Multidirectional parabolic prediction-based interpolation-free sub-pixel motion estimation
Rui Fan 0002, Yongfei Zhang, Bo Li 0006, Gang Wang 0023 |
Signal Process. Image Commun. | 4 |