Gaofeng Meng

dblp:78/6915 · DBLP profile ↗
← Back
99ranked-venue papers
13as first author
45since 2021 · last 2026
0000-0002-7103-6321ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 11 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 6 first-author · 29 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021
YearPublicationVenuePosition
2026 Practical Continual Forgetting for Pre-Trained Vision Models
abstract
For privacy and security concerns, the need to erase unwanted information from pre-trained vision models is becoming evident nowadays. In real-world scenarios, erasure requests originate at any time from both users and model owners, and these requests usually form a sequence. Therefore, under such a setting, selective information is expected to be continuously removed from a pre-trained model while maintaining the rest. We define this problem as continual forgetting and identify three key challenges. (i) For unwanted knowledge, efficient and effective deleting is crucial. (ii) For remaining knowledge, the impact brought by the forgetting procedure should be minimal. (iii) In real-world scenarios, the training samples may be scarce or partially missing during the process of forgetting. To address them, we first propose Group Sparse LoRA (GS-LoRA). Specifically, towards (i), we introduce Low-Rank Adaptation (LoRA) modules to fine-tune the Feed-Forward Network (FFN) layers in Transformer blocks for each forgetting task independently, and towards (ii), a simple group sparse regularization is adopted, enabling automatic selection of specific LoRA groups and zeroing out the others. To further extend GS-LoRA to more practical scenarios, we incorporate prototype information as additional supervision and introduce a more practical approach, GS-LoRA++. For each forgotten class, we move the logits away from its original prototype. For the remaining classes, we pull the logits closer to their respective prototypes. We conduct extensive experiments on face recognition, object detection and image classification and demonstrate that our method manages to forget specific classes with minimal impact on other classes.
Hongbo Zhao 0006, Fei Zhu 0004, Bolin Ni, Gaofeng Meng, Zhaoxiang Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Pareto Continual Learning: Preference-Conditioned Learning and Adaption for Dynamic Stability-Plasticity Trade-off
abstract
Continual learning aims to learn multiple tasks sequentially. A key challenge in continual learning is balancing between two objectives: retaining knowledge from old tasks (stability) and adapting to new tasks (plasticity). Experience replay methods, which store and replay past data alongside new data, have become a widely adopted approach to mitigate catastrophic forgetting. However, these methods neglect the dynamic nature of the stability-plasticity trade-off and aim to find a fixed and unchanging balance, resulting in suboptimal adaptation during training and inference. In this paper, we propose Pareto Continual Learning (ParetoCL), a novel framework that reformulates the stability-plasticity trade-off in continual learning as a multi-objective optimization (MOO) problem. ParetoCL introduces a preference-conditioned model to efficiently learn a set of Pareto optimal solutions representing different trade-offs and enables dynamic adaptation during inference. From a generalization perspective, ParetoCL can be seen as an objective augmentation approach that learns from different objective combinations of stability and plasticity. Extensive experiments across multiple datasets and settings demonstrate that ParetoCL outperforms state-of-the-art methods and adapts to diverse continual learning scenarios.
Song Lai 0001, Zhe Zhao 0008, Fei Zhu 0004, Xi Lin 0001, Qingfu Zhang 0001, Gaofeng Meng
AAAI6
2025 Occlusion-aware Non-Rigid Point Cloud Registration via Unsupervised Neural Deformation Correntropy
abstract
Non-rigid alignment of point clouds is crucial for scene understanding, reconstruction, and various computer vision and robotics tasks. Recent advancements in implicit deformation networks for non-rigid registration have significantly reduced the reliance on large amounts of annotated training data. However, existing state-of-the-art methods still face challenges in handling occlusion scenarios. To address this issue, this paper introduces an innovative unsupervised method called Occlusion-Aware Registration (OAR) for non-rigidly aligning point clouds. The key innovation of our method lies in the utilization of the adaptive correntropy function as a localized similarity measure, enabling us to treat individual points distinctly. In contrast to previous approaches that solely minimize overall deviations between two shapes, we combine unsupervised implicit neural representations with the maximum correntropy criterion to optimize the deformation of unoccluded regions. This effectively avoids collapsed, tearing, and other physically implausible results. Moreover, we present a theoretical analysis and establish the relationship between the maximum correntropy criterion and the commonly used Chamfer distance, highlighting that the correntropy-induced metric can be served as a more universal measure for point cloud analysis. Additionally, we introduce locally linear reconstruction to ensure that regions lacking correspondences between shapes still undergo physically natural deformations. Our method achieves superior or competitive performance compared to existing approaches, particularly when dealing with occluded geometries. We also demonstrate the versatility of our method in challenging tasks such as large deformations, shape interpolation, and shape completion under occlusion disturbances.
Mingyang Zhao 0001, Gaofeng Meng, Dong-Ming Yan 0001
ICLR2
2025 Agent Reviewers: Domain-specific Multimodal Agents with Shared Memory for Paper Review
abstract
Feedback from peer review is essential to improve the quality of scientific articles. However, at present, many manuscripts do not receive sufficient external feedback for refinement before or during submission. Therefore, a system capable of providing detailed and professional feedback is crucial for enhancing research efficiency. In this paper, we have compiled the largest dataset of paper reviews to date by collecting historical open-access papers and their corresponding review comments and standardizing them using LLM. We then developed a multi-agent system that mimics real human review processes, based on LLMs. This system, named Agent Reviewers, includes the innovative introduction of multimodal reviewers to provide feedback on the visual elements of papers. Additionally, a shared memory pool that stores historical papers' metadata is preserved, which supplies reviewer agents with background knowledge from different fields. Our system is evaluated using ICLR 2024 papers and achieves superior performance compared to existing AI-based review systems. Comprehensive ablation studies further demonstrate the effectiveness of each module and agent in this system.
Shixiong Xu, Jinqiu Li, Kun Ding 0001, Gaofeng Meng
ICML5
2025 F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery
Lumin Chen, Zhiying Wu, Tianye Lei, Xuexue Bai, Ming Feng, Yuxi Wang 0001, Gaofeng Meng, Zhen Lei 0001, Hongbin Liu 0001
MICCAI (9)7
2025 MedICL: In-Context Learning for Semantically Enhanced AKI Prediction in Cardiac Surgery
Chenyang Su, Yishun Wang, Boqiang Xu, Rong Feng, Hongbin Liu 0001, Gaofeng Meng
MICCAI (11)7
2025 Gradient-Guided Epsilon Constraint Method for Online Continual Learning
abstract
Online Continual Learning (OCL) requires models to learn sequentially from data streams with limited memory. Rehearsal-based methods, particularly Experience Replay (ER), are commonly used in OCL scenarios. This paper revisits ER through the lens of $\epsilon$-constraint optimization, revealing that ER implicitly employs a soft constraint on past task performance, with its weighting parameter post-hoc defining a slack variable. While effective, ER's implicit and fixed slack strategy has limitations: it can inadvertently lead to updates that negatively impact generalization, and its fixed trade-off between plasticity and stability may not optimally balance current streaming with memory retention, potentially overfitting to the memory buffer. To address these shortcomings, we propose the \textbf{G}radient-Guided \textbf{E}psilon \textbf{C}onstraint (\textbf{GEC}) method for online continual learning. GEC explicitly formulates the OCL update as an $\epsilon$-constraint optimization problem, which minimize the loss on the current task data and transform the stability objective as constraints and propose a gradient-guided method to dynamically adjusts the update direction based on whether the performance on memory samples violates a predefined slack tolerance $\bar{\varepsilon}$: if forgetting exceeds this tolerance, GEC prioritizes constraint satisfaction; otherwise, it focuses on the current task while controlling the rate of increase in memory loss. Empirical evaluations on standard OCL benchmarks demonstrate GEC's ability to achieve a superior trade-off, leading to improved overall performance. Code is available at https://github.com/laisong-22004009/GEC_OCL.
Song Lai 0001, Changyi Ma, Fei Zhu 0004, Zhe Zhao 0008, Xi Lin 0001, Gaofeng Meng, Qingfu Zhang 0001
NeurIPS6
2025 Practical incremental learning: Striving for better performance-efficiency trade-off
Shixiong Xu, Bolin Ni, Xing Nie, Fei Zhu 0004, Jianlong Chang, Gaofeng Meng
Neurocomputing7
2024 Defying Imbalanced Forgetting in Class Incremental Learning
abstract
We observe a high level of imbalance in the accuracy of different learned classes in the same old task for the first time. This intriguing phenomenon, discovered in replay-based Class Incremental Learning (CIL), highlights the imbalanced forgetting of learned classes, as their accuracy is similar before the occurrence of catastrophic forgetting. This discovery remains previously unidentified due to the reliance on average incremental accuracy as the measurement for CIL, which assumes that the accuracy of classes within the same task is similar. However, this assumption is invalid in the face of catastrophic forgetting. Further empirical studies indicate that this imbalanced forgetting is caused by conflicts in representation between semantically similar old and new classes. These conflicts are rooted in the data imbalance present in replay-based CIL methods. Building on these insights, we propose CLass-Aware Disentanglement (CLAD) as a means to predict the old classes that are more likely to be forgotten and enhance their accuracy. Importantly, CLAD can be seamlessly integrated into existing CIL methods. Extensive experiments demonstrate that CLAD consistently improves current replay-based methods, resulting in performance gains of up to 2.56%.
Shixiong Xu, Gaofeng Meng, Xing Nie, Bolin Ni, Bin Fan 0001, Shiming Xiang
AAAI2
2024 Continual Forgetting for Pre-Trained Vision Models
abstract
For privacy and security concerns, the need to erase un-wanted information from pre-trained vision models is becoming evident nowadays. In real-world scenar-ios, erasure requests originate at any time from both users and model owners. These requests usually form a sequence. Therefore, under such a setting, selective information is expected to be continuously removed from a pre-trained model while maintaining the rest. We define this problem as continual forgetting and identify two key challenges. (i) For unwanted knowledge, efficient and effective deleting is crucial. (ii) For remaining knowledge, the impact brought by the forgetting procedure should be minimal. To address them, we propose Group Sparse LoRA (GS-LoRA). Specifically, towards (i), we use LoRA modules to fine-tune the FFN layers in Transformer blocks for each forgetting task independently, and towards (ii), a simple group sparse regularization is adopted, enabling automatic selection of specific LoRA groups and zeroing out the others. GS-LoRA is effective, parameter-efficient, data-efficient, and easy to implement. We conduct extensive experiments on face recognition, object detection and image classification and demonstrate that GS-LoRA manages to forget specific classes with minimal impact on other classes. Codes will be released on https://github.com/bjzhb666/GS-LoRA.
Hongbo Zhao 0006, Bolin Ni, Junsong Fan, Yuxi Wang 0001, Yuntao Chen, Gaofeng Meng, Zhaoxiang Zhang 0001
CVPR6
2024 Enhancing Visual Continual Learning with Language-Guided Supervision
abstract
Continual learning (CL) aims to empower models to learn new tasks without forgetting previously acquired knowledge. Most prior works concentrate on the techniques of architectures, replay data, regularization, etc. However, the category name of each class is largely neglected. Existing methods commonly utilize the one-hot labels and randomly initialize the classifier head. We argue that the scarce semantic information conveyed by the one-hot labels hampers the effective knowledge transfer across tasks. In this paper, we revisit the role of the classifier head within the CL paradigm and replace the classifier with semantic knowledge from pretrained language models (PLMs). Specifically, we use PLMs to generate semantic targets for each class, which are frozen and serve as supervision signals during training. Such targets fully consider the semantic correlation between all classes across tasks. Empirical studies show that our approach mitigates forgetting by alleviating representation drifting and facilitating knowledge transfer across tasks. The proposed method is simple to implement and can seamlessly be plugged into existing methods with negligible adjustments. Extensive experiments based on eleven mainstream baselines demonstrate the effectiveness and generalizability of our approach to various protocols. For example, under the class-incremental learning setting on ImageNet-100, our method significantly improves the Top-1 accuracy by 3.2% to 6.1% while reducing the forgetting rate by 2.6% to 13.1%.
Bolin Ni, Hongbo Zhao 0006, Chenghao Zhang 0003, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang
CVPR5
2024 Correspondence-Free Non-Rigid Point Set Registration Using Unsupervised Clustering Analysis
abstract
This paper presents a novel non-rigid point set registration method that is inspired by unsupervised clustering analysis. Unlike previous approaches that treat the source and target point sets as separate entities, we develop a holistic framework where they are formulated as clustering centroids and clustering members, separately. We then adopt Tikhonov regularization with an$\ell_{1}$-induced Laplacian kernel instead of the commonly used Gaussian kernel to ensure smooth and more robust displacement fields. Our formulation delivers closed-form solutions, theoretical guarantees, independence from dimensions, and the ability to handle large deformations. Subsequently, we introduce a clustering-improved Nyström method to effectively reduce the computational complexity and storage of the Gram matrix to linear, while providing a rigorous bound for the low-rank approximation. Our method achieves high accuracy results across various scenarios and surpasses competitors by a significant margin, particularly on shapes with sub-stantial deformations. Additionally, we demonstrate the versatility of our method in challenging tasks such as shape transfer and medical registration. [Code release]
Mingyang Zhao 0001, Jingen Jiang 0001, Lei Ma 0008, Shi-Qing Xin, Gaofeng Meng, Dong-Ming Yan 0001
CVPR5
2024 AddressCLIP: Empowering Vision-Language Models for City-Wide Image Address Localization
Shixiong Xu, Chenghao Zhang 0003, Lubin Fan, Gaofeng Meng, Shiming Xiang, Jieping Ye
ECCV (28)4
2024 SFD: Similar Frame Dataset for Content-Based Video Retrieval
abstract
Content-based video retrieval aims to retrieve near-duplicate entries from a database of a given query video. It plays an important role in combating video piracy. Robustness to video temporal dynamics is crucial for a representation model in video retrieval, as frames extracted from two copied videos are hardly temporally aligned in actual situations. However, current image retrieval datasets have difficulty in evaluating this robustness. To address this issue, we collect Similar Frame Dataset (SFD), which consists of 32,923 query-target pairs with 128,240 distraction images. The task of SFD is to retrieve the target frame from all items given a query frame. SFD is constructed by sampling frames from Kinetics-700 action classification dataset. An object detection model (Faster R-CNN) and a Multimodal Large Language Model (BLIP2) are used during sampling to select those valid frames. Besides, we propose Adjacent Frames Contrastive Learning (AFCL) framework. In AFCL, adjacent frames are sampled from unlabeled videos as positive pairs. An image representation model with robustness to changing frames can be trained under AFCL framework and achieve the state-of-the-art performance on SFD. The code will be released at https://github.com/Chuan-shanjia/Similar-Frame-Dataset.
Chaowei Han, Gaofeng Meng, Chunlei Huo
ICIP2
2024 A Multimodal Transformer for Live Streaming Highlight Prediction
abstract
Recently, live streaming platforms have gained immense popularity. Traditional video highlight detection mainly focuses on visual features and utilizes both past and future content for prediction. However, live streaming requires models to infer without future frames and process complex multimodal interactions, including images, audio and text comments. To address these issues, we propose a multimodal transformer that incorporates historical look-back windows. We introduce a novel Modality Temporal Alignment Module to handle the temporal shift of cross-modal signals. Additionally, using existing datasets with limited manual annotations is insufficient for live streaming whose topics are constantly updated and changed. Therefore, we propose a novel Border-aware Pairwise Loss to learn from a large-scale dataset and utilize user implicit feedback as a weak supervision signal. Extensive experiments show our model outperforms various strong baselines on both real-world scenarios and public datasets. And we will release our dataset and code to better assess this topic.
Jiaxin Deng, Shiyao Wang 0001, Dong Shen 0003, Liqin Zhao, Fan Yang 0094, Guorui Zhou, Gaofeng Meng
ICME7
2024 MMBee: Live Streaming Gift-Sending Recommendations via Multi-Modal Fusion and Behaviour Expansion
abstract
Live streaming services are becoming increasingly popular due to real-time interactions and entertainment. Viewers can chat and send comments or virtual gifts to express their preferences for the streamers. Accurately modeling the gifting interaction not only enhances users' experience but also increases streamers' revenue. Previous studies on live streaming gifting prediction treat this task as a conventional recommendation problem, and model users' preferences using categorical data and observed historical behaviors. However, it is challenging to precisely describe the real-time content changes in live streaming using limited categorical information. Moreover, due to the sparsity of gifting behaviors, capturing the preferences and intentions of users is quite difficult. In this work, we propose MMBee based on real-time Multi-Modal Fusion and Behaviour Expansion to address these issues. Specifically, we first present a Multi-modal Fusion Module with Learnable Query (MFQ) to perceive the dynamic content of streaming segments and process complex multi-modal interactions, including images, text comments and speech. To alleviate the sparsity issue of gifting behaviors, we present a novel Graph-guided Interest Expansion (GIE) approach that learns both user and streamer representations on large-scale gifting graphs with multi-modal attributes. It consists of two main parts: graph node representations pre-training and metapath-based behavior expansion, all of which help model jump out of the specific historical gifting behaviors for exploration and largely enrich the behavior representations. Comprehensive experiment results show that MMBee achieves significant performance improvements on both public datasets and Kuaishou real-world streaming datasets and the effectiveness has been further validated through online A/B experiments. MMBee has been deployed and is serving hundreds of millions of users at Kuaishou.
Jiaxin Deng, Shiyao Wang 0001, Jiansong Qi, Liqin Zhao, Guorui Zhou, Gaofeng Meng
KDD7
2024 Force Sensing Guided Artery-Vein Segmentation via Sequential Ultrasound Images
Yimeng Geng, Gaofeng Meng, Mingcong Chen, Guanglin Cao, Mingyang Zhao 0001, Hongbin Liu 0001
MICCAI (4)2
2024 EchoMEN: Combating Data Imbalance in Ejection Fraction Regression via Multi-expert Network
Song Lai 0001, Mingyang Zhao 0001, Zhe Zhao 0008, Shi Chang, Xiaohua Yuan, Hongbin Liu 0001, Qingfu Zhang 0001, Gaofeng Meng
MICCAI (4)8
2024 OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction
abstract
In this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving.
Hongbo Zhao 0006, Lue Fan, Yuntao Chen, Yuran Yang, Xiaojuan Jin, Gaofeng Meng, Zhaoxiang Zhang 0001
NeurIPS8
2024 DDGPnP: Differential degree graph based PnP solution to handle outliers
Zhichao Cui, Zeqi Chen, Chi Zhang 0020, Gaofeng Meng, Yuehu Liu, Xiangmo Zhao
Comput. Vis. Image Underst.4
2024 Reusable Architecture Growth for Continual Stereo Matching
abstract
The remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks, this needs gathering training data that covers a number of heterogeneous scenes at deployment time. However, training samples are typically acquired continuously in practical applications, making the capability to learn new scenes continually even more crucial. For this purpose, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at inference. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth to learn new scenes continually in both supervised and self-supervised manners. It can maintain high reusability during growth by reusing previous units while obtaining good performance. Additionally, we present a Scene Router module to adaptively select the scene-specific architecture path at inference. Comprehensive experiments on numerous datasets show that our framework performs impressively in various weather, road, and city circumstances and surpasses the state-of-the-art methods in more challenging cross-dataset settings. Further experiments also demonstrate the adaptability of our method to unseen scenes, which can facilitate end-to-end stereo architecture learning and practical deployment.
Chenghao Zhang 0003, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Shiming Xiang, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 MoBoo: Memory-Boosted Vision Transformer for Class-Incremental Learning
abstract
Continual learning strives to acquire knowledge across sequential tasks without forgetting previously assimilated knowledge. Current state-of-the-art methodologies utilize dynamic architectural strategies to increase the network capacity for new tasks. However, these approaches often suffer from a rapid growth in the number of parameters. While some methods introduce an additional network compression stage to address this, they tend to construct complex and hyperparameter-sensitive systems. In this work, we introduce a novel solution to this challenge by proposing Memory-Boosted transformer (MoBoo), instead of conventional architecture expansion and compression. Specifically, we design a memory-augmented attention mechanism by establishing a memory bank where the “key” and “value” linear projections are stored. This memory integration prompts the model to leverage previously learned knowledge, thereby enhancing stability during training at a marginal cost. The memory bank is lightweight and can be easily managed with a straightforward queue. Moreover, to increase the model’s plasticity, we design a memory-attentive aggregator, which leverages the cross-attention mechanism to adaptively summarize the image representation from the encoder output that has historical knowledge involved. Extensive experiments on challenging benchmarks demonstrate the effectiveness of our method. For example, on ImageNet-100 under 10 tasks, our method outperforms the current state-of-the-art methods by +3.74% in average accuracy and using fewer parameters.
Bolin Ni, Xing Nie, Chenghao Zhang 0003, Shixiong Xu, Xin Zhang 0093, Gaofeng Meng, Shiming Xiang
IEEE Trans. Circuits Syst. Video Technol.6
2024 Pro-Tuning: Unified Prompt Tuning for Vision Tasks
abstract
In computer vision, fine-tuning is the de-facto approach to leverage pre-trained vision models to perform downstream tasks. However, deploying it in practice is quite challenging, due to adopting parameter inefficient global update and heavily relying on high-quality downstream data. Recently, prompt-based learning, which adds the task-relevant prompt to adapt the pre-trained models to downstream tasks, has drastically boosted the performance of many natural language downstream tasks. In this work, we extend this notable transfer ability benefited from prompt into vision models as an alternative to fine-tuning. To this end, we propose parameter-efficient Prompt tuning (Pro-tuning) to adapt diverse frozen pre-trained models to a wide variety of downstream vision tasks. The key to Pro-tuning is prompt-based tuning, i.e., learning task-specific vision prompts for downstream input images with the pre-trained model frozen. By only training a small number of additional parameters, Pro-tuning can generate compact and robust downstream models both for CNN-based and transformer-based network architectures. Comprehensive experiments evidence that the proposed Pro-tuning outperforms fine-tuning on a broad range of vision tasks and scenarios, including image classification (under generic objects, class imbalance, image corruption, natural adversarial examples, and out-of-distribution generalization), and dense prediction tasks such as object detection and semantic segmentation.
Xing Nie, Bolin Ni, Jianlong Chang, Gaofeng Meng, Chunlei Huo, Shiming Xiang, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Active Disparity Sampling for Stereo Matching With Adjoint Network
abstract
The sparse signals provided by external sources have been leveraged as guidance for improving dense disparity estimation. However, previous methods assume depth measurements to be randomly sampled, which restricts performance improvements due to under-sampling in challenging regions and over-sampling in well-estimated areas. In this work, we introduce an Active Disparity Sampling problem that selects suitable sampling patterns to enhance the utility of depth measurements given arbitrary sampling budgets. We achieve this goal by learning an Adjoint Network for a deep stereo model to measure its pixel-wise disparity quality. Specifically, we design a hard-soft prior supervision mechanism to provide hierarchical supervision for learning the quality map. A Bayesian optimized disparity sampling policy is further proposed to sample depth measurements with the guidance of the disparity quality. Extensive experiments on standard datasets with various stereo models demonstrate that our method is suited and effective in different stereo architectures and outperforms existing fixed and adaptive sampling methods under different sampling rates. Remarkably, the proposed method makes substantial improvements when generalized to heterogeneous unseen domains.
Chenghao Zhang 0003, Gaofeng Meng, Bolin Ni, Shiming Xiang
IEEE Trans. Image Process.2
2023 Domain Decorrelation with Potential Energy Ranking
abstract
Machine learning systems, especially the methods based on deep learning, enjoy great success in modern computer vision tasks under ideal experimental settings. Generally, these classic deep learning methods are built on the i.i.d. assumption, supposing the training and test data are drawn from the same distribution independently and identically. However, the aforementioned i.i.d. assumption is, in general, unavailable in the real-world scenarios, and as a result, leads to sharp performance decay of deep learning algorithms. Behind this, domain shift is one of the primary factors to be blamed. In order to tackle this problem, we propose using Potential Energy Ranking (PoER) to decouple the object feature and the domain feature in given images, promoting the learning of label-discriminative representations while filtering out the irrelevant correlations between the objects and the background. PoER employs the ranking loss in shallow layers to make features with identical category and domain labels close to each other and vice versa. This makes the neural networks aware of both objects and background characteristics, which is vital for generating domain-invariant features. Subsequently, with the stacked convolutional blocks, PoER further uses the contrastive loss to make features within the same categories distribute densely no matter domains, filtering out the domain information progressively for feature alignment. PoER reports superior performance on domain generalization benchmarks, improving the average top-1 accuracy by at least 1.20% compared to the existing methods. Moreover, we use PoER in the ECCV 2022 NICO Challenge, achieving top place with only a vanilla ResNet-18 and winning the jury award. The code has been made publicly available at: https://github.com/ForeverPs/PoER.
Sen Pei, Jiaxi Sun, Shiming Xiang, Gaofeng Meng
AAAI5
2023 Robust Feature Rectification of Pretrained Vision Models for Object Recognition
abstract
Pretrained vision models for object recognition often suffer a dramatic performance drop with degradations unseen during training. In this work, we propose a RObust FEature Rectification module (ROFER) to improve the performance of pretrained models against degradations. Specifically, ROFER first estimates the type and intensity of the degradation that corrupts the image features. Then, it leverages a Fully Convolutional Network (FCN) to rectify the features from the degradation by pulling them back to clear features. ROFER is a general-purpose module that can address various degradations simultaneously, including blur, noise, and low contrast. Besides, it can be plugged into pretrained models seamlessly to rectify the degraded features without retraining the whole model. Furthermore, ROFER can be easily extended to address composite degradations by adopting a beam search algorithm to find the composition order. Evaluations on CIFAR-10 and Tiny-ImageNet demonstrate that the accuracy of ROFER is 5% higher than that of SOTA methods on different degradations. With respect to composite degradations, ROFER improves the accuracy of a pretrained CNN by 10% and 6% on CIFAR-10 and Tiny-ImageNet respectively.
Shengchao Zhou, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang
AAAI2
2023 Bilateral Memory Consolidation for Continual Learning
abstract
Humans are proficient at continuously acquiring and integrating new knowledge. By contrast, deep models forget catastrophically, especially when tackling highly long task sequences. Inspired by the way our brains constantly rewrite and consolidate past recollections, we propose a novel Bilateral Memory Consolidation (BiMeCo) framework that focuses on enhancing memory interaction capabilities. Specifically, BiMeCo explicitly decouples model parameters into short-term memory module and long-term memory module, responsible for representation ability of the model and generalization over all learned tasks, respectively. BiMeCo encourages dynamic interactions between two memory modules by knowledge distillation and momentum-based updating for forming generic knowledge to prevent forgetting. The proposed BiMeCo is parameter-efficient and can be integrated into existing methods seamlessly. Extensive experiments on challenging benchmarks show that BiMeCo significantly improves the performance of existing continual learning methods. For example, combined with the state-of-the-art method CwD [55], BiMeCo brings in significant gains of around 2% to 6% while using 2x fewer parameters on CIFAR-100 under ResNet-18.
Xing Nie, Shixiong Xu, Xiyan Liu, Gaofeng Meng, Chunlei Huo, Shiming Xiang
CVPR4
2023 Generalization Across Subjects and Sessions for EEG-based Emotion Recognition Using Multi-source Attention-based Dynamic Residual Transfer
abstract
As an important element of emotional brain-computer interfaces, electroencephalography (EEG) signals have made significant progress in emotion recognition due to their high temporal resolution and reliability. However, EEG signals vary widely among individuals and do not satisfy temporal non-stationarity. Furthermore, trained models cannot maintain good classification accuracy for new individuals or new sessions during the inference stage. Although domain adaptation has been employed to address these issues, most approaches that consider different subjects or sessions as a single source domain ignore the large discrepancies between source domains, while methods that consider multi-source domains need to construct a domain adaptation branch for each source domain. Here, we propose a novel emotion recognition method, i.e., multi-source attention-based dynamic residual transfer (MS-ADRT). We introduce a dynamic feature extractor, in which the model uses an attention module to induce parameters to vary with the sample, implicitly enabling multi-source domain adaptation by adapting to the sample, thus reducing multi-source domain adaptation to single-source domain adaptation. Maximum mean discrepancy (MMD) and maximum classifier discrepancy (MCD)-based adversarial training are also used to narrow distances between source and target domains and facilitate the feature extractor to mine domain-invariant and sentiment-distinguishable features. We compared our algorithm with representative methods using the SEED and SEED-IV datasets, and experimentally verified that our method outperforms other state-of-the-art approaches. The proposed method provides a more effective transfer learning pathway for EEG-based sentiment analysis under multi-source scenarios.
Wanqing Jiang, Gaofeng Meng, Tianzi Jiang, Nianming Zuo
IJCNN2
2023 A Unified Model for Video Understanding and Knowledge Embedding with Heterogeneous Knowledge Graph Dataset
abstract
Video understanding is an important task in short video business platforms and it has a wide application in video recommendation and classification. Most of the existing video understanding works only focus on the information that appeared within the video content, including the video frames, audio and text. However, introducing common sense knowledge from the external Knowledge Graph (KG) dataset is essential for video understanding when referring to the content which is less relevant to the video. Owing to the lack of video knowledge graph dataset, the work which integrates video understanding and KG is rare. In this paper, we propose a heterogeneous dataset that contains the multi-modal video entity and fruitful common sense relations. This dataset also provides multiple novel video inference tasks like the Video-Relation-Tag (VRT) and Video-Relation-Video (VRV) tasks. Furthermore, based on this dataset, we propose an end-to-end model that jointly optimizes the video understanding objective with knowledge graph embedding, which can not only better inject factual knowledge into video understanding but also generate effective multi-modal entity embedding for KG. Comprehensive experiments indicate that combining video understanding embedding with factual knowledge benefits the content-based video retrieval performance. Moreover, it also helps the model generate better knowledge graph embedding which outperforms traditional KGE-based methods on VRT and VRV tasks with at least 42.36% and 17.73% improvement in [email protected].
Jiaxin Deng, Dong Shen 0003, Haojie Pan, Ximan Liu, Gaofeng Meng, Fan Yang 0094, Tingting Gao, Ruiji Fu, Zhongyuan Wang 0006
ICMR6
2022 Continual Stereo Matching of Continuous Driving Scenes with Growing Architecture
abstract
The deep stereo models have achieved state-of-the-art performance on driving scenes, but they suffer from severe performance degradation when tested on unseen scenes. Although recent work has narrowed this performance gap through continuous online adaptation, this setup requires continuous gradient updates at inference and can hardly deal with rapidly changing scenes. To address these challenges, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at deployment. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth for continual learning of new scenes. During growth, it can maintain high reusability by reusing previous neural units while achieving good performance. A module named Scene Router is further introduced to adaptively select the scene-specific architecture path at inference. Experimental results demonstrate that our method achieves compelling performance in various types of challenging driving scenes.
Chenghao Zhang 0003, Bin Fan 0001, Gaofeng Meng, Zhaoxiang Zhang 0001, Chunhong Pan
CVPR4
2022 Expanding Language-Image Pretrained Models for General Video Recognition
Bolin Ni, Houwen Peng, Songyang Zhang 0004, Gaofeng Meng, Jianlong Fu, Shiming Xiang, Haibin Ling
ECCV (4)5
2022 Out-of-distribution Detection with Boundary Aware Learning
Sen Pei, Xin Zhang 0093, Bin Fan 0001, Gaofeng Meng
ECCV (24)4
2022 Stereo Depth Estimation with Echoes
Chenghao Zhang 0003, Bolin Ni, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Chunhong Pan
ECCV (27)4
2022 Spatiotemporal Contextual Consistency Network for Precipitation Nowcasting
abstract
Precipitation nowcasting is forecasting rainfall in the short-term conditioned by the known meteorological parameters. Recently, deep neural networks (DNNs) have shown outstanding performance in this task. But, there are several challenges imposed by the multiple meteorological elements, including the multimodal modeling, the considerable variation in scales of precipitation region, as well as the long-tailed distribution of rainfall data. To solve these problems, this paper proposes Spatiotemporal Contextual Consistency Network (SCCN) for learning from the multi meteorological elements. Architecturally, a parameter-shared multimodal fusion CNN encoder, which dynamically exchanges features between different modalities, is used to encode the multimodal meteorological data. To improve the spatial modeling of the multiple meteorological features, we compose the multi-scale filters and deconstruction convolution to modify the gate operators in ConvLSTM to propose a spatial contextual consistency ConvLSTM (SCC-ConvLSTM). Furthermore, considering the temporal consistency in rainfall, a temporal consistency module (TCM) is designed to gear to long-tailed distribution. Under this module, different long-tailed meteorological elements are calculated to encode features and residuals fused with the previous precipitation distribution in sequence. The experimental results of precipitation nowcasting demonstrate the effectiveness of our method on the ERA5 dataset and WeatherBench dataset.
Xinyu Xiao, Qizhao Jin, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICDM3
2022 Components Regulated Generation of Handwritten Chinese Text-lines in Arbitrary Length
abstract
Generating readable images of handwritten Chinese text-lines is very challenging due to complicated topological structures in Chinese. To address this problem, we propose a components regulated model named HCT-GAN to generate the entire lines of Chinese handwriting from text-line labels. Specifically, HCT-GAN is designed as a CGAN-based architecture that additionally integrates a Chinese text encoder (CTE), a sequence recognition module(SRM), and a spatial perception module (SPM). Compared with the one-hot embedding, CTE learns the latent content representation by reusing the structure and component embedding shared among the Chinese characters. SRM provides sequence-level constraints to the generated images. SPM can adaptively constrain the spatial correlation between the generated components, which facilitates the modeling of characters with complicated topological structures. Benefiting from such artful modeling, our model suffices to generate images of handwritten Chinese text-lines in arbitrary length. Extensive experimental results demonstrate that our model achieves state-of-the-art performance in handwritten Chinese lines generation.
Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICPR3
2022 Learning adversarial point-wise domain alignment for stereo matching
Chenghao Zhang 0003, Gaofeng Meng, Shiming Xiang, Chunhong Pan
Neurocomputing2
2022 Monocular contextual constraint for stereo matching with adaptive weights assignment
Chenghao Zhang 0003, Gaofeng Meng, Shiming Xiang, Chunhong Pan
Image Vis. Comput.2
2022 Density-Aware Haze Image Synthesis by Self-Supervised Content-Style Disentanglement
abstract
The key procedure of haze image synthesis with adversarial training lies in the disentanglement of the feature involved only in haze synthesis, i.e.,the style feature, from the feature representing the invariant semantic content, i.e.,the content feature. Previous methods introduced a binary classifier to constrain the domain membership from being distinguished through the learned content feature during the training stage, thereby the style information is separated from the content feature. However, we find that these methods cannot achieve complete content-style disentanglement. The entanglement of the flawed style feature with content information inevitably leads to the inferior rendering of haze images. To address this issue, we propose a self-supervised style regression model with stochastic linear interpolation that can suppress the content information in the style feature. Ablative experiments demonstrate the disentangling completeness and its superiority in density-aware haze image synthesis. Moreover, the synthesized haze data are applied to test the generalization ability of vehicle detectors. Further study on the relation between haze density and detection performance shows that haze has an obvious impact on the generalization ability of vehicle detectors and that the degree of performance degradation is linearly correlated to the haze density, which in turn validates the effectiveness of the proposed method.
Chi Zhang 0020, Zihang Lin, Liheng Xu, Zongliang Li, Wei Tang 0016, Yuehu Liu, Gaofeng Meng, Le Wang 0003, Li Li 0013
IEEE Trans. Circuits Syst. Video Technol.7
2022 Decoupled Representation Learning for Character Glyph Synthesis
abstract
Character glyph synthesis is still an open challenging problem, which involves two related aspects,i.e., font style transfer and content consistency. In this paper, we propose a novel model named FontGAN, which integrates the character structure stylization, de-stylization and texture transfer into a unified framework. Specifically, we decouple character images into style representation and content representation, which offers fine-grained control of these two types of variables, thus improving the quality of the generated results. To effectively capture the style information, a style consistency module (SCM) is introduced. Technically, SCM exploits category-guided Kullback-Leibler divergence to explicitly model the style representation into different prior distributions. In this way, our model is capable of implementing transformations between multiple domains in one framework. In addition, we propose content prior module (CPM) to provide content prior for the model to guide the content encoding process and alleviates the problem of stroke deficiency during structure de-stylization. Benefiting from the idea of decoupling and regrouping, our FontGAN suffices to achieve many-to-many translation tasks for glyph structure. Experimental results demonstrate that the proposed FontGAN achieves the state-of-the-art performance in character glyph synthesis.
Xiyan Liu, Gaofeng Meng, Jianlong Chang, Ruiguang Hu, Shiming Xiang, Chunhong Pan
IEEE Trans. Multim.2
2021 Enhanced Boundary Learning for Glass-like Object Segmentation
abstract
Glass-like objects such as windows, bottles, and mirrors exist widely in the real world. Sensing these objects has many applications, including robot navigation and grasping. However, this task is very challenging due to the arbitrary scenes behind glass-like objects. This paper aims to solve the glass-like object segmentation problem via enhanced boundary learning. In particular, we first propose a novel refined differential module that outputs finer boundary cues. We then introduce an edge-aware point-based graph convolution network module to model the global shape along the boundary. We use these two modules to design a decoder that generates accurate and clean segmentation results, especially on the object contours. Both modules are lightweight and effective: they can be embedded into various segmentation models. In extensive experiments on three recent glass-like object segmentation datasets, including Trans10k, MSD, and GDD, our approach establishes new state-of-the-art results. We also illustrate the strong generalization properties of our method on three generic segmentation datasets, including Cityscapes, BDD, and COCO Stuff. Code and models will be available for further research.
Hao He 0015, Xiangtai Li, Jianping Shi, Yunhai Tong, Gaofeng Meng, Véronique Prinet, Lubin Weng
ICCV6
2021 Differentiable Convolution Search for Point Cloud Processing
abstract
Exploiting convolutional neural networks for point cloud processing is quite challenging, due to the inherent irregular distribution and discrete shape representation of point clouds. To address these problems, many handcrafted convolution variants have sprung up in recent years. Though with elaborate design, these variants could be far from optimal in sufficiently capturing diverse shapes formed by discrete points. In this paper, we propose PointSeaConv, i.e., a novel differential convolution search paradigm on point clouds. It can work in a purely data-driven manner and thus is capable of auto-creating a group of suitable convolutions for geometric shape modeling. We also propose a joint optimization framework for simultaneous search of internal convolution and external architecture, and introduce epsilon-greedy algorithm to alleviate the effect of discretization error. As a result, PointSeaNet, a deep network that is sufficient to capture geometric shapes at both convolution level and architecture level, can be searched out for point cloud processing. Extensive experiments strongly evidence that our proposed PointSeaNet surpasses current handcrafted deep models on challenging benchmarks across multiple tasks with remarkable margins.
Xing Nie, Yongcheng Liu, Shaohong Chen, Jianlong Chang, Chunlei Huo, Gaofeng Meng, Qi Tian 0001, Chunhong Pan
ICCV6
2021 EAT-NAS: elastic architecture transfer for accelerating large-scale neural architecture search
Jiemin Fang, Yukang Chen, Xinbang Zhang, Qian Zhang 0009, Chang Huang, Gaofeng Meng, Wenyu Liu 0001, Xinggang Wang
Sci. China Inf. Sci.6
2021 Dynamic camera configuration learning for high-confidence active object detection
Nuo Xu 0006, Chunlei Huo, Xin Zhang 0093, Gaofeng Meng, Chunhong Pan
Neurocomputing5
2021 DATA: Differentiable ArchiTecture Approximation With Distribution Guided Sampling
abstract
Neural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap effectively, we develop Differentiable ArchiTecture Approximation (DATA) with Ensemble Gumbel-Softmax (EGS) estimator and Architecture Distribution Constraint (ADC) to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients reversely, reducing the estimation bias in a differentiable way. To narrow the distribution gap between sampled architectures and supernet, further, the ADC is introduced to reduce the variance of sampling during searching. Benefiting from such modeling, architecture probabilities and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep neural architectures in an extended search space. Conclusively, in the validating process, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on various tasks including image classification, few-shot learning, unsupervised clustering, semantic segmentation and language modeling strongly demonstrate that DATA is capable of discovering high-performance architectures while guaranteeing the required efficiency. Code is available at https://github.com/XinbangZhang/DATA-NAS.
Xinbang Zhang, Jianlong Chang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Zhouchen Lin, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Handwritten Text Generation via Disentangled Representations
abstract
Automatically generating handwritten text images is a challenging task due to the diverse handwriting styles and the irregular writing in natural scenes. In this paper, we propose an effective generative model called HTG-GAN to synthesize handwritten text images from latent prior. Unlike single-character synthesis, our method is capable of generating images of sequence characters with arbitrary length, which pays more attention to the structural relationship between characters. We model the structural relationship as the style representation to avoid explicitly modeling the stroke layout. Specifically, the text image is disentangled into style representation and content representation, where the style representation is mapped into Gaussian distribution and the content representation is embedded using character index. In this way, our model can generate new handwritten text images with specified contents and various styles to perform data augmentation, thereby boosting handwritten text recognition (HTR). Experimental results show that our method achieves state-of-the-art performance in handwritten text generation.
Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan
IEEE Signal Process. Lett.2
2020 Spatio-Temporal Graph Structure Learning for Traffic Forecasting
abstract
As an indispensable part in Intelligent Traffic System (ITS), the task of traffic forecasting inherently subjects to the following three challenging aspects. First, traffic data are physically associated with road networks, and thus should be formatted as traffic graphs rather than regular grid-like tensors. Second, traffic data render strong spatial dependence, which implies that the nodes in the traffic graphs usually have complex and dynamic relationships between each other. Third, traffic data demonstrate strong temporal dependence, which is crucial for traffic time series modeling. To address these issues, we propose a novel framework named Structure Learning Convolution (SLC) that enables to extend the traditional convolutional neural network (CNN) to graph domains and learn the graph structure for traffic forecasting. Technically, SLC explicitly models the structure information into the convolutional operation. Under this framework, various non-Euclidean CNN methods can be considered as particular instances of our formulation, yielding a flexible mechanism for learning on the graph. Along this technical line, two SLC modules are proposed to capture the global and local structures respectively and they are integrated to construct an end-to-end network for traffic forecasting. Additionally, in this process, Pseudo three Dimensional convolution (P3D) networks are combined with SLC to capture the temporal dependencies in traffic data. Extensively comparative experiments on six real-world datasets demonstrate our proposed approach significantly outperforms the state-of-the-art ones.
Jianlong Chang, Gaofeng Meng, Shiming Xiang, Chunhong Pan
AAAI3
2020 Deep Self-Evolution Clustering
abstract
Clustering is a crucial but challenging task in pattern analysis and machine learning. Existing methods often ignore the combination between representation learning and clustering. To tackle this problem, we reconsider the clustering task from its definition to develop Deep Self-Evolution Clustering (DSEC) to jointly learn representations and cluster data. For this purpose, the clustering task is recast as a binary pairwise-classification problem to estimate whether pairwise patterns are similar. Specifically, similarities between pairwise patterns are defined by the dot product between indicator features which are generated by a deep neural network (DNN). To learn informative representations for clustering, clustering constraints are imposed on the indicator features to represent specific concepts with specific representations. Since the ground-truth similarities are unavailable in clustering, an alternating iterative algorithm called Self-Evolution Clustering Training (SECT) is presented to select similar and dissimilar pairwise patterns and to train the DNN alternately. Consequently, the indicator features tend to be one-hot vectors and the patterns can be clustered by locating the largest response of the learned indicator features. Extensive experiments strongly evidence that DSEC outperforms current models on twelve popular image, text and audio datasets consistently.
Jianlong Chang, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Local-Aggregation Graph Networks
abstract
Convolutional neural networks (CNNs) provide a dramatically powerful class of models, but are subject to traditional convolution that can merely aggregate permutation-ordered and dimension-equal local inputs. It causes that CNNs are allowed to only manage signals on Euclidean or grid-like domains (e.g., images), not ones on non-Euclidean or graph domains (e.g., traffic networks). To eliminate this limitation, we develop a local-aggregation function, a sharable nonlinear operation, to aggregate permutation-unordered and dimension-unequal local inputs on non-Euclidean domains. In the context of the function approximation theory, the local-aggregation function is parameterized with a group of orthonormal polynomials in an effective and efficient manner. By replacing the traditional convolution in CNNs with the parameterized local-aggregation function, Local-Aggregation Graph Networks (LAGNs) are readily established, which enable to fit nonlinear functions without activation functions and can be expediently trained with the standard back-propagation. Extensive experiments on various datasets strongly demonstrate the effectiveness and efficiency of LAGNs, leading to superior performance on numerous pattern recognition and machine learning tasks, including text categorization, molecular activity detection, taxi flow prediction, and image classification.
Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Baselines Extraction from Curved Document Images via Slope Fields Recovery
abstract
Baselines estimation is a critical preprocessing step for many tasks of document image processing and analysis. The problem is very challenging due to arbitrarily complicated page layouts and various types of image quality degradations. This paper proposes a method based on slope fields recovery for curved baseline extraction from a distorted document image captured by a hand-held camera. Our method treats the curved baselines as the solution curves of an ordinary differential equation defined on a slope field. By assuming the page shape is a smooth and developable surface, we investigate a type of intrinsic geometric constraints of baselines to estimate the latent slope field. The curved baselines are finally obtained by solving an ordinary differential equation through the Euler method. Unlike the traditional text-lines based methods, our method is free from text-lines detection and segmentation. It can exploit multiple visual cues other than horizontal text-lines available in images for baselines extraction and is quite robust to document scripts, various types of image quality degradation (e.g., image distortion, blur and non-uniform illumination), large areas of non-textual objects and complex page layouts. Extensive experiments on synthetic and real-captured document images are implemented to evaluate the performance of the proposed method.
Gaofeng Meng, Chunhong Pan, Shiming Xiang, Ying Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Geometric rectification of document images using adversarial gated unwarping network
Xiyan Liu, Gaofeng Meng, Bin Fan 0001, Shiming Xiang, Chunhong Pan
Pattern Recognit.2
2019 No-Reference Image Quality Assessment with Reinforcement Recursive List-Wise Ranking
abstract
Opinion-unaware no-reference image quality assessment (NR-IQA) methods have received many interests recently because they do not require images with subjective scores for training. Unfortunately, it is a challenging task, and thus far no opinion-unaware methods have shown consistently better performance than the opinion-aware ones. In this paper, we propose an effective opinion-unaware NR-IQA method based on reinforcement recursive list-wise ranking. We formulate the NR-IQA as a recursive list-wise ranking problem which aims to optimize the whole quality ordering directly. During training, the recursive ranking process can be modeled as a Markov decision process (MDP). The ranking list of images can be constructed by taking a sequence of actions, and each of them refers to selecting an image for a specific position of the ranking list. Reinforcement learning is adopted to train the model parameters, in which no ground-truth quality scores or ranking lists are necessary for learning. Experimental results demonstrate the superior performance of our approach compared with existing opinion-unaware NR-IQA methods. Furthermore, our approach can compete with the most effective opinion-aware methods. It improves the state-of-the-art by over 2% on the CSIQ benchmark and outperforms most compared opinion-aware models on TID2013.
Jie Gu 0002, Gaofeng Meng, Cheng Da, Shiming Xiang, Chunhong Pan
AAAI2
2019 RENAS: Reinforced Evolutionary Neural Architecture Search
abstract
Neural Architecture Search (NAS) is an important yet challenging task in network design due to its high computational consumption. To address this issue, we propose the Reinforced Evolutionary Neural Architecture Search (RENAS), which is an evolutionary method with reinforced mutation for NAS. Our method integrates reinforced mutation into an evolution algorithm for neural architecture exploration, in which a mutation controller is introduced to learn the effects of slight modifications and make mutation actions. The reinforced mutation controller guides the model population to evolve efficiently. Furthermore, as child models can inherit parameters from their parents during evolution, our method requires very limited computational resources. In experiments, we conduct the proposed search method on CIFAR-10 and obtain a powerful network architecture, RENASNet. This architecture achieves a competitive result on CIFAR-10. The explored network architecture is transferable to ImageNet and achieves a new state-of-the-art accuracy, i.e., 75.7% top-1 accuracy with 5.36M parameters on mobile ImageNet. We further test its performance on semantic segmentation with DeepLabv3 on the PASCAL VOC. RENASNet outperforms MobileNet-v1, MobileNet-v2 and NASNet. It achieves 75.83% mIOU without being pretrained on COCO.
Yukang Chen, Gaofeng Meng, Qian Zhang 0009, Shiming Xiang, Chang Huang, Lisen Mu, Xinggang Wang
CVPR2
2019 DensePoint: Learning Densely Contextual Representation for Efficient Point Cloud Processing
abstract
Point cloud processing is very challenging, as the diverse shapes formed by irregular points are often indistinguishable. A thorough grasp of the elusive shape requires sufficiently contextual semantic information, yet few works devote to this. Here we propose DensePoint, a general architecture to learn densely contextual representation for point cloud processing. Technically, it extends regular grid CNN to irregular point configuration by generalizing a convolution operator, which holds the permutation invariance of points, and achieves efficient inductive learning of local patterns. Architecturally, it finds inspiration from dense connection mode, to repeatedly aggregate multi-level and multi-scale semantics in a deep hierarchy. As a result, densely contextual information along with rich semantics, can be acquired by DensePoint in an organic manner, making it highly effective. Extensive experiments on challenging benchmarks across four tasks, as well as thorough model analysis, verify DensePoint achieves the state of the arts.
Yongcheng Liu, Bin Fan 0001, Gaofeng Meng, Jiwen Lu, Shiming Xiang, Chunhong Pan
ICCV3
2019 GSTNet: Global Spatial-Temporal Network for Traffic Flow Prediction
abstract
Predicting traffic flow on traffic networks is a very challenging task, due to the complicated and dynamic spatial-temporal dependencies between different nodes on the network. The traffic flow renders two types of temporal dependencies, including short-term neighboring and long-term periodic dependencies. What's more, the spatial correlations over different nodes are both local and non-local. To capture the global dynamic spatial-temporal correlations, we propose a Global Spatial-Temporal Network (GSTNet), which consists of several layers of spatial-temporal blocks. Each block contains a multi-resolution temporal module and a global correlated spatial module in sequence, which can simultaneously extract the dynamic temporal dependencies and the global spatial correlations. Extensive experiments on the real world datasets verify the effectiveness and superiority of the proposed method on both the public transportation network and the road network.
Shen Fang, Gaofeng Meng, Shiming Xiang, Chunhong Pan
IJCAI3
2019 DATA: Differentiable ArchiTecture Approximation
abstract
Neural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap, we develop Differentiable ArchiTecture Approximation (DATA) with an Ensemble Gumbel-Softmax (EGS) estimator to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients from binary codes to probability vectors. Benefiting from such modeling, in searching, architecture parameters and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep models in a large enough search space. Conclusively, during validating, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on a variety of popular datasets strongly evidence that our method is capable of discovering high-performance architectures for image classification, language modeling and semantic segmentation, while guaranteeing the requisite efficiency during searching.
Jianlong Chang, Xinbang Zhang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Chunhong Pan
NeurIPS4
2019 DetNAS: Backbone Search for Object Detection
abstract
Object detectors are usually equipped with backbone networks designed for image classification. It might be sub-optimal because of the gap between the tasks of image classification and object detection. In this work, we present DetNAS to use Neural Architecture Search (NAS) for the design of better backbones for object detection. It is non-trivial because detection training typically needs ImageNetpre-training while NAS systems require accuracies on the target detection task as supervisory signals. Based on the technique of one-shot supernet, which contains all possible networks in the search space, we propose a framework for backbone search on object detection. We train the supernet under the typical detector training schedule: ImageNet pre-training and detection fine-tuning. Then, the architecture search is performed on the trained supernet, using the detection task as the guidance. This framework makes NAS on backbones very efficient. In experiments, we show the effectiveness of DetNAS on various detectors, for instance, one-stage RetinaNetand the two-stage FPN. We empirically find that networks searched on object detection shows consistent superiority compared to those searched on ImageNet classification. The resulting architecture achieves superior performance than hand-crafted networks on COCO with much less FLOPs complexity.
Yukang Chen, Tong Yang 0005, Xiangyu Zhang 0005, Gaofeng Meng, Xinyu Xiao, Jian Sun 0001
NeurIPS4
2019 Scene text detection and recognition with advances in deep learning: a survey
Xiyan Liu, Gaofeng Meng, Chunhong Pan
Int. J. Document Anal. Recognit.2
2019 Nonlinear Asymmetric Multi-Valued Hashing
abstract
Most existing hashing methods resort to binary codes for large scale similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose Nonlinear Asymmetric Multi-Valued Hashing (NAMVH) supported by two distinct non-binary embeddings. Specifically, a real-valued embedding is used for representing the newly-coming query by an ideally nonlinear transformation. Besides, a multi-integer-embedding is employed for compressing the whole database, which is modeled by Binary Sparse Representation (BSR) with fixed sparsity. With these two non-binary embeddings, NAMVH preserves more precise similarities between data points and enables access to the incremental extension with database samples evolving dynamically. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the pairwise label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by a well-designed alternative optimization method. Extensive experiments on seven large scale datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency.
Cheng Da, Gaofeng Meng, Shiming Xiang, Kun Ding 0001, Shibiao Xu, Qing Yang 0002, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Blind image quality assessment via learnable attention-based pooling
Jie Gu 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan
Pattern Recognit.2
2019 Learning graph structure via graph convolutional networks
Jianlong Chang, Gaofeng Meng, Shibiao Xu, Shiming Xiang, Chunhong Pan
Pattern Recognit.3
2019 Deep Video Dehazing With Semantic Segmentation
abstract
Recent research have shown the potential of using convolutional neural networks (CNNs) to accomplish single image dehazing. In this work, we take one step further to explore the possibility of exploiting a network to perform haze removal for videos. Unlike single image dehazing, video based approaches can take advantage of the abundant information that exists across neighboring frames. In this work, assuming that a scene point yields highly correlated transmission values between adjacent video frames, we develop a deep learning solution for video dehazing, where a CNN is trained end-to-end to learn how to accumulate information across frames for transmission estimation. The estimated transmission map is subsequently used to recover a haze-free frame via atmospheric scattering model. In addition, as the semantic information of a scene provides a strong prior for image restoration, we propose to incorporate global semantic priors as input to regularize the transmission maps so that the estimated maps can be smooth in the regions of the same object and only discontinuous across the boundaries of different objects. To train this network, we generate a dataset consisted of synthetic hazy and haze-free videos for supervision based on the NYU depth dataset. We show that the features learned from this dataset are capable of removing haze that arises in outdoor scenes in a wide range of videos. Extensive experiments demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods on both synthetic and real-world videos.
Wenqi Ren, Jingang Zhang, Xiangyu Xu 0002, Lin Ma 0002, Xiaochun Cao, Gaofeng Meng, Wei Liu 0005
IEEE Trans. Image Process.6
2018 Exploiting Vector Fields for Geometric Rectification of Distorted Document Images
Gaofeng Meng, Yuanqi Su, Ying Wu 0001, Shiming Xiang, Chunhong Pan
ECCV (16)1
2018 Semantic Image Synthesis via Conditional Cycle-Generative Adversarial Networks
abstract
Traditional approaches for semantic image synthesis mainly focus on text descriptions while ignoring the related structures and attributes in the original images. Therefore, some critical information, e.g., the style, backgrounds, objects shapes and pose, is missed in the generated images. In this paper, we propose a novel framework called Conditional Cycle-Generative Adversarial Network (CCGAN) to address this issue. Our model can generate photo-realistic images conditioned on the given text descriptions, while maintaining the attributes of the original images. The framework mainly consists of two coupled conditional adversarial networks, which are able to learn a desirable image mapping that can keep the structures and attributes in the images. We introduce a conditional cycle consistency loss to prevent the contradiction between two generators. This loss allows the generated images to retain most of the features of the original image, so as to improve the stability of network training. Moreover, benefiting from the mechanism of circular training, the proposed networks can learn the semantic information of the text much accurately. Experiments on Caltech-UCSD Bird dataset and Oxford-102 flower dataset demonstrate that the proposed method significantly outperforms the existing methods in terms of image details reconstruction and semantic information expression.
Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICPR2
2018 Structure-Aware Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) are inherently subject to invariable filters that can only aggregate local inputs with the same topological structures. It causes that CNNs are allowed to manage data with Euclidean or grid-like structures (e.g., images), not ones with non-Euclidean or graph structures (e.g., traffic networks). To broaden the reach of CNNs, we develop structure-aware convolution to eliminate the invariance, yielding a unified mechanism of dealing with both Euclidean and non-Euclidean structured data. Technically, filters in the structure-aware convolution are generalized to univariate functions, which are capable of aggregating local inputs with diverse topological structures. Since infinite parameters are required to determine a univariate function, we parameterize these filters with numbered learnable parameters in the context of the function approximation theory. By replacing the classical convolution in CNNs with the structure-aware convolution, Structure-Aware Convolutional Neural Networks (SACNNs) are readily established. Extensive experiments on eleven datasets strongly evidence that SACNNs outperform current models on various machine learning tasks, including image classification and clustering, text categorization, skeleton-based action recognition, molecular activity detection, and taxi flow prediction.
Jianlong Chang, Jie Gu 0002, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan
NeurIPS4
2018 Facade repetition detection in a fronto-parallel view with fiducial lines extraction
Hongfei Xiao, Gaofeng Meng, Lingfeng Wang 0002, Chunhong Pan
Neurocomputing2
2018 Deep unsupervised learning with consistent inference of latent representations
Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan
Pattern Recognit.3
2018 Blind Image Quality Assessment via Vector Regression and Object Oriented Pooling
abstract
This paper presents an effective method based on vector regression and object oriented pooling for blind image quality assessment. Unlike previous models that map the extracted features directly to a quality score, the proposed vector regression framework yields a vector of belief scores for the input image. We explore the uncertainty factors in quality assessment and design the belief scores to measure the confidences of an image to be assigned to the corresponding quality grades. Moreover, we propose an object oriented pooling strategy to further improve the performance by incorporating semantic information of image contents. According to this strategy, regions occupied by objects will be assigned more weights in the pooling phase, leading to a more accurate quality assessment. Extensive experiments on benchmark datasets demonstrate that our approach achieves state-of-the-art performance and shows a great generalization ability.
Jie Gu 0002, Gaofeng Meng, Judith Redi, Shiming Xiang, Chunhong Pan
IEEE Trans. Multim.2
2017 AMVH: Asymmetric Multi-Valued hashing
abstract
Most existing hashing methods resort to binary codes for similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose an asymmetric multi-valued hashing method supported by two different non-binary embeddings. (1) A real-valued embedding is used for representing the newly-coming query. (2) A multi-integer-embedding is employed for compressing the whole database, which is modeled by binary sparse representation with fixed sparsity. With these two non-binary embeddings, the similarities between data points can be preserved precisely. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by alternative optimization. Extensive experiments on three multilabel datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency.
Cheng Da, Shibiao Xu, Kun Ding 0001, Gaofeng Meng, Shiming Xiang, Chunhong Pan
CVPR4
2017 Learning deep vector regression model for no-reference image quality assessment
abstract
The goal of no-reference image quality assessment (NR-IQA) is to estimate human perceived image quality without access to either reference image or prior knowledge about distortion type. Previous approaches for this problem are typically based on a regression framework that maps the image features directly to a quality score. In contrast, psychological evidence shows that humans prefer to evaluate visual quality with qualitative descriptions, e.g., using a five-grade ordinal scale: “excellent”, “good”, “fair”, “poor” and “bad”. Based on this observation, we propose a vector regression model that predicts five belief scores rather than a single quality score. The belief scores are designed to indicate the confidences of the test image being assigned with these five quality grades. In addition, with the purpose of more extensive applications, a saliency-based pooling strategy is presented to convert the predicted confidences into objective quality scores. Extensive experiments performed on two benchmark datasets demonstrate that our approach achieves state-of-the-art performance and shows great generalization ability.
Jie Gu 0002, Gaofeng Meng, Lingfeng Wang 0002, Chunhong Pan
ICASSP2
2017 Deep Adaptive Image Clustering
abstract
Image clustering is a crucial but challenging task in machine learning and computer vision. Existing methods often ignore the combination between feature learning and clustering. To tackle this problem, we propose Deep Adaptive Clustering (DAC) that recasts the clustering problem into a binary pairwise-classification framework to judge whether pairs of images belong to the same clusters. In DAC, the similarities are calculated as the cosine distance between label features of images which are generated by a deep convolutional network (ConvNet). By introducing a constraint into DAC, the learned label features tend to be one-hot vectors that can be utilized for clustering images. The main challenge is that the ground-truth similarities are unknown in image clustering. We handle this issue by presenting an alternating iterative Adaptive Learning algorithm where each iteration alternately selects labeled samples and trains the ConvNet. Conclusively, images are automatically clustered based on the label features. Experimental results show that DAC achieves state-of-the-art performance on five popular datasets, e.g., yielding 97.75% clustering accuracy on MNIST, 52.18% on CIFAR-10 and 46.99% on STL-10.
Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICCV3
2017 Deep Networks for Degraded Document Image Binarization through Pyramid Reconstruction
abstract
Binarization of document images is an important processing step for document images analysis and recognition. However, this problem is quite challenging in some cases because of the quality degradation of document images, such as varying illumination, complicated backgrounds, image noises due to ink spots, water stains or document creases. In this paper, we propose a framework based on deep convolutional neural-network (DCNN) for adaptive binarization of degraded document images. The basic idea of our method is to decompose a degraded document image into a spatial pyramid structure by using DCNN, with each layer at different scale. Then the foreground image is sequentially reconstructed from these layers in a coarse-to-fine manner by using deconvolutional network. Such kind of decomposition is quite beneficial, since multi-resolution supervision information can be directly introduced into network learning. We also define several loss functions about label consistency and foregrounds smoothing to further regularize the training of the network. Experimental results demonstrate the effectiveness of the proposed method.
Gaofeng Meng, Kun Yuan 0003, Ying Wu 0001, Shiming Xiang, Chunhong Pan
ICDAR1
2017 Learnable contextual regularization for semantic segmentation of indoor scene images
abstract
Semantic segmentation of indoor scene images has a wide range of applications. However, due to a large number of classes and uneven distribution in indoor scenes, mislabels are often made when facing small objects or boundary regions. Technically, contextual information may benefit for segmentation results, but has not yet been exploited sufficiently. In this paper, we propose a learnable contextual regularization model for enhancing the semantic segmentation results of color indoor scene images. This regularization model is combined with a deep convolutional segmentation network without significantly increasing the number of additional parameters. Our model, derived from the inherent contextual regularization on the indoor scene objects, benefits much from the learnable constraint layers bridging the lower layers and the higher layers in the deep convolutional network. The constraint layers are further integrated with a weighted L1-norm based contextual regularization between the neighboring pixels of RGB values to improve the segmentation results. Experimental results on NYUDv2 indoor scene dataset demonstrate the effectiveness and efficiency of the proposed method.
Gaofeng Meng, Lingfeng Wang 0002, Chunhong Pan
ICIP3
2017 Image super-resolution via deep dilated convolutional networks
abstract
Deep learning techniques have been successfully applied in single image super-resolution (SR). Recently, researches have shown that increasing the depth of network can significantly improve SR performance. Very deep networks for SR achieved a large improvement than former methods. However, simply increasing depths basically introduce more parameters and this lead to cumbersome computational cost. In this paper, we present a general and effective method to accelerate very deep networks for single image SR. Our method is based on dilated convolution operation, which support exponential expansion of the receptive field without increasing filter size. With the help of dilated convolution, shallow networks can achieve large receptive field and exploit contextual information in an efficient way. Based on a very deep network, we propose a 12 layers dilated convolutional network for SR (DCNSR). While accelerating 2x speed, our shallow network achieves better performance than original deep networks and shows state-of-the-art reconstructed results.
Zehao Huang, Lingfeng Wang 0002, Gaofeng Meng, Chunhong Pan
ICIP3
2017 Efficient cloud detection in remote sensing images using edge-aware segmentation network and easy-to-hard training strategy
abstract
Detecting cloud regions in remote sensing image (RSI) is very challenging yet of great importance to meteorological forecasting and other RSI-related applications. Technically, this task is typically implemented as a pixel-level segmentation. However, traditional methods based on handcrafted or low-level cloud features often fail to achieve satisfactory performances from images with bright non-cloud and/or semitransparent cloud regions. What is more, the performances could be further degraded due to the ambiguous boundaries caused by complicated textures and non-uniform distribution of intensities. In this paper, we propose a multi-task based deep neural network for cloud detection in RSIs. Architecturally, our network is designed to combine the two tasks of cloud segmentation and cloud edge detection together to encourage a better detection near cloud boundaries, resulting in an end-to-end approach for accurate cloud detection. Accordingly, an efficient sample selection strategy is proposed to train our network in an easy-to-hard manner, in which the number of the selected samples is governed by a weight that is annealed until the entire training samples have been considered. Both visual and quantitative comparisons are conducted on RSIs collected from Google Earth. The experimental results indicate that our method can yield superior performance over the state-of-the-art methods.
Kun Yuan 0003, Gaofeng Meng, Dongcai Cheng, Shiming Xiang, Chunhong Pan
ICIP2
2017 Active Rectification of Curved Document Images Using Structured Beams
Gaofeng Meng, Shiming Xiang, Chunhong Pan, Nanning Zheng 0001
Int. J. Comput. Vis.1
2017 SeNet: Structured Edge Network for Sea-Land Segmentation
abstract
Separating an optical remote sensing image into sea and land areas is very challenging yet of great importance to coastline extraction and subsequent object detection. Traditional methods based on handcrafted feature extraction and image processing often face this dilemma when confronting high-resolution remote sensing images for their complicated texture and intensity distribution. In this letter, we apply the prevalent deep convolutional neural networks to the sea–land segmentation problem and make two innovations on top of the traditional structure. First, we propose a local smooth regularization to achieve better spatially consistent results, which frees us from the complicated morphological operations that are commonly used in traditional methods. Second, we use a multitask loss to simultaneously obtain the segmentation and edge detection results. The attached structured edge detection branch can further refine the segmentation result and dramatically improve edge accuracy. Experiments on a set of natural-colored images from Google Earth demonstrate the effectiveness of our approach in terms of quantitative and visual performances compared with state-of-the-art methods.
Dongcai Cheng, Gaofeng Meng, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.2
2016 Efficient sea-land segmentation using seeds learning and edge directed graph cut
Dongcai Cheng, Gaofeng Meng, Shiming Xiang, Chunhong Pan
Neurocomputing2
2015 Extraction of Virtual Baselines from Distorted Document Images Using Curvilinear Projection
abstract
The baselines of a document page are a set of virtual horizontal and parallel lines, to which the printed contents of document, e.g., text lines, tables or inserted photos, are aligned. Accurate baseline extraction is of great importance in the geometric correction of curved document images. In this paper, we propose an efficient method for accurate extraction of these virtual visual cues from a curved document image. Our method comes from two basic observations that the baselines of documents do not intersect with each other and that within a narrow strip, the baselines can be well approximated by linear segments. Based upon these observations, we propose a curvilinear projection based method and model the estimation of curved baselines as a constrained sequential optimization problem. A dynamic programming algorithm is then developed to efficiently solve the problem. The proposed method can extract the complete baselines through each pixel of document images in a high accuracy. It is also scripts insensitive and highly robust to image noises, non-textual objects, image resolutions and image quality degradation like blurring and non-uniform illumination. Extensive experiments on a number of captured document images demonstrate the effectiveness of the proposed method.
Gaofeng Meng, Zuming Huang, Yonghong Song, Shiming Xiang, Chunhong Pan
ICCV1
2015 Image Deblurring with Coupled Dictionary Learning
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang
Int. J. Comput. Vis.2
2014 Active Flattening of Curved Document Images via Two Structured Beams
abstract
Document images captured by a digital camera often suffer from serious geometric distortions. In this paper, we propose an active method to correct geometric distortions in a camera-captured document image. Unlike many passive rectification methods that rely on text-lines or features extracted from images, our method uses two structured beams illuminating upon the document page to recover two spatial curves. A developable surface is then interpolated to the curves by finding the correspondence between them. The developable surface is finally flattened onto a plane by solving a system of ordinary differential equations. Our method is a content independent approach and can restore a corrected document image of high accuracy with undistorted contents. Experimental results on a variety of real-captured document images demonstrate the effectiveness and efficiency of the proposed method.
Gaofeng Meng, Ying Wang 0008, Shenquan Qu, Shiming Xiang, Chunhong Pan
CVPR1
2014 Facade repetition extraction using block matrix based model
abstract
Repetition extraction plays an important role in facade image analysis. In this paper, this task is handled within the graph cut based image segmentation framework. To model the repetitions, generalized translation symmetry (GTS) is introduced to enable aperiodic repetition layouts. More importantly, GTS is explicitly formulated in terms of matrix multiplication. That is, GTS is viewed as the product of a repetitive pattern and two block matrices. These two block matrices are employed to represent the vertical and horizontal symmetry respectively. On this basis, repetition extraction is formulated as a GTS constrained energy minimization problem. An alternatively optimization algorithm based on graph cut and dynamic programming is finally developed to solve the problem. Experimental results demonstrate the validity of our method.
Hongfei Xiao, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan
ICIP2
2014 Facade Labeling via Explicit Matrix Factorization
Hongfei Xiao, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICISP3
2014 Multifocus image fusion via focus segmentation and region reconstruction
Jiangyong Duan, Gaofeng Meng, Shiming Xiang, Chunhong Pan
Neurocomputing2
2014 Spectral Unmixing via Data-Guided Sparsity
abstract
Hyperspectral unmixing, the process of estimating a common set of spectral bases and their corresponding composite percentages at each pixel, is an important task for hyperspectral analysis, visualization, and understanding. From an unsupervised learning perspective, this problem is very challenging-both the spectral bases and their composite percentages are unknown, making the solution space too large. To reduce the solution space, priors. In practice, these priors would easily lead to some unsuitable solution. This is because they are achieved by applying an identical strength of constraints to all the factors, which does not hold in practice. To overcome this limitation, we propose a novel sparsity-based method by learning a data-guided map (DgMap) to describe the individual mixed level of each pixel. Through this DgMap, the l(p) (0 < p < 1) constraint is applied in an adaptive manner. Such implementation not only meets the practical situation, but also guides the spectral bases toward the pixels under highly sparse constraint. What is more, an elegant optimization scheme as well as its convergence proof have been provided in this paper. Extensive experiments on several datasets also demonstrate that the DgMap is feasible, and high quality unmixing results could be obtained by our method.
Feiyun Zhu, Ying Wang 0008, Bin Fan 0001, Shiming Xiang, Gaofeng Meng, Chunhong Pan
IEEE Trans. Image Process.5
2013 Removing out-of-focus blur from similar image pairs
abstract
This paper presents a new deblurring method to remove the out-of-focus blur from similar image pairs. The method is motivated by an observation that a blurred structure appearing in one image can often have its corresponding clear one in the similar clear images. Our method first extracts the patch pairs from input images by SIFT matching. Then the constraints on the patch pairs are used to estimate the blur kernel via the RANSAC algorithm. Finally, the non-blind deconvolution is adopted to restore the blurred image. The main advantage is that we can improve the deblurring results with the help of additional similar clear images in many practical applications. Our method is validated on synthetic and real images by comparing with state-of-the-art methods.
Jiangyong Duan, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICASSP2
2013 Efficient Image Dehazing with Boundary Constraint and Contextual Regularization
abstract
Images captured in foggy weather conditions often suffer from bad visibility. In this paper, we propose an efficient regularization method to remove hazes from a single input image. Our method benefits much from an exploration on the inherent boundary constraint on the transmission function. This constraint, combined with a weighted L_1-norm based contextual regularization, is modeled into an optimization problem to estimate the unknown scene transmission. A quite efficient algorithm based on variable splitting is also presented to solve the problem. The proposed method requires only a few general assumptions and can restore a high-quality haze-free image with faithful colors and fine image details. Experimental results on a variety of haze images demonstrate the effectiveness and efficiency of the proposed method.
Gaofeng Meng, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan
ICCV1
2013 Nonparametric Illumination Correction for Scanned Document Images via Convex Hulls
abstract
A scanned image of an opened book page often suffers from various scanning artifacts known as scanning shading and dark borders noises. These artifacts will degrade the qualities of the scanned images and cause many problems to the subsequent process of document image analysis. In this paper, we propose an effective method to rectify these scanning artifacts. Our method comes from two observations: that the shading surface of most scanned book pages is quasi-concave and that the document contents are usually printed on a sheet of plain and bright paper. Based on these observations, a shading image can be accurately extracted via convex hulls-based image reconstruction. The proposed method proves to be surprisingly effective for image shading correction and dark borders removal. It can restore a desired shading-free image and meanwhile yield an illumination surface of high quality. More importantly, the proposed method is nonparametric and thus does not involve any user interactions or parameter fine-tuning. This would make it very appealing to nonexpert users in applications. Extensive experiments based on synthetic and real-scanned document images demonstrate the efficiency of the proposed method.
Gaofeng Meng, Shiming Xiang, Nanning Zheng 0001, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Level set evolution with locally linear classification for image segmentation
Ying Wang 0008, Shiming Xiang, Chunhong Pan, Lingfeng Wang 0002, Gaofeng Meng
Pattern Recognit.5
2013 Edge-Directed Single-Image Super-Resolution Via Adaptive Gradient Magnitude Self-Interpolation
abstract
Super-resolution from a single image plays an important role in many computer vision systems. However, it is still a challenging task, especially in preserving local edge structures. To construct high-resolution images while preserving the sharp edges, an effective edge-directed super-resolution method is presented in this paper. An adaptive self-interpolation algorithm is first proposed to estimate a sharp high-resolution gradient field directly from the input low-resolution image. The obtained high-resolution gradient is then regarded as a gradient constraint or an edge-preserving constraint to reconstruct the high-resolution image. Extensive results have shown both qualitatively and quantitatively that the proposed method can produce convincing super-resolution images containing complex and sharp features, as compared with the other state-of-the-art super-resolution algorithms.
Lingfeng Wang 0002, Shiming Xiang, Gaofeng Meng, Huai-Yu Wu, Chunhong Pan
IEEE Trans. Circuits Syst. Video Technol.3
2012 Image Guided Tone Mapping with Locally Nonlinear Model
Huxiang Gu, Ying Wang 0008, Shiming Xiang, Gaofeng Meng, Chunhong Pan
ECCV (4)4
2012 Metric Rectification of Curved Document Images
abstract
In this paper, we propose a metric rectification method to restore an image from a single camera-captured document image. The core idea is to construct an isometric image mesh by exploiting the geometry of page surface and camera. Our method uses a general cylindrical surface (GCS) to model the curved page shape. Under a few proper assumptions, the printed horizontal text lines are shown to be line convergent symmetric. This property is then used to constrain the estimation of various model parameters under perspective projection. We also introduce a paraperspective projection to approximate the nonlinear perspective projection. A set of close-form formulas is thus derived for the estimate of GCS directrix and document aspect ratio. Our method provides a straightforward framework for image metric rectification. It is insensitive to camera positions, viewing angles, and the shapes of document pages. To evaluate the proposed method, we implemented comprehensive experiments on both synthetic and real-captured images. The results demonstrate the efficiency of our method. We also carried out a comparative experiment on the public CBDAR2007 data set. The experimental results show that our method outperforms the state-of-the-art methods in terms of OCR accuracy and rectification errors.
Gaofeng Meng, Chunhong Pan, Shiming Xiang, Jiangyong Duan
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Image deblurring with matrix regression and gradient evolution
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang
Pattern Recognit.2
2012 Discriminative Least Squares Regression for Multiclass Classification and Feature Selection
abstract
This paper presents a framework of discriminative least squares regression (LSR) for multiclass classification and feature selection. The core idea is to enlarge the distance between different classes under the conceptual framework of LSR. First, a technique called ε-dragging is introduced to force the regression targets of different classes moving along opposite directions such that the distances between classes can be enlarged. Then, the ε-draggings are integrated into the LSR model for multiclass classification. Our learning framework, referred to as discriminative LSR, has a compact model form, where there is no need to train two-class machines that are independent of each other. With its compact form, this model can be naturally extended for feature selection. This goal is achieved in terms of L2,1 norm of matrix, generating a sparse learning model for feature selection. The model for multiclass classification and its extension for feature selection are finally solved elegantly and efficiently. Experimental evaluation over a range of benchmark datasets indicates the validity of our method.
Shiming Xiang, Feiping Nie 0001, Gaofeng Meng, Chunhong Pan, Changshui Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2010 Skew Estimation of Document Images Using Bagging
abstract
This paper proposes a general-purpose method for estimating the skew angles of document images. Rather than to derive a skew angle merely from text lines, the proposed method exploits various types of visual cues of image skew available in local image regions. The visual cues are extracted by Radon transform and then outliers of them are iteratively rejected through a floating cascade. A bagging (bootstrap aggregating) estimator is finally employed to combine the estimations on the local image blocks. Our experimental results show significant improvements against the state-of-the-art methods, in terms of execution speed and estimation accuracy, as well as the robustness to short and sparse text lines, multiple different skews and the presence of nontextual objects of various types and quantities.
Gaofeng Meng, Chunhong Pan, Nanning Zheng 0001
IEEE Trans. Image Process.1
2008 Illumination transition image: Parameter-based illumination estimation and re-rendering
abstract
Varying illumination condition is a challenging problem for face recognition and synthesis. The illumination re-rendering technique allows aligning the illumination effects of facial images or relighting them as expected. In this paper, we propose an improved illumination re-rendering method based on more accurate mapping of facial images in the parametric illumination space. This will make the parameter-based illumination alignment more reliable. A clustering-based criterion is designed to evaluate its parameter estimation precision. To guide the image re-rendering between any a parameter pair, an intermediate image called the illumination transition image (ITI) is defined to represent both the illumination variation information and the person-specific facial shape features. The extensive experimental results verify the proposed method outperforms the quotient image approach on both parameter precision and rendering quality.
Nanning Zheng 0001, Gaofeng Meng, Shaoyi Du
ICPR4
2008 Affine Registration of Point Sets Using ICP and ICA
abstract
This letter proposes a novel algorithm for affine registration of point sets in the way of incorporating an affine transformation into the iterative closest point (ICP) algorithm. At each iterative step of this algorithm, a closed-form solution of the affine transformation is derived. Similar to the ICP algorithm, this new algorithm converges monotonically to a local minimum from any given initial parameters. To get the best affine registration result, good initial parameters are required which are successfully estimated by using independent component analysis (ICA). Experimental results demonstrate the robustness and high accuracy of this algorithm.
Shaoyi Du, Nanning Zheng 0001, Gaofeng Meng, Zejian Yuan
IEEE Signal Process. Lett.3
2008 Shading Extraction and Correction for Scanned Book Images
abstract
When one scans document pages from a bound book, shading artifacts are commonly occurred in the book spine area. In this letter, we propose a general-purpose method for image shading correction based on an assumption that the reflectance function of the page surface is piecewise constant and the illumination function is smooth. The proposed method is able to completely correct more general types of shading artifacts which are nonuniformly distributed along the book spine. Comparison experiments on a synthetic and a variety of real scanned book images demonstrate the feasibility and effectiveness of the proposed method.
Gaofeng Meng, Nanning Zheng 0001, Shaoyi Du, Yonghong Song, Yuanlin Zhang 0001
IEEE Signal Process. Lett.1
2007 Document Images Retrieval Based on Multiple Features Combination
abstract
Retrieving the relevant document images from a great number of digitized pages with different kinds of artificial variations and documents quality deteriorations caused by scanning and printing is a meaningful and challenging problem. We attempt to deal with this problem by combining up multiple different kinds of document features in a hybrid way. Firstly, two new kinds of document image features based on the projection histograms and crossings number histograms of an image are proposed. Secondly, the proposed two features, together with density distribution feature and local binary pattern feature, are combined in a multistage structure to develop a novel document image retrieval system. Experimental results show that the proposed novel system is very efficient and robust for retrieving different kinds of document images, even if some of them are severely degraded.
Gaofeng Meng, Nanning Zheng 0001, Yonghong Song, Yuanlin Zhang 0001
ICDAR1
2007 Circular Noises Removal from Scanned Document Images
abstract
Defects inspection and correction is an important topic in the fields of scanned documents preprocessing. In this paper, a very fast and robust algorithm is proposed for locating and removing a special kind of circular noises caused by scanning documents with punched holes. Firstly, original image is reduced according to an elaborately selected ratio. Punched holes after reduction will leave some distinctive small regions. By examining such small regions, holes noises can be fast detected and located. To diminish false detections, Hough transformation is applied to the roughly located regions to further confirm the located holes. Finally, circular noise is eliminated by fitting a bi-linear blending Coons surface which interpolates along the four edges of noisy region. Experiments on a variety of scanned documents with punched holes demonstrate the feasibility and efficiency of the proposed algorithm.
Gaofeng Meng, Nanning Zheng 0001, Yuanlin Zhang 0001, Yonghong Song
ICDAR1