EDBT 2026 Demo / reviewers in the wild / expert
Anlong Ming
dblp:52/3276
· DBLP profile ↗
74ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0003-2952-7757ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 37 · 1 first-author · 27 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Thinking Aesthetics Assessment of Image Color Temperature: Models, Datasets and BenchmarksabstractColor temperature, as a crucial attribute influencing image color, plays a critical role in Image Aesthetics Assessment (IAA). Yet, within the existing IAA field, little light has been shed on assessing the aesthetic quality of image color temperature. To bridge this gap, we introduce a new task: Image Color Temperature Aesthetics Assessment (ICTAA). However, this task poses the following challenges: 1) Perceptual Sensitivity: humans exhibit high sensitivity to subtle shifts in color temperature, necessitating a model to enable fine-grained discrimination; 2) Spectral Continuity: The theoretical modeling of color temperature aesthetics requires continuous labels; however, the just-noticeable-difference property of human perception makes continuous labeling infeasible, necessitating a well-designed labeling strategy. To address the aforementioned challenges, we make the following efforts. First, we propose a multi-modal contrastive learning framework, ICTA2Net, that models color temperature differences between image pairs while strictly controlling other visual attributes. Second, leveraging color temperature transitivity, we design a weakly supervised strategy that discretely samples images based on anchor images and human perception to build contrastive relations across color temperatures, enabling learning from discrete labels. Thirdly, we construct a color temperature aesthetics dataset, ICTAA240K, and a benchmark for validation. Additionally, we propose a new metric, Information Entropy-weighted Accuracy (IEA), which weights accuracy by the degree of annotation disagreement to reflect model performance across varying sample difficulties, complementing existing evaluation metrics. Experiments show our method outperforms existing state-of-the-art IAA methods on ICTAA240K, thereby setting an effective roadmap for ICTAA. Jinguang Cheng, Taiyu Chen, Anlong Ming |
AAAI | 5 |
| 2026 | Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language ModelsabstractImage Aesthetics Assessment (IAA) evaluates visual quality through user-centered perceptual analysis and can guide various applications. Recent advances in Multimodal Large Language Models (MLLMs) have sparked interest in adapting them for IAA. However, two critical limitations persist in applying MLLMs to IAA: 1) the tokenization strategy leads to insensitivity to scores, and 2) the classification-based decoding mechanisms introduce score quantization errors. Current MLLM-based IAA methods treat the task as coarse rating classification followed by probability-to-score mapping, which loses fine-grained information. To address these challenges, we propose ROC4MLLM, offering complementary solutions from two perspectives:1) Representation: We separate scores from the word token space to avoid tokenizing scores as text. An independent position token bridges these spaces, improving the sensitivity of the model to score positions in text. 2) Computation: We apply distinct loss functions for text and score predictions to enhance the sensitivity of the model to score gradients. Decoupling scores from text ensures effective supervision while preventing interference between scores and text in the loss computation. Extensive experiments across five datasets demonstrate that ROC4MLLM achieves state-of-the-art performance without requiring additional training data. Additionally, its plug-and-play design ensures seamless integration with existing MLLMs, boosting their IAA performance. Xingyuan Ma, Anlong Ming, Haobin Zhong, Huadong Ma |
AAAI | 3 |
| 2026 | Direction-aware deep policy learning for efficient capacitated arc routing
Feng Xue 0001, Runze Guo, Anlong Ming, Nicu Sebe |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Exercise quality assessment in monocular video streaming
Boxuan Xu, Zhaowen Lin, Anlong Ming |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | ELTA 2.0: Rethinking Long-Tail for Image Aesthetics Assessment
Anlong Ming, Huadong Ma |
Int. J. Comput. Vis. | 4 |
| 2026 | M2Beats 2.0: When motion meets beats in short-form videos twice
Ao Lv, Anlong Ming |
Pattern Recognit. | 5 |
| 2025 | Rethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An ExemplificationabstractThe Personalized Aesthetics Assessment (PAA) aims to accurately predict an individual’s unique perception of aesthetics. With the surging demand for customization, PAA enables applications to generate personalized outcomes by aligning with individual aesthetic preferences. The prevailing PAA paradigm involves two stages: pre-training and fine-tuning, but it faces three inherent challenges: 1) The model is pre-trained using datasets of the Generic Aesthetics Assessment (GAA), but the collective preferences of GAA lead to conflicts in individualized aesthetic predictions. 2) The scope and stage of personalized surveys are related to both the user and the assessed object; however, the prevailing personalized surveys fail to adequately address assessed objects’ characteristics. 3) During application usage, the cumulative multimodal feedback from an individual holds great value that should be considered for improving the PAA model but unfortunately attracts insufficient attention. To address the aforementioned challenges, we introduce a new PAA paradigm called PAA+, which is structured into three distinct stages: pre-training, fine-tuning, and continual learning. Furthermore, to better reflect individual differences, we employ a familiar and intuitive application, physique aesthetics assessment (PhysiqueAA), to validate the PAA+ paradigm. We propose a dataset called PhysiqueAA50K, consisting of over 50,000 annotated physique images. Furthermore, we develop a PhysiqueAA framework (PhysiqueFrame) and conduct a large-scale benchmark, achieving state-of-the-art (SOTA) performance. Our research is expected to provide an innovative roadmap and application for the PAA community. The code and dataset are available in here. Haobin Zhong, Anlong Ming, Huadong Ma |
CVPR | 3 |
| 2025 | IFS-Light: An Interactive Framework for Single-view Face Relighting with both Facial and Lighting Consistency
Anlong Ming |
ACM Multimedia | 3 |
| 2025 | HDR-NFlow: High Dynamic Range imaging with normalizing flow
Shuaikang Shang, Xuejing Kang, Anlong Ming |
Neurocomputing | 3 |
| 2025 | DA3Attacker: A Diffusion-Based Attacker Against Aesthetics-Oriented Black-Box ModelsabstractThe adage "Beautiful Outside But Ugly Inside" resonates with the security and explainability challenges encountered in image aesthetics assessment (IAA). Although deep neural networks (DNNs) have demonstrated remarkable performance in various IAA tasks, how to probe, explain, and enhance aesthetics-oriented "black-box" models has not yet been investigated to our knowledge. This lack of investigation has significantly impeded the commercial application of IAA. In this paper, we investigate the susceptibility of current IAA models to adversarial attacks and aim to elucidate the underlying mechanisms that contribute to their vulnerabilities. To address this, we propose a novel diffusion-based framework as an attacker (DA3Attacker), capable of generating adversarial examples (AEs) to deceive diverse black-box IAA models. DA3Attacker employs a dedicated Attack Diffusion Transformer, equipped with modular aesthetics-oriented filters. By undergoing two unsupervised training stages, it constructs a latent space to generate AEs and facilitates two distinct yet controllable attack modes: restricted and unrestricted. Extensive experiments on 26 baseline models demonstrate that our method effectively explores the vulnerabilities of these IAA models, while also providing multi-attribute explanations for their feature dependencies. To facilitate further research, we contribute the evaluation tools and four metrics for measuring adversarial robustness, as well as a dataset of 60,000 re-labeled AEs for fine-tuning IAA models. The resources are available here. Shuntian Zheng, Anlong Ming, Yanni Wang, Huadong Ma |
IEEE Trans. Image Process. | 3 |
| 2025 | Evidence-Based Real-Time Road Segmentation With RGB-D Data AugmentationabstractDespite significant progress in RGB-D based road segmentation in recent years, the latest methods cannot achieve both state-of-the-art accuracy and real time due to the high-performance reliance on heavy structures. We argue that this reliance is due to unsuitable multimodal fusion. To be specific, RGB and depth data in road scenes are each sensitive to different regions, but current RGB-D based road segmentation methods generally combine features within sensitive regions which preserves false road representation from one of the data. Based on such findings, we design an Evidence-based Road Segmentation Method (Evi-RoadSeg), which incorporates prior knowledge of the modal-specific characteristics. Firstly, we abandon the cross-modal fusion operation commonly used in existing multimodal based methods. Instead, we collect the road evidence from RGB and depth inputs separately via two low-latency subnetworks, and fuse the road representation of the two subnetworks by taking both modalities’ evidence as a measure of confidence. Secondly, we propose an RGB-D data augmentation scheme tailored to road scenes to enhance the unique properties of RGB and depth data. It facilitates learning by adding more sensitive regions to the samples. Finally, the proposed method is evaluated on the widely used KITTI-road, ORFD, and R2D datasets. Our method achieves state-of-the-art accuracy at over 70 FPS, 5$\times$faster than comparable RGB-D methods. Furthermore, extensive experiments illustrate that our method can be deployed on a Jetson Nano 2GB with a speed of 8$+$FPS. The code will be released in https://github.com/xuefeng-cvr/Evi-RoadSeg. Feng Xue 0001, Yicong Chang, Wenzhuang Xu, Wenteng Liang, Fei Sheng, Anlong Ming |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | ELTA: An Enhancer against Long-Tail for Aesthetics-oriented ModelsabstractReal-world datasets often exhibit long-tailed distributions, compromising the generalization and fairness of learning-based models. This issue is particularly pronounced in Image Aesthetics Assessment (IAA) tasks, where such imbalance is difficult to mitigate due to a severe distribution mismatch between features and labels, as well as the great sensitivity of aesthetics to image variations. To address these issues, we propose an Enhancer against Long-Tail for Aesthetics-oriented models (ELTA). ELTA first utilizes a dedicated mixup technique to enhance minority feature representation in high-level space while preserving their intrinsic aesthetic qualities. Next, it aligns features and labels through a similarity consistency approach, effectively alleviating the distribution mismatch. Finally, ELTA adopts a specific strategy to refine the output distribution, thereby enhancing the quality of pseudo-labels. Experiments on four representative datasets (AVA, AADB, TAD66K, and PARA) show that our proposed ELTA achieves state-of-the-art performance by effectively mitigating the long-tailed issue in IAA datasets. Moreover, ELTA is designed with plug-and-play capabilities for seamless integration with existing methods. To our knowledge, this is the first contribution in the IAA community addressing long-tail. All resources are available in here. Anlong Ming, Huadong Ma |
ICML | 3 |
| 2024 | M2Beats: When Motion Meets Beats in Short-form Videos
Dongxiang Jiang, Anlong Ming |
IJCAI | 4 |
| 2024 | AK4Prompts: Aesthetics-driven Automatically Keywords-Ranking for Prompts in Text-To-Image Models
Mengchao Wang, Anlong Ming |
IJCAI | 4 |
| 2024 | Thinking Temporal Automatic White Balance: Datasets, Models and Benchmarks
Xuejing Kang, Anlong Ming |
ACM Multimedia | 4 |
| 2024 | SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One ModelabstractIn the last year, universal monocular metric depth estimation (universal MMDE) has gained considerable attention, serving as the foundation model for various multimedia tasks, such as video and image editing. Nonetheless, current approaches face challenges in maintaining consistent accuracy across diverse scenes without scene-specific parameters and pre-training, hindering the practicality of MMDE. Furthermore, these methods rely on extensive datasets comprising millions, if not tens of millions, of data for training, leading to significant time and hardware expenses. This paper presents SM4Depth, a model that seamlessly works for both indoor and outdoor scenes, without needing extensive training data and GPU clusters. Firstly, to obtain consistent depth across diverse scenes, we propose a novel metric scale modeling, i.e., variation- based unnormalized depth bins. It reduces the ambiguity of the conventional metric bins and enables better adaptation to large depth gaps of scenes during training. Secondly, we propose a ''divide and conquer'' solution to reduce reliance on massive training data. Instead of estimating directly from the vast solution space, the metric bins are estimated from multiple solution sub-spaces to reduce complexity. Additionally, we introduce an uncut depth dataset, BUPT Depth, to evaluate the depth accuracy and consistency across various indoor and outdoor scenes. Trained on a consumer-grade GPU using just 150K RGB-D pairs, SM4Depth achieves outstanding performance on the most never-before-seen datasets, especially maintaining consistent accuracy across indoors and outdoors. The code can be found here. Feng Xue 0001, Anlong Ming, Mingshuai Zhao, Huadong Ma, Nicu Sebe |
ACM Multimedia | 3 |
| 2024 | "Special Relativity" of Image Aesthetics Assessment: a Preliminary Empirical PerspectiveabstractImage aesthetics assessment (IAA) primarily examines image quality from a user-centric perspective and can be applied to guide various applications, including image capture, recommendation, and enhancement. The fundamental issue in IAA revolves around the quantification of image aesthetics. Existing methodologies rely on assigning a scalar (or a distribution) to represent aesthetic value based on conventional practices, which confines this scalar within a specific range and artificially labels it. However, conventional methods rarely incorporate research on interpretability, particularly lacking systematic responses to the following three fundamental questions: 1) Can aesthetic qualities be quantified? 2) What is the nature of quantifying aesthetics? 3) How can aesthetics be accurately quantified? In this paper, we present a law called "Special Relativity" of IAA (SR-IAA) that addresses the aforementioned core questions. We have developed a Multi-Attribute IAA Framework (MAINet), which serves as a preliminary validation for SR-IAA within the existing datasets and achieves state-of-the-art (SOTA) performance. Specifically, our metrics on multi-attribute assessment outperform the second-best performance by 8.06% (AADB), 1.67% (PARA), and 2.44% (SPAQ) in terms of SRCC. We anticipate that our research will offer innovative theoretical guidance to the IAA research community. All resources are available here. Anlong Ming, Huadong Ma |
ACM Multimedia | 2 |
| 2024 | Rethinking No-reference Image Exposure Assessment from Holism to Pixel: Models, Datasets and BenchmarksabstractThe past decade has witnessed an increasing demand for enhancing image quality through exposure, and as a crucial prerequisite in this endeavor, Image Exposure Assessment (IEA) is now being accorded serious attention. However, IEA encounters two persistent challenges that remain unresolved over the long term: the accuracy and generalizability of No-reference IEA are inadequate for practical applications; the scope of IEA is confined to qualitative and quantitative analysis of the entire image or subimage, such as providing only a score to evaluate the exposure level, thereby lacking intuitive and precise fine-grained evaluation for complex exposure conditions. The objective of this paper is to address the persistent bottleneck challenges from three perspectives: model, dataset, and benchmark. 1) Model-level: we propose a Pixel-level IEA Network (P-IEANet) that utilizes Haar discrete wavelet transform (DWT) to analyze, decompose, and assess exposure from both lightness and structural perspectives, capable of generating pixel-level assessment results under no-reference scenarios. 2) Dataset-level: we elaborately build an exposure-oriented dataset, IEA40K, containing 40K images, covering 17 typical lighting scenarios, 27 devices, and 50+ scenes, with each image densely annotated by more than 10 experts with pixel-level labels. 3) Benchmark-level: we develop a comprehensive benchmark of 19 methods based on IEA40K. Our P-IEANet not only achieves state-of-the-art (SOTA) performance on all metrics but also seamlessly integrates with existing exposure correction and lighting enhancement methods. To our knowledge, this is the first work that explicitly emphasizes assessing complex image exposure problems at a pixel level, providing a significant boost to the IEA and exposure-related community. The code and dataset are available in \href{https://github.com/mRobotit/Pixel-level-No-reference-Image-Exposure-Assessment}{\textcolor{red} {here}}. Shuntian Zheng, Anlong Ming, Banyu Wu, Huadong Ma |
NeurIPS | 3 |
| 2024 | Indoor Obstacle Discovery on Reflective Ground via Monocular Camera
Feng Xue 0001, Yicong Chang, Tianxi Wang, Yu Zhou 0016, Anlong Ming |
Int. J. Comput. Vis. | 5 |
| 2023 | SWBNet: A Stable White Balance Network for sRGB ImagesabstractThe white balance methods for sRGB images (sRGB-WB) aim to directly remove their color temperature shifts. Despite achieving promising white balance (WB) performance, the existing methods suffer from WB instability, i.e., their results are inconsistent for images with different color temperatures. We propose a stable white balance network (SWBNet) to alleviate this problem. It learns the color temperature-insensitive features to generate white-balanced images, resulting in consistent WB results. Specifically, the color temperatureinsensitive features are learned by implicitly suppressing lowfrequency information sensitive to color temperatures. Then, a color temperature contrastive loss is introduced to facilitate the most information shared among features of the same scene and different color temperatures. This way, features from the same scene are more insensitive to color temperatures regardless of the inputs. We also present a color temperature sensitivity-oriented transformer that globally perceives multiple color temperature shifts within an image and corrects them by different weights. It helps to improve the accuracy of stabilized SWBNet, especially for multiillumination sRGB images. Experiments indicate that our SWBNet achieves stable and remarkable WB performance. Xuejing Kang, Anlong Ming |
AAAI | 4 |
| 2023 | Unknown Sniffer for Object Detection: Don't Turn a Blind Eye to Unknown ObjectsabstractThe recently proposed open-world object and open-set detection have achieved a breakthrough in finding never-seen-before objects and distinguishing them from known ones. However, their studies on knowledge transfer from known classes to unknown ones are not deep enough, resulting in the scanty capability for detecting unknowns hidden in the background. In this paper, we propose the unknown sniffer (UnSniffer) to find both unknown and known objects. Firstly, the generalized object confidence (GOC) score is introduced, which only uses known samples for supervision and avoids improper suppression of unknowns in the back-ground. Significantly, such confidence score learned from known objects can be generalized to unknown ones. Additionally, we propose a negative energy suppression loss to further suppress the non-object samples in the background. Next, the best box of each unknown is hard to obtain during inference due to lacking their semantic information in training. To solve this issue, we introduce a graph-based determination scheme to replace hand-designed non-maximum suppression (NMS) post-processing. Finally, we present the Unknown Object Detection Benchmark, the first publicly benchmark that encompasses precision evaluation for unknown detection to our knowledge. Experiments show that our method is far better than the existing state-of-the-art methods. Code is available at: https://github.com/Went-Liang/UnSniffer. Wenteng Liang, Feng Xue 0001, Guofeng Zhong, Anlong Ming |
CVPR | 5 |
| 2023 | Thinking Image Color Aesthetics Assessment: Models, Datasets and BenchmarksabstractWe present a comprehensive study on a new task named image color aesthetics assessment (ICAA), which aims to assess color aesthetics based on human perception. ICAA is important for various applications such as imaging measurement and image analysis. However, due to the highly diverse aesthetic preferences and numerous color combinations, ICAA presents more challenges than conventional image quality assessment tasks. To advance ICAA research, 1) we propose a baseline model called the Delegate Transformer, which not only deploys deformable transformers to adaptively allocate interest points, but also learns human color space segmentation behavior by the dedicated module. 2) We elaborately build a color-oriented dataset, ICAA17K, containing 17K images, covering 30 popular color combinations, 80 devices and 50 scenes, with each image densely annotated by more than 1,500 people. Moreover, we develop a large-scale benchmark of 15 methods, the most comprehensive one thus far based on two datasets, SPAQ and ICAA17K. Our work, not only achieves state-of-the-art performance, but more importantly offers the community a roadmap to explore solutions for ICAA. Code and dataset are available in here. Anlong Ming, Jinyuan Sun, Shuntian Zheng, Huadong Ma |
ICCV | 2 |
| 2023 | ICDA: Illumination-Coupled Domain Adaptation Framework for Unsupervised Nighttime Semantic SegmentationabstractThe performance of nighttime semantic segmentation has been significantly improved thanks to recent unsupervised methods. However, these methods still suffer from complex domain gaps, i.e., the challenging illumination gap and the inherent dataset gap. In this paper, we propose the illumination-coupled domain adaptation framework(ICDA) to effectively avoid the illumination gap and mitigate the dataset gap by coupling daytime and nighttime images as a whole with semantic relevance. Specifically, we first design a new composite enhancement method(CEM) that considers not only illumination but also spatial consistency to construct the source and target domain pairs, which provides the basic adaptation unit for our ICDA. Next, to avoid the illumination gap, we devise the Deformable Attention Relevance(DAR) module to capture the semantic relevance inside each domain pair, which can couple the daytime and nighttime images at the feature level and adaptively guide the predictions of nighttime images. Besides, to mitigate the dataset gap and acquire domain-invariant semantic relevance, we propose the Prototype-based Class Alignment(PCA) module, which improves the usage of category information and performs fine-grained alignment. Extensive experiments show that our method reduces the complex domain gaps and achieves state-of-the-art performance for nighttime semantic segmentation. Our code is available at https://github.com/chenghaoDong666/ICDA. Chenghao Dong, Xuejing Kang, Anlong Ming |
IJCAI | 3 |
| 2023 | WBFlow: Few-shot White Balance for sRGB Images via Reversible Neural FlowsabstractThe sRGB white balance methods aim to correct the nonlinear color cast of sRGB images without accessing raw values. Although existing methods have achieved increasingly better results, their generalization to sRGB images from multiple cameras is still under explored. In this paper, we propose the network named WBFlow that not only performs superior white balance for sRGB images but also generalizes well to multiple cameras. Specifically, we take advantage of neural flow to ensure the reversibility of WBFlow, which enables lossless rendering of color cast sRGB images back to pseudo raw features for linear white balancing and thus achieves superior performance. Furthermore, inspired by camera transformation approaches, we have designed a camera transformation (CT) in pseudo raw feature space to generalize WBFlow for different cameras via few shot learning. By utilizing a few sRGB images from an untrained camera, our WBFlow can perform well on this camera by learning the camera specific parameters of CT. Extensive experiments show that WBFlow achieves superior camera generalization and accuracy on three public datasets as well as our rendered multiple camera sRGB dataset. Our code is available at https://github.com/ChunxiaoLe/WBFlow. Xuejing Kang, Anlong Ming |
IJCAI | 3 |
| 2023 | EAT: An Enhancer for Aesthetics-Oriented TransformersabstractTransformers have shown great potential in various vision tasks, but none of them have surpassed the best CNN model on image aesthetics assessment (IAA) tasks. IAA is a challenging task in multimedia systems that requires attention to both foreground and background, as well as robustness to noisy and redundant labels. The global and dense attention mechanism of Transformers, designed for saliency-oriented tasks, may miss important aesthetic information in the background, increase the computational cost and slow down the convergence on IAA tasks. To address these issues, we propose an Enhancer for Aesthetics-Oriented Transformers (EAT). EAT uses a deformable, sparse and data-dependent attention mechanism that learns where to focus and how to refine attention by offsets. EAT also guides the offsets to balance the attention between foreground and background according to dedicated rules. Our EAT-enhanced Transformers outperform the previous methods on four representative datasets with fewer training epochs. Code is available in https://github.com/woshidandan/Image-Aesthetics-Assessment Anlong Ming, Shuntian Zheng, Haobin Zhong, Huadong Ma |
ACM Multimedia | 2 |
| 2023 | MARF: Multiscale Adaptive-Switch Random Forest for Leg Detection With 2-D Laser ScannersabstractFor the 2-D laser-based tasks, e.g., people detection and people tracking, leg detection is usually the first step. Thus, it carries great weight in determining the performance of people detection and people tracking. However, many leg detectors ignore the inevitable noise and the multiscale characteristics of the laser scan, which makes them sensitive to the unreliable features of point cloud and further degrades the performance of the leg detector. In this article, we propose a multiscale adaptive-switch random forest (MARF) to overcome these two challenges. First, the adaptive-switch decision tree is designed to use noise-sensitive features to conduct weighted classification and noise-invariant features to conduct binary classification, which makes our detector perform more robust to noise. Second, considering the multiscale property that the sparsity of the 2-D point cloud is proportional to the length of laser beams, we design a multiscale random forest structure to detect legs at different distances. Moreover, the proposed approach allows us to discover a sparser human leg from point clouds than others. Consequently, our method shows an improved performance compared to other state-of-the-art leg detectors on the challenging Moving Legs dataset and retains the entire pipeline at a speed of 60+ FPS on low-computational laptops. Moreover, we further apply the proposed MARF to the people detection and tracking system, achieving a considerable gain in all metrics. Tianxi Wang, Feng Xue 0001, Yu Zhou 0016, Anlong Ming |
IEEE Trans. Cybern. | 4 |
| 2023 | Vine Spread for Superpixel SegmentationabstractSuperpixel is the over-segmentation region of an image, whose basic units "pixels" have similar properties. Although many popular seeds-based algorithms have been proposed to improve the segmentation quality of superpixels, they still suffer from the seeds initialization problem and the pixel assignment problem. In this paper, we propose Vine Spread for Superpixel Segmentation (VSSS) to form superpixel with high quality. First, we extract image color and gradient features to define the soil model that establishes a "soil" environment for vine, and then we define the vine state model by simulating the vine "physiological" state. Thereafter, to catch more image details and twigs of the object, we propose a new seeds initialization strategy that perceives image gradients at the pixel-level and without randomness. Next, to balance the boundary adherence and the regularity of the superpixel, we define a three-stage "parallel spreading" vine spread process as a novel pixel assignment scheme, in which the proposed nonlinear velocity for vines helps to form the superpixel with regular shape and homogeneity, the crazy spreading mode for vines and the soil averaging strategy help to enhance the boundary adherence of superpixel. Finally, a series of experimental results demonstrate that our VSSS offers competitive performance in the seed-based methods, especially in catching object details and twigs, balancing boundary adherence and obtaining regular shape superpixels. Xuejing Kang, Anlong Ming |
IEEE Trans. Image Process. | 3 |
| 2022 | Transfer Learning for Color Constancy via Statistic PerspectiveabstractColor Constancy aims to correct image color casts caused by scene illumination. Recently, although the deep learning approaches have remarkably improved on single-camera data, these models still suffer from the seriously insufficient data problem, resulting in shallow model capacity and degradation in multi-camera settings. In this paper, to alleviate this problem, we present a Transfer Learning Color Constancy (TLCC) method that leverages cross-camera RAW data and massive unlabeled sRGB data to support training. Specifically, TLCC consists of the Statistic Estimation Scheme (SE-Scheme) and Color-Guided Adaption Branch (CGA-Branch). SE-Scheme builds a statistic perspective to map the camera-related illumination labels into camera-agnostic form and produce pseudo labels for sRGB data, which greatly expands data for joint training. Then, CGA-Branch further promotes efficient transfer learning from sRGB to RAW data by extracting color information to regularize the backbone's features adaptively. Experimental results show the TLCC has overcome the data limitation and model degradation, outperforming the state-of-the-art performance on popular benchmarks. Moreover, the experiments also prove the TLCC is capable of learning new scenes information from sRGB data to improve accuracy on the RAW images with similar scenes. Xuejing Kang, Zhaowen Lin, Anlong Ming |
AAAI | 5 |
| 2022 | Fast Road Segmentation via Uncertainty-aware Symmetric NetworkabstractThe high performance of RGB-D based road segmentation methods contrasts with their rare application in commercial autonomous driving, which is owing to two reasons: 1) the prior methods cannot achieve high inference speed and high accuracy in both ways; 2) the different properties of RGB and depth data are not well-exploited, limiting the reliability of predicted road. In this paper, based on the evidence theory, an uncertainty-aware symmetric network (USNet) is proposed to achieve a trade-off between speed and accuracy by fully fusing RGB and depth data. Firstly, cross-modal feature fusion operations, which are indispensable in the prior RGB-D based methods, are abandoned. We instead separately adopt two light-weight subnetworks to learn road representations from RGB and depth inputs. The light-weight structure guarantees the real-time inference of our method. Moreover, a multi-scale evidence collection (MEC) module is designed to collect evidence in multiple scales for each modality, which provides sufficient evidence for pixel class determination. Finally, in uncertainty-aware fusion (UAF) module, the uncertainty of each modality is perceived to guide the fusion of the two sub-networks. Experimental results demonstrate that our method achieves a state-of-the-art accuracy with real-time inference speed of$43+$FPS. The source code is available at https://github.com/morancyc/USNet. Yicong Chang, Feng Xue 0001, Fei Sheng, Wenteng Liang, Anlong Ming |
ICRA | 5 |
| 2022 | Monocular Depth Distribution Alignment with Low ComputationabstractThe performance of monocular depth estimation generally depends on the amount of parameters and computational cost. It leads to a large accuracy contrast between light-weight networks and heavy-weight networks, which limits their application in the real world. In this paper, we model the majority of accuracy contrast between them as the difference of depth distribution, which we call 'Distribution drift'. To this end, a distribution alignment network (DANet) is proposed. We firstly design a pyramid scene transformer (PST) module to capture inter-region interaction in multiple scales. By perceiving the difference of depth features between every two regions, DANet tends to predict a reasonable scene structure, which fits the shape of distribution to ground truth. Then, we propose a local-global optimization (LGO) scheme to realize the supervision of global range of scene depth. Thanks to the alignment of depth distribution shape and scene depth range, DANet sharply alleviates the distribution drift, and achieves a comparable performance with prior heavy-weight methods, but uses only 1% floating-point operations per second (FLOPs) of them. The experiments on two datasets, namely the widely used NYUDv2 dataset and the more challenging iBims-1 dataset, demonstrate the effectiveness of our method. The source code is available at https://github.com/YiLiM1/DANet. Fei Sheng, Feng Xue 0001, Yicong Chang, Wenteng Liang, Anlong Ming |
ICRA | 5 |
| 2022 | Rethinking Image Aesthetics Assessment: Models, Datasets and BenchmarksabstractChallenges in image aesthetics assessment (IAA) arise from that images of different themes correspond to different evaluation criteria, and learning aesthetics directly from images while ignoring the impact of theme variations on human visual perception inhibits the further development of IAA; however, existing IAA datasets and models overlook this problem. To address this issue, we show that a theme-oriented dataset and model design are effective for IAA. Specifically, 1) we elaborately build a novel dataset, called TAD66K, that contains 66K images covering 47 popular themes, and each image is densely annotated by more than 1200 people with dedicated theme evaluation criteria. 2) We develop a baseline model, TANet, which can effectively extract theme information and adaptively establish perception rules to evaluate images with different themes. 3) We develop a large-scale benchmark (the most comprehensive thus far) by comparing 17 methods with TANet on three representative datasets: AVA, FLICKR-AES and the proposed TAD66K, TANet achieves state-of-the-art performance on all three datasets. Our work offers the community an opportunity to explore more challenging directions; the code, dataset and supplementary material are available at https://github.com/woshidandan/TANet. Dongxiang Jiang, Anlong Ming |
IJCAI | 5 |
| 2022 | Domain Adversarial Learning for Color ConstancyabstractColor Constancy aims to eliminate the color cast of RAW images caused by non-neutral illuminants. Though contemporary approaches based on convolutional neural networks significantly improve illuminant estimation, they suffer from the seriously insufficient data problem. To solve this problem by effectively utilizing multi-domain data, we propose the Domain Adversarial Learning Color Constancy (DALCC) which consists of the Domain Adversarial Learning Branch (DALB) and the Feature Reweighting Module (FRM). In DALB, the Camera Domain Classifier and the feature extractor compete against each other in an adversarial way to encourage the emergence of domain-invariant features. At the same time, the Illuminant Transformation Module performs color space conversion to solve the inconsistent color space problem caused by those domain-invariant features. They collaboratively avoid model degradation of multi-device training caused by the domain discrepancy of feature distribution, which enables our DALCC to benefit from multi-domain data. Besides, to better utilize multi-domain data, we propose the FRM that reweights the feature map to suppress Non-Primary Illuminant regions, which reduces the influence of misleading illuminant information. Experiments show that the proposed DALCC can more effectively take advantage of multi-domain data and thus achieve state-of-the-art performance on commonly used benchmark datasets. Xuejing Kang, Anlong Ming |
IJCAI | 3 |
| 2022 | Explored Normalized Cut With Random Walk Refining Term for Image SegmentationabstractThe Normalized Cut (NCut) model is a popular graph-based model for image segmentation. But it suffers from the excessive normalization problem and weakens the small object and twig segmentation. In this paper, we propose an Explored Normalized Cut (ENCut) model that establishes a balance graph model by adopting a meaningful-loop and a k-step random walk, which reduces the energy of small salient region, so as to enhance the small object segmentation. To improve the twig segmentation, our ENCut model is further enhanced by a new Random Walk Refining Term (RWRT) that adds local attention to our model with the help of an un-supervising random walk. Finally, a move-making based strategy is developed to efficiently solve the ENCut model with RWRT. Experiments on three standard datasets indicate that our model can achieve state-of-the-art results among the NCut-based segmentation models. Lei Zhu 0012, Xuejing Kang, Lizhu Ye, Anlong Ming |
IEEE Trans. Image Process. | 4 |
| 2021 | MT-ORL: Multi-Task Occlusion Relationship LearningabstractRetrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely occlusion boundary extraction and occlusion orientation prediction, and secondly, improper representation of occlusion orientation. In this paper, we propose a novel architecture called Occlusion-shared and Path-separated Network (OPNet), which solves the first issue by exploiting rich occlusion cues in shared high-level features and structured spatial information in task-specific low-level features. We then design a simple but effective orthogonal occlusion representation (OOR) to tackle the second issue. Our method surpasses the state-of-the-art methods by 6.1%/8.3% Boundary-AP and 6.5%/10% Orientation-AP on standard PIOD/BSDS ownership datasets. Code is available at https://github.com/fengpanhe/MT-ORL. Panhe Feng, Qi She, Lei Zhu 0012, Lin Zhang 0040, Zijian Feng, Changhu Wang, Chunpeng Li, Xuejing Kang, Anlong Ming |
ICCV | 10 |
| 2021 | HDA-Net: Horizontal Deformable Attention Network for Stereo MatchingabstractStereo matching is a fundamental and challenging task which has various applications in autonomous driving, dense reconstruction and other depth related tasks. Contextual information with discriminative features is crucial for accurate stereo matching in the ill-posed regions (textureless, occlusion, etc.). In this paper, we propose an efficient horizontal attention module to adaptively capture the global correspondence clues. Compared with the popular non-local attention, our horizontal attention is more effective for stereo matching with better performance and lower consumption of computation and memory. We further introduce a deformable module to refine the contextual information in the disparity discontinuous areas such as the boundary of objects. Learning-based method is adopted to construct the cost volume by concatenating the features of two branches. In order to offer explicit similarity measure to guide learning-based volume for obtaining more reasonable unimodal matching cost distribution we additionally combine the learning-based volume with the improved zero-centered group-wise correlation volume. Finally, we regularize the 4D joint cost volume by a 3D CNN module and generate the final output by disparity regression. The experimental results show that our proposed HDA-Net achieves the state-of-the-art performance on the Scene Flow dataset and obtains competitive performance on the KITTI datasets compared with the relevant networks. Xuesong Zhang 0001, Anlong Ming |
ACM Multimedia | 5 |
| 2021 | HDNet: Hybrid Distance Network for semantic segmentation
Chunpeng Li, Xuejing Kang, Lei Zhu 0012, Lizhu Ye, Panhe Feng, Anlong Ming |
Neurocomputing | 6 |
| 2021 | Boundary-induced and scene-aggregated network for monocular depth prediction
Feng Xue 0001, Junfeng Cao, Yu Zhou 0016, Fei Sheng, Anlong Ming |
Pattern Recognit. | 6 |
| 2020 | High Accuracy Compressive Chromo-Tomography Reconstruction via Convolutional Sparse CodingabstractOver the last decade various compressive snapshot hyperspectral imaging methods have been proposed. The limited reconstruction quality from severely compressed measurements, however, has been a practical barrier to real applications. This paper proposes a compressive chromo-tomography framework that incorporates the convolutional sparse coding (CSC) prior into the classical total variation and L1regularization functionals. Such a combination allows excellent high-frequency recovery capabilities of CSC, while effectively suppressing ghost artifacts in tomographic reconstructions. Since nondifferentiable regularizers are employed, we propose a preconditioned alternating direction method of multipliers (ADMM) for flexible and efficient solutions, both for the reconstruction task and for hyperspectral convolutional dictionary learning. We demonstrate in our numerical experiments that just 25 learned 3D CSC filters can fulfill a rather effective hyperspectral imagery representation and that the proposed method is capable of high accuracy reconstructions. Xuesong Zhang 0001, Jing Jiang 0017, Anlong Ming |
ICME | 6 |
| 2020 | BP-net: deep learning-based superpixel segmentation for RGB-D imageabstractIn this paper, we propose a deep learning-based su-perpixel segmentation algorithm for RGB-D image. The proposed deep neural network called BP-net is composed of boundary detection network (B-net) that exploits multiscale information from depth image to extract the geometry edge of objects, and pixel labeling network (P-net) that extracts pixel features and generates superpixels. A boundary pass filter is proposed to combine the edge information and pixel features and ensures superpixels adhere better to geometry edges. To generate regular superpixels, we design a loss function which takes the shape regularity error and superpixel accuracy into account. In addition, for providing reasonable initial seeds, a new seeds initialization strategy is proposed, in which the density of seeds is investigated from a 2-manifolds space to reduce the number of superpixels that cover multiple objects in the region of rich texture. Experimental results demonstrate that our algorithm outperforms the existing state-of-the-art algorithms in terms of accuracy and shape regularity on the RGB-D dataset. Xuejing Kang, Anlong Ming |
ICPR | 3 |
| 2020 | Dynamic Random Walk for Superpixel SegmentationabstractIn this paper, we propose a novel random walk model, called Dynamic Random Walk (DRW), which adds a new type of dynamic node to the original RW model and reduces redundant calculation by limiting the walk range. To solve the seed-lacking problem of the proposed DRW, we redefine the energy function of the original RW and use the first arrival probability among each node pair to avoid the interference for each partition. Relaxation of our DRW is performed with the help of a greedy strategy and the Weighted Random Walk Entropy(WRWE) that uses the gradient feature to approximate the stationary distribution. The proposed DRW not only can enhance the boundary adherence but also can run with linear time complexity. To extend our DRW for superpixel segmentation, a seed initialization strategy is proposed. It can evenly distribute seeds in both 2D and 3D space and generate superpixels in only one iteration. The experimental results demonstrate that our DRW is faster than existing RW models and better than the state-of-the-art superpixel segmentation algorithms with respect to both efficiency and segmentation effects. Xuejing Kang, Lei Zhu 0012, Anlong Ming |
IEEE Trans. Image Process. | 3 |
| 2020 | Tiny Obstacle Discovery by Occlusion-Aware Multilayer RegressionabstractEdges are the fundamental visual element for discovering tiny obstacles using a monocular camera. Nevertheless, tiny obstacles often have weak and inconsistent edge cues due to various properties such as small size and similar appearance to the free space, making it hard to capture them. To this end, we propose an occlusion-based multilayer approach, which specifies the scene prior as multilayer regions and utilizes these regions in each obstacle discovery module, i.e., edge detection and proposal extraction. Firstly, an obstacle-aware occlusion edge is generated to accurately capture the obstacle contour by fusing the edge cues inside all the multilayer regions, which intensifies the object characteristics of these obstacles. Then, a multistride sliding window strategy is proposed for capturing proposals that enclose the tiny obstacles as completely as possible. Moreover, a novel obstacle-aware regression model is proposed for effectively discovering obstacles. It is formed by a primary-secondary regressor, which can learn two dissimilarities between obstacles and other categories separately, and eventually generate an obstacle-occupied probability map. The experiments are conducted on two datasets to demonstrate the effectiveness of our approach under different scenarios. And the results show that the proposed method can approximately improve accuracy by 19% over FPHT and PHT, and achieves comparable performance to MergeNet. Furthermore, multiple experiments with different variants validate the contribution of our method. The source code is available at https://github.com/XuefengBUPT/TOD_OMR. Feng Xue 0001, Anlong Ming, Yu Zhou 0016 |
IEEE Trans. Image Process. | 2 |
| 2019 | A Novel Super-resolution Method Based on Patch Reconstruction with Simk Clustering and Nonlinear MappingabstractIn this paper, we propose a patch-wise super-resolution (SR) method that combines an external-sample classification tree and a nonlinear-mapping learning stage to simultaneously guarantees reconstruction quality and speed at the stage of patch representation and mapping. We use the low-resolution (LR) to high-resolution (HR) mapping kernel of each patch-pair sample (called SIMK) to complete classification by binary tree branching and provide reasonable training sets for mapping-learning. Then a high accuracy but low cost lightweight network is learned for each tree node to choose the reasonable branch path for the testing LR patches. In the mapping-learning stage, the nonlinear mapping for each class is represented as a full-connected network, which provides satisfying generalization ability for LR patch reconstruction. Comparing with state-of-the-art methods, our approach achieves real-time (>24fps) SR of realistic vision and high quality for different upscaling factors. Anlong Ming, Xuesong Zhang 0001, Xuejing Kang |
ICASSP | 2 |
| 2019 | Spatio-spectral Modulation Using a Binary Photomask for Compressive ChromotomographyabstractRecent advances in compressive spectral imagers have demonstrated the potential of spatio-spectral modulation (SSM) for improved reconstruction performance. Existing SSM techniques, however, use either a color filter array or a complex optical arrangement, both of which can only provide limited modulation bandwidth in the spectral dimension. This paper proposes a practical SSM method to help address the "missing cone" problem of chromo-tomography. A high-resolution binary coded aperture is used to modulate the dispersed images, which in the Fourier domain fulfills a 3D convolution of the probed spectrum with the aperture's wide spectrum. This spectrum spreading process facilitates the compressed sensing strategy determined by the Fourier Slice Theorem and we demonstrate the advantages of the proposed approach with numerical experiments. Xuesong Zhang 0001, Jing Jiang 0017, Anlong Ming, Xuejing Kang, Gonzalo R. Arce |
ICASSP | 3 |
| 2019 | Occlusion-Shared and Feature-Separated Network for Occlusion Relationship ReasoningabstractOcclusion relationship reasoning demands closed contour to express the object, and orientation of each contour pixel to describe the order relationship between objects. Current CNN-based methods neglect two critical issues of the task: (1) simultaneous existence of the relevance and distinction for the two elements, i.e, occlusion edge and occlusion orientation; and (2) inadequate exploration to the orientation features. For the reasons above, we propose the Occlusion-shared and Feature-separated Network (OFNet). On one hand, considering the relevance between edge and orientation, two sub-networks are designed to share the occlusion cue. On the other hand, the whole network is split into two paths to learn the high semantic features separately. Moreover, a contextual feature for orientation prediction is extracted, which represents the bilateral cue of the foreground and background areas. The bilateral cue is then fused with the occlusion cue to precisely locate the object regions. Finally, a stripe convolution is designed to further aggregate features from surrounding scenes of the occlusion edge. The proposed OFNet remarkably advances the state-of-the-art approaches on PIOD and BSDS ownership dataset. Feng Xue 0001, Menghan Zhou, Anlong Ming, Yu Zhou 0016 |
ICCV | 4 |
| 2019 | ACPNP: an Efficient Solution for Absolute Camera Pose Estimation from Two Affine CorrespondencesabstractIn this paper, a novel algorithm to estimate the absolute camera pose is proposed using two affine correspondences (ACs). Exploring the relationship between the affine transformation and the projection equation, six linear constraints are derived, and only two ACs are sufficient to recover the pose. Even though perspective cameras are assumed, the constraints can straightforwardly be generalized to other camera models since they describe the relationship between local affinities and projection. Benefiting from the requirement of less correspondences, the proposed algorithm needs less sampling times when robust estimators like RANSAC are applied, and still performs stably with rather limited number of correspondences. For the improvement of robustness, the affine transformation is further optimized via photometric and epipolar constraints. The proposed method was validated on both synthetic and real-world datasets, which demonstrates that the proposed method yields results superior to the state-of-the-art in terms of accuracy. Hengsong Li, Anlong Ming |
ICIP | 4 |
| 2019 | A Unified Unsupervised Learning Framework for Stereo Matching and Ego-Motion EstimationabstractLearning to estimate depth and ego-motion from video sequences via deep convolutional networks is attracting significant attention for potentially wide computer vision applications. Most prior work in unsupervised depth learning use monocular video sequences as the input of their networks. However, their results need a scale factor that is computed frame-to-frame to maintain a stable relative scale. In this paper, we propose an unsupervised learning framework for the task of joint depth and ego-motion estimation from stereo sequences. The usage of stereo sequences can provide both spatial (left to right) and temporal (forward to back-ward) photometric warping constrains for supervised learning and allow for an absolute scale factor for the scene depth and camera pose, which is of great significance for vision guidance. Experiments on the KITTI driving dataset reveal that our framework outperforms state-of-the-art results employing unsupervised neural networks. Hengsong Li, Yuanqi Wang, Anlong Ming |
ICIP | 4 |
| 2019 | Real-Time Light Field Depth Estimation via GPU-Accelerated Muti-View Semi-Global MatchingabstractThe structured and redundant imagery of light field cameras can provide more robust depth estimation results while on the other hand demands a huge computation power, which limits its real-time applications, such as online industrial monitoring, 3D endoscopic surgery etc. This paper extends the classical SGM(Semi-global matching) [1] algorithm to light filed multi-view stereo framework, which can acquire sub-pixel level disparity estimation to cope with the micro-baseline of light field cameras. The whole algorithm is tailored for parallelization on GPU exploiting multi-stream asynchronization, multi-thread allocation, and multi-type memory management. The results show that our method’s execution time is less than 50ms on NVIDIA 1080Ti, which to our knowledge is the fastest among reported geometry based methods while keeping a comparative accuracy performance. Yuanqi Wang, Hengsong Li, Anlong Ming |
ICIP | 4 |
| 2019 | Adaptive Occlusion Boundary Extraction for Depth InferenceabstractIn this paper, we propose an adaptive occlusion boundary extraction method for depth inference based on an adaptive segmentation and classification. First, an Adaptive DRW is proposed to generate more precise seeds and adaptive segmentation results, which can improve the feature quality and lower the boundary imbalance degree. Then, to deal with the imbalanced classification, we design a cost-sensitive boosting method-Adaptive AdaCost to better classify the imbalanced boundary, which can further improve overall performance and lower the cumulative misclassification cost and cost upper bound. Benefited from our Adaptive DRW and AdaCost, we extract more reliable and precise occlusion boundaries and use them for depth inference. The experiment results demonstrate that the combination of our Adaptive DRW and Adaptive AdaCost can produce more precise occlusion boundaries, and the depth inference result with our occlusion boundaries can be greatly improved. Lizhu Ye, Lei Zhu 0012, Xuejing Kang, Anlong Ming |
ICIP | 4 |
| 2019 | Context-Constrained Accurate Contour Extraction for Occlusion Edge DetectionabstractOcclusion edge detection requires both accurate locations and context constraints of the contour. Existing CNN-based pipeline does not utilize adaptive methods to filter the noise introduced by low-level features. To address this dilemma, we propose a novel Context-constrained accurate Contour Extraction Network (CCENet). Spatial details are retained and contour-sensitive context is augmented through two extraction blocks, respectively. Then, an elaborately designed fusion module is available to integrate features, which plays a complementary role to restore details and remove clutter. Weight response of attention mechanism is eventually utilized to enhance occluded contours and suppress noise. The proposed CCENet significantly surpasses state-of-the-art methods on PIOD and BSDS ownership dataset of object edge detection and occlusion orientation detection. Menghan Zhou, Anlong Ming, Yu Zhou 0016 |
ICME | 3 |
| 2019 | A Novel Multi-layer Framework for Tiny Obstacle DiscoveryabstractFor tiny obstacle discovery in a monocular image, edge is a fundamental visual element. Nevertheless, because of various reasons, e.g., noise and similar color distribution with background, it is still difficult to detect the edges of tiny (b) obstacles at long distance. In this paper, we propose an obstacle-aware discovery method to recover the missing contours of these obstacles, which helps to obtain obstacle proposals as much as possible. First, by using visual cues in monocular images, several multi-layer regions are elaborately inferred to reveal the distances from the camera. Second, several novel obstacle-aware occlusion edge maps are constructed to well capture the contours of tiny obstacles, which combines cues from each layer. Third, to ensure the existence of the tiny obstacle proposals, the maps from all layers are used for proposals extraction. Finally, based on these proposals containing tiny obstacles, a novel obstacle-aware regressor is proposed to generate an obstacle occupied probability map with high confidence. The convincing experimental results with comparisons on the Lost and Found dataset demonstrate the effectiveness of our approach, achieving around 9.5% improvement on the accuracy than FPHT and PHT, it even gets comparable performance to MergeNet. Moreover, our method outperforms the state-of-the-art algorithms and significantly improves the discovery ability for tiny obstacles at long distance. Feng Xue 0001, Anlong Ming, Menghan Zhou, Yu Zhou 0016 |
ICRA | 2 |
| 2019 | Reality-Preserving Multiple Parameter Discrete Fractional Angular Transform and Its Application to Color Image EncryptionabstractIn this paper, we first define a reality-preserving multiple parameter fractional angular transform (RPMPDFrAT), which is a useful tool for image encryption. Then, we propose a new color image encryption algorithm based on the defined RPMPDFrAT. The encryption process consists of two phases: encryption in the spatial domain and RPMPDFrAT domain. In the spatial domain, three color components of the plain image are mapped by dual cylindrical transform, which can nonlinearly hide the original color information. Then, the intermediate output is scrambled by a coupled logistic map to reduce the correlation of adjacent pixels and uniformly distribute the image energy of different color components. Thereafter, the scrambled image is transformed by the proposed RPMPDFrAT, which can ensure that we obtain the real-value output. Finally, a process similar to the spatial domain is performed in the RPMPDFrAT domain to further improve the security of the cryptosystem. Numerical simulations are performed and demonstrate that the proposed image encryption algorithm is effective and sensitive to keys. Moreover, some potential attacks are tested to verify the robustness of the proposed method, and the performance of our method outperforms previously published ones. Xuejing Kang, Anlong Ming, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Dynamic Random Walk for Superpixel Segmentation
Lei Zhu 0012, Xuejing Kang, Anlong Ming, Xuesong Zhang 0001 |
ACCV (6) | 3 |
| 2018 | Objectness-Aware Tracking via Double-Layer ModelabstractThe prediction drifts to the non-object backgrounds is a critical issue in conversional correlation filter (CF) based trackers. The key insight of this paper is to propose a double-layer model to address this problem. Specifically, the first layer is a CF tracker, which is employed to predict a rough position of the target, and the objectness layer, which is regarded as the second layer, is utilized to reveal the object characteristics of the predicted target. The novel objectness layer firstly constructs a set of target-related object proposals, which satisfy both the spatial and temporal constraints. And then an objectness classifier is learned upon the proposal set to best separate the target from the noise background proposals. The convincing experimental results on the challenging OTB100 and TC128 dataset demonstrate the effectiveness of the presented approach. Menghan Zhou, Jianxiang Ma, Anlong Ming, Yu Zhou 0016 |
ICIP | 3 |
| 2018 | A New Single Image Super-resolution Method Using SIMK-based Classification and ISRM TechniqueabstractSingle image super-resolution (SR) technique is widely used to estimate high-resolution (HR) images from low-resolution (LR) ones. As a research hotspot, many example-based SR methods achieve superior results by learning class-mapping-kernels from classified external LR-HR patch-pair samples. However, in these methods, the classification of samples is generally based on the features of LR patch, and the interference of ill-samples to learn class-mapping-kernels is ignored as well. In this paper, we propose a new SR method with Sample Individual Mapping-Kernel (SIMK) based classification and Ill-Sample Removal Mechanism (ISRM) in learning LR-HR mapping. In the proposed sample classification, we use the SIMK feature which is the LR-to-HR mapping kernel of each sample, to classify samples and obtain more reasonable sample sets for mapping-learning. To prevent overfitting and reduce the complexity of SIMK-based-classification, samples are pre-categorized by relative pixel values of LR patch. In the mapping-learning process, the ill-samples which are far away from the classification center are removed to improve the validity of class-mapping-kernels. In addition, for each testing LR patch, the optimal class is assigned reasonably based on a probabilistic decision model learned from Naive Bayes Classifier. Comparing with state-of-the-art methods, our SR method achieves both visual and performance improvement. Anlong Ming, Xuejing Kang |
ICPR | 2 |
| 2018 | Learning Training Samples for Occlusion Edge Detection and Its Application in Depth Ordering InferenceabstractThis paper studies the problem of occlusion edge detection, which is applied to infer the depth order of objects in a monocular image. The key observation is that, given the fixed regression objective, the accuracy of occlusion edge detection is effectively boosted by selecting appropriate training samples in a discriminative feature subspace. Specifically, the ℓ1-regularized logistic regression is employed to learn a more sparse yet discriminative feature subspace, while the training sample selection is formulated as a quadratic optimization with the robust Huber loss. The presented formulation avoids the noises efficiently, and hence the desirable occlusion edges can be detected. We validate the effectiveness of our approach on depth order inference problem. Experiments are conducted on two famous datasets, i.e., the Cornell depth-order dataset and the NYU2 dataset. Promising results demonstrate the superiority of our approach over the state-of-the-art approaches. Yu Zhou 0016, Jianxiang Ma, Anlong Ming, Xiang Bai |
ICPR | 3 |
| 2017 | Object-Level ProposalsabstractEdge and surface are two fundamental visual elements of an object. The majority of existing object proposal approaches utilize edge or edge-like cues to rank candidates, while we consider that the surface cue containing the 3D characteristic of objects should be captured effectively for proposals, which has been rarely discussed before. In this paper, an object-level proposal model is presented, which constructs an occlusion-based objectness taking the surface cue into account. Specifically, the better detection of occlusion edges is focused on to enrich the surface cue into proposals, namely, the occlusion-dominated fusion and normalization criterion are designed to obtain the approximately overall contour information, to enhance the occlusion edge map at utmost and thus boost proposals. Experimental results on the PASCAL VOC 2007 and MS COCO 2014 dataset demonstrate the effectiveness of our approach, which achieves around 6% improvement on the average recall than Edge Boxes at 1000 proposals and also leads to a modest gain on the performance of object detection. Jianxiang Ma, Anlong Ming, Xinggang Wang, Yu Zhou 0016 |
ICCV | 2 |
| 2017 | A non-rigid 3D model retrieval method based on scale-invariant heat kernel signature features
Pengjie Li, Huadong Ma, Anlong Ming |
Multim. Tools Appl. | 3 |
| 2016 | A energy efficient multi-dimension model for system control in smart environment systemsabstractA smart environment system should automatically control the devices according to the sensing information and users' requirements so as to keep the environmental elements (e.g., temperature, light) within the desired range. System control with minimum power is one key issue in such a system. In this paper, we propose a multi-dimension model for system control. In this model, each environmental element is abstracted into a dimension, such that a service with conditions and targets can be formulated as a multi-dimensional service space, and a smart environment with many services may map to a comprehensive multi-dimensional service space through space computation. Based on this model, we propose a minimum power adjustment algorithm for energy-efficient scheduling in smart environment, which transforms the optimal control problem into the problem of the shortest weighted distance of point-to-polygonal in multi-dimensional space. Theoretical analysis and experimental results show that the proposed model is effective and efficient in energy-efficient system control. It is important to point out that the proposed algorithms are scalable when the number of dimensions or services increases. Anlong Ming, Yanchen Ren, Zhibo Pang, Kim Fung Tsang |
INDIN | 1 |
| 2016 | Human action recognition with skeleton induced discriminative approximate rigid part model
Yu Zhou 0016, Anlong Ming |
Pattern Recognit. Lett. | 2 |
| 2015 | Learning discriminative occlusion feature for depth ordering inference on monocular imageabstractIn this paper, a novel depth ordering inference approach is presented. Our main insight is to integrate the discriminative feature selection, occlusion feature learning and same-layer (S-L) relationship judgement into a uniform sparsity based classification objective, which cannot only supply the precise segmentation for the occlusion edge, but also reduce the solution space for the depth ordering inference efficiently. In addition, a novel triple descriptor is adopted to judge the foreground relationship, which is more discriminative than conversional local cues and can further reduce the solution space. The inference is executed by finding a valid path on a directed graph model. We validate our approach on the Cornell depth-order dataset and the NYU 2 dataset, and the convincing experimental results demonstrate the effectiveness of our approach. Anlong Ming, Baofeng Xun, Jia Ni, Mingfei Gao, Yu Zhou 0016 |
ICIP | 1 |
| 2013 | Combining topological and view-based features for 3D model retrieval
Pengjie Li, Huadong Ma, Anlong Ming |
Multim. Tools Appl. | 3 |
| 2012 | Robust tracking by accounting for hard negatives explicitly
Tianfu Wu 0004, Mingtao Pei, Anlong Ming, Zhenyu Yao |
ICPR | 4 |
| 2012 | A binary-classification-tree based framework for distributed target classification in multimedia sensor networksabstractWith rapid improvements and miniaturization in hardware, sensor nodes equipped with acoustic and visual information collection modules promise an unprecedented opportunity for target surveillance applications. This paper investigates a critical task of target surveillance, multi-class classification, in distributed multimedia sensor networks. We first analyze the procedure of target classification utilizing the acoustic and visual information. Then, we propose a binary classification tree based framework for distributed target classification in multimedia sensor networks. The proposed framework includes three main components: Generation of binary classification tree, Division of binary classification tree, and Selection of multimedia sensor nodes. Finally, we conduct an experimental application of target classification and extensive simulations to validate and evaluate our proposed framework and related schemes. Liang Liu 0001, Anlong Ming, Huadong Ma, Xi Zhang 0005 |
INFOCOM | 2 |
| 2011 | View-based 3D model retrieval using two-level spatial structureabstractRecently, the view-based 3D model retrieval methods have received great research attentions. However, these methods are difficult to preserve the spatial structure of 3D models. In this paper, we propose a novel view-based 3D model retrieval method to solve this problem. Our method is based on the two-level (3D model-level and 2D image-level) spatial structure. Firstly, we extract the spatial structure circular descriptor (SSCD) images from 3D models. The SSCD images can preserve the spatial structure on the 3D model-level. Then, we modify the bag-of-features (BOF) method to extract view-based features from these SSCD images. The modified BOF method can preserve the spatial structure on the 2D image-level. Finally, we calculate the similarity between the query model and the models in the databases by adapting the earth mover distance method. Experimental results show that our method can achieve satisfactory retrieval performance for both the articulated models and the rigid models. Pengjie Li, Huadong Ma, Anlong Ming |
ICIP | 3 |
| 2011 | EGMM: An enhanced Gaussian mixture model for detecting moving objects with intermittent stopsabstractMoving object detection is one of the most important tasks in intelligent visual surveillance systems. Gaussian Mixture Model (GMM) has been most widely used for moving object detection, because of its robustness to variable scenes. However, to the best of our knowledge, existing GMM based methods can not detect moving objects which gradually stop and keep still state for a while. In this paper, we present an Enhanced Gaussian MixtureModel, called EGMM, to handle this problem. We integrate an Initial Gaussian Background Model (IGBM) and an extended Kalman filter based tracker with GMM, to enhance its performance. Experimental results show that our EGMM based method has a lower miss rate at the same false positives per image comparing to GMM based method for moving pedestrian detection, and it also has a higher detection rate for abandoned object detection comparing to GMM based method. Huiyuan Fu, Huadong Ma, Anlong Ming |
ICME | 3 |
| 2011 | View-based 3D model retrieval with topological structureabstractWith the rapidly increasing of 3D models, the view-based 3D model retrieval methods have received significant research attention. The previous view-based methods can achieve ideal retrieval result for the rigid models, but they only obtain poor retrieval result for the deformable models because they can not preserve the topological structure well. In this paper, we propose a view-based 3D model retrieval algorithm using topological structure. We extract the view-based features from the images rendered at the salient topological points. To preserve the topological structure of the 3D model, a multiresolutional reeb graph (MRG) is constructed according to the salient topological points. We take the view-based features as the attribute information of the corresponding MRG nodes. The comparison between two 3D models is transformed to compute the similarity of the corresponding MRGs. Experimental results on two standard benchmarks show that our algorithm can achieve satisfactory retrieval performance for both the deformable models and the rigid models. Pengjie Li, Huadong Ma, Anlong Ming |
ICME | 3 |
| 2011 | Fast accurate pedestrian detection using a MPL-Boosted cascade of weak FIK-SVM classifiersabstractWe address the problem of pedestrian detection in still images. Current pedestrian detection systems are hard to improve both speed and accuracy simultaneously. In order to achieve a balance between speed and accuracy, we propose a novel MPL-Boosted cascade of weak FIK-SVM classifiers. Our method achieves high recall while taking the speed-advantage of cascade-of-rejectors approach. Each feature in our algorithm corresponds to a 66-D HOG-LBP feature vector that describe a block. The weak classifiers we use are the separating hyper-plane computed by using a FIK-SVM. We use MPL-Boost to select features from a large set of possible blocks. The integral image and convoluted trilinear interpolation are used for rapid calculation of block feature. For a 320×240 image, the system can process 16 frames per second with sparse scan, while defeat the accuracy level of existing methods. Junqiang Wang, Huadong Ma, Anlong Ming |
ICME | 3 |
| 2011 | Non-rigid 3D model retrieval using multi-scale local featuresabstractThe number of available non-rigid 3D models in various areas increases steadily. The local features are more effective than global features for the search of these non-rigid 3D models. Global descriptors fail to consistently compensate for the intra-class variability of non-rigid 3D models. To solve this problem, we propose a non-rigid 3D model retrieval method based on multi-scale local features. Firstly, we extract keypoints at multiple scales automaticlly. Then, the Heat Kernel Signature (HKS) local descriptors are computed for each keypoint. However, the HKS descriptors are sensitive to scale. In order to solve this problem, the HKS descriptors are put into the Bag-of-Features (BOF) framework. In the BOF framework, we use a kind of histogram equalization technique to make our feature descriptor robust to model scaling. Experimental results on two public benchmarks show that our algorithm can achieve satisfactory retrieval performance for the non-rigid 3D models. Pengjie Li, Huadong Ma, Anlong Ming |
ACM Multimedia | 3 |
| 2010 | Fast human detection using mi-sVM and a cascade of HOG-LBP featuresabstractThis paper presents a human detection approach which can process images rapidly and detect the objects accurately. The features used in our system are the cascade of the HOG (Histograms of Oriented Gradients) and LBP (Local Binary Pattern). In order to achieve high recall at each stage of the cascade, we modify the mi-SVM (Support Vector Machine for multiple instance learning) to train the HOG and LBP features respectively. In this way, we implement a novel cascade-ofrejectors method to detect the human fast, while maintaining a similar accuracy reported in previous methods. Experimental results show our method can process frames at 5 to 10 frames per second, depending on the scanning density in the image. Chengbin Zeng, Huadong Ma, Anlong Ming |
ICIP | 3 |
| 2009 | A collaborative surveillance system for role identificationabstractMany works of conventional surveillance have focused on people tracking, behavior or event detection, gait or face based recognition, etc. However, role identification is also very important in video surveillance but usually paid less attention. In this paper, we propose a collaborative multi-camera system to identify people with specific roles using a causal network to form a best identification result from the evidences known so far. Not only visual features but also spatio-temporal features are used in our method, as well as some object specific features. Collaborative multiple cameras benefit locating the position of moving objects and overcoming occlusions. Experimental results demonstrates the effectiveness of our method. Anlong Ming, Huadong Ma |
CSCWD | 1 |
| 2009 | A Coverage-Enhancing Method for 3D Directional Sensor NetworksabstractIn conventional directional sensor networks, coverage control for each sensor is based on a 2D directional sensing model. However, 2D directional sensing model cannot accurately characterize the actual application scene of image/video sensor networks. To remedy this deficiency, we propose a 3D directional sensor coverage-control model with tunable orientations. In order to improve the efficiency of target-detecting, we develop a virtual potential-field based coverage-enhancing scheme to improve the coverage performance. Furthermore, we apply the simulated annealing algorithm for objective optimization. The extensive simulations show the effectiveness of our proposed 3D sensing model and coverage enhancing method. Huadong Ma, Xi Zhang 0005, Anlong Ming |
INFOCOM | 3 |
| 2008 | Frame-skipping tracking for single object with global motion detectionabstractFrame-skipping videos usually appear in wireless video sensor networks which have wirelessly interconnected devices that are able to ubiquitously retrieve video content from the environment. Frame-skipping videos bring to difficulties in getting the transition model (how objects move between frames). We propose a particle filter with global motion detection requiring no offline or online learning. Experimental results show the proposed approach improves the tracking accuracy in comparison with the existing conventional methods, under the condition of frame skipping data and motion of both targets and video sensors. Anlong Ming, Huadong Ma |
ICPR | 1 |
| 2007 | An Algorithm Testbed for the Biometrics Grid
Anlong Ming, Huadong Ma |
GPC | 1 |
| 2006 | An Improved Approach to the Line-Based Face RecognitionabstractThe line-based face recognition method is distinguished by its features, but its development and application is limited to some inherent drawbacks. This paper propose a method for decreasing the influence under variable illumination intensity by using the line-based singular value (LSV) feature vector instead of image gray-level value to calculate "distance" between two lines. We prove that our method is invariant to the illumination intensity. Finally, we suggest a distributed computing algorithm using grid computing to solve the multi-scale computation. Experimental results show our approach is effective Anlong Ming, Huadong Ma |
ICME | 1 |