EDBT 2026 Demo / reviewers in the wild / expert
Chaomin Shen 0001
dblp:32/3402-1
· DBLP profile ↗
41ranked-venue papers
6as first author
19since 2021 · last 2025
0000-0001-9389-6472ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorSystems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Comprehensive Overhaul of Multimodal Assistant with Small Language ModelsabstractMultimodal Large Language Models (MLLMs) have showcased impressive skills in tasks related to visual understanding and reasoning. Yet, their widespread application faces obstacles due to the high computational demands during both the training and inference phases, restricting their use to a limited audience within the research and user communities. In this paper, we investigate the design aspects of Multimodal Small Language Models (MSLMs) and propose an efficient multimodal assistant named Mipha, which is designed to create synergy among various aspects: visual representation, language models, and optimization strategies. We show that without increasing the volume of training data, our Mipha-3B outperforms the state-of-the-art large MLLMs, especially LLaVA-1.5-13B, on multiple benchmarks. Through detailed discussion, we provide insights and guidelines for developing strong MSLMs that rival the capabilities of MLLMs. Minjie Zhu, Yichen Zhu 0001, Ning Liu 0007, Xin Liu 0086, Chaomin Shen 0001, Yaxin Peng |
AAAI | 6 |
| 2025 | ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action ModelabstractZhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhongyi Zhou, Yichen Zhu 0001, Minjie Zhu, Ning Liu 0007, Weibin Meng, Yaxin Peng, Chaomin Shen 0001, Feifei Feng |
EMNLP | 9 |
| 2025 | Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual LearningabstractContinual Learning enables models to learn and adapt to new tasks while retaining prior knowledge. Introducing new tasks, however, can naturally lead to feature entanglement across tasks, limiting the model’s capability to distinguish between new domain data. In this work, we propose a method called Feature Realignment through Experts on hyperSpHere in Continual Learning (Fresh-CL). By leveraging predefined and fixed simplex equiangular tight frame (ETF) classifiers on a hypersphere, our model improves feature separation both intra and inter tasks. However, the projection to a simplex ETF shifts with new tasks, disrupting structured feature representation of previous tasks and degrading performance. Therefore, we propose a dynamic extension of ETF through mixture of experts, enabling adaptive projections onto diverse subspaces to enhance feature representation. Experiments on 11 datasets demonstrate a 2% improvement in accuracy compared to the strongest baseline, particularly in fine-grained datasets, confirming the efficacy of combining ETF and MoE to improve feature distinction in continual learning scenarios. Zhongyi Zhou, Yaxin Peng, Pin Yi, Minjie Zhu, Chaomin Shen 0001 |
ICASSP | 5 |
| 2025 | DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and AutoregressionabstractIn this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inherently lack reasoning capabilities. Central to our approach is autoregressive reasoning — a task decomposition and explanation process enabled by a pre-trained VLM — to guide diffusion-based action policies. To tightly couple reasoning with action generation, we introduce a reasoning injection module that directly embeds self-generated reasoning phrases into the policy learning process. The framework is simple, flexible, and efficient, enabling seamless deployment across diverse robotic platforms.
We conduct extensive experiments using multiple real robots to validate the effectiveness of DiVLA. Our tests include a challenging factory sorting task, where DiVLA successfully categorizes objects, including those not seen during training. The reasoning injection module enhances interpretability, enabling explicit failure diagnosis by visualizing the model’s decision process. Additionally, we test DiVLA on a zero-shot bin-picking task, achieving \textbf{63.7\% accuracy on 102 previously unseen objects}. Our method demonstrates robustness to visual changes, such as distractors and new backgrounds, and easily adapts to new embodiments. Furthermore, DiVLA can follow novel instructions and retain conversational ability. Notably, DiVLA is data-efficient and fast at inference; our smallest DiVLA-2B runs 82Hz on a single A6000 GPU. Finally, we scale the model from 2B to 72B parameters, showcasing improved generalization capabilities with increased model size. Yichen Zhu 0001, Minjie Zhu, Zhibin Tang, Zhongyi Zhou, Chaomin Shen 0001, Yaxin Peng, Feifei Feng |
ICML | 8 |
| 2025 | Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic ManipulationabstractDiffusion Policy is a powerful technique tool for learning end-to-end visuomotor robot control. It is expected that Diffusion Policy possesses scalability, a key attribute for deep neural networks, typically suggesting that increasing model size would lead to enhanced performance. However, our observations indicate that Diffusion Policy in transformer architecture (DP-T) struggles to scale effectively; even minor additions of layers can deteriorate training outcomes. To address this issue, we introduce Scalable Diffusion Transformer Policy for visuomotor learning. Our proposed method, namely ScaleDP, introduces two modules that improve the training dynamic of Diffusion Policy and allow the network to better handle multimodal action distribution. First, we identify that DPT suffers from large gradient issues, making the optimization of Diffusion Policy unstable. To resolve this issue, we factorize the feature embedding of observation into multiple affine layers, and integrate it into the transformer blocks. Additionally, our utilize non-causal attention which allows the policy network to “see” future actions during prediction, helping to reduce compounding errors. We demonstrate that our proposed method successfully scales the Diffusion Policy from 10 million to 1 billion parameters. This new model, named ScaleDP, can effectively scale up the model size with improved performance and generalization. We benchmark ScaleDP across 50 different tasks from MetaWorld and find that our largest ScaleDP outperforms DP-T with an average improvement of 21.6%. Across 7 real-world robot tasks, our ScaleDP demonstrates an average improvement of 36. 25% over DP-T on four single-arm tasks and 75% on three bimanual tasks. We believe our work paves the way for scaling up models for visuomotor learning. The project page is available at https://scaling-diffusion-policy.github.io/. Minjie Zhu, Yichen Zhu 0001, Ning Liu 0007, Chaomin Shen 0001, Yaxin Peng, Feifei Feng, Jian Tang 0008 |
ICRA | 8 |
| 2025 | Efficient Feature Fusion for UAV Object DetectionabstractObject detection in unmanned aerial vehicle (UAV) remote sensing images poses significant challenges due to unstable image quality, small object sizes, complex backgrounds, and environmental occlusions. Small objects, in particular, occupy small portions of images, making their accurate detection highly difficult. Existing multi-scale feature fusion methods address these challenges to some extent by aggregating features across different resolutions. However, they often fail to effectively balance the classification and localization performance for small objects, primarily due to insufficient feature representation and imbalanced network information flow. In this paper, we propose a novel feature fusion framework specifically designed for UAV object detection tasks to enhance both localization accuracy and classification performance. The proposed framework integrates hybrid upsampling and downsampling modules, enabling feature maps from different network depths to be flexibly adjusted to arbitrary resolutions. This design facilitates cross-layer connections and multi-scale feature fusion, ensuring improved representation of small objects. Our approach leverages hybrid downsampling to enhance fine-grained feature representation, improving spatial localization of small targets, even under complex conditions. Simultaneously, the upsampling module aggregates global contextual information, optimizing feature consistency across scales and enhancing classification robustness in cluttered scenes. Experimental results on two public UAV datasets demonstrate the effectiveness of the proposed framework. Integrated into the YOLO-v10 model, our method achieves a 2 percentage points improvement in average precision (AP) compared to the baseline YOLO-v10 model, while maintaining the same number of parameters. These results highlight the potential of our framework for accurate and efficient UAV object detection. Yaxin Peng, Chaomin Shen 0001 |
IJCNN | 3 |
| 2025 | ChatVLA-2: Vision-Language-Action Model with Open-World ReasoningabstractVision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks. We argue that a generalizable VLA model should retain and expand upon the VLM's core competencies: 1) **Open-world reasoning** - the VLA should inherit the knowledge from VLM, i.e., recognize anything that the VLM can recognize, capable of solving math problems, possessing visual-spatial intelligence, 2) **Reasoning following** – effectively translating the open-world reasoning into actionable steps for the robot. In this work, we introduce **ChatVLA-2**, a novel mixture-of-expert VLA model coupled with a specialized three-stage training pipeline designed to preserve the VLM’s original strengths while enabling actionable reasoning. To validate our approach, we design a math-matching task wherein a robot interprets math problems written on a whiteboard and picks corresponding number cards from a table to solve equations. Remarkably, our method exhibits exceptional mathematical reasoning and OCR capabilities, despite these abilities not being explicitly trained within the VLA. Furthermore, we demonstrate that the VLA possesses strong spatial reasoning skills, enabling it to interpret novel directional instructions involving previously unseen objects. Overall, our method showcases reasoning and comprehension abilities that significantly surpass state-of-the-art imitation learning methods such as OpenVLA, DexVLA, and $\pi_0$. This work represents a substantial advancement toward developing truly generalizable robotic foundation models endowed with robust reasoning capacities. Zhongyi Zhou, Yichen Zhu 0001, Zhibin Tang, Yaxin Peng, Chaomin Shen 0001 |
NeurIPS | 7 |
| 2025 | SAB Net: A Semantic Attention Boosting Framework for Semantic SegmentationabstractSemantic segmentation has achieved great progress by effectively fusing features of contextual information. In this article, we propose an end-to-end semantic attention boosting (SAB) framework to adaptively fuse the contextual information iteratively across layers with semantic regularization. Specifically, we first propose a pixelwise semantic attention (SAP) block, with a semantic metric representing the pixelwise category relationship, to aggregate the nonlocal contextual information. In addition, we improve the computation complexity of SAP block from to for images with size . Second, we present a categorywise semantic attention (SAC) block to adaptively balance the nonlocal contextual dependencies and the local consistency with a categorywise weight, overcoming the contextual information confusion caused by the feature imbalance within intra-category. Furthermore, we propose the SAB module to refine the segmentation with SAC and SAP blocks. By applying the SAB module iteratively across layers, our model shrinks the semantic gap and enhances the structure reasoning by fully utilizing the coarse segmentation information. Extensive quantitative evaluations demonstrate that our method significantly improves the segmentation results and achieves superior performance on the PASCAL VOC 2012, Cityscapes, PASCAL Context, and ADE20K datasets. Xiaofeng Ding 0003, Chaomin Shen 0001, Tieyong Zeng, Yaxin Peng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Exploring Gradient Explosion in Generative Adversarial Imitation Learning: A Probabilistic PerspectiveabstractGenerative Adversarial Imitation Learning (GAIL) stands as a cornerstone approach in imitation learning. This paper investigates the gradient explosion in two types of GAIL: GAIL with deterministic policy (DE-GAIL) and GAIL with stochastic policy (ST-GAIL). We begin with the observation that the training can be highly unstable for DE-GAIL at the beginning of the training phase and end up divergence. Conversely, the ST-GAIL training trajectory remains consistent, reliably converging. To shed light on these disparities, we provide an explanation from a theoretical perspective. By establishing a probabilistic lower bound for GAIL, we demonstrate that gradient explosion is an inevitable outcome for DE-GAIL due to occasionally large expert-imitator policy disparity, whereas ST-GAIL does not have the issue with it. To substantiate our assertion, we illustrate how modifications in the reward function can mitigate the gradient explosion challenge. Finally, we propose CREDO, a simple yet effective strategy that clips the reward function during the training phase, allowing the GAIL to enjoy high data efficiency and stable trainability. Wanying Wang, Yichen Zhu 0001, Yirui Zhou, Chaomin Shen 0001, Jian Tang 0008, Yaxin Peng, Yangchun Zhang |
AAAI | 4 |
| 2024 | Harmonizing Knowledge Transfer in Neural Network with Unified Distillation
Yaomin Huang, Zaomin Yan, Chaomin Shen 0001, Faming Fang, Guixu Zhang |
ECCV (33) | 3 |
| 2024 | Object-Centric Instruction Augmentation for Robotic ManipulationabstractHumans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as "pick and place", understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the literature that uses the large language model to enrich the text descriptions, the latter remains underexplored. In this work, we introduce the Object-Centric Instruction Augmentation (OCI) framework to augment highly semantic and information-dense language instruction with position cues. We utilize a Multi-modal Large Language Model (MLLM) to weave knowledge of object locations into natural language instruction, thus aiding the policy network in mastering actions for versatile manipulation. Additionally, we present a feature reuse mechanism to integrate the vision-language features from off-the-shelf pre-trained MLLM into policy networks. Through a series of simulated and real-world robotic tasks, we demonstrate that robotic manipulator imitation policies trained with our enhanced instructions outperform those relying solely on traditional language instructions. Yichen Zhu 0001, Minjie Zhu, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008 |
ICRA | 7 |
| 2024 | Language-Conditioned Robotic Manipulation with Fast and Slow ThinkingabstractThe language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple "pick-and-place" to tasks requiring intent recognition and visual reasoning. Inspired by the dual-process theory in cognitive science—which suggests two parallel systems of fast and slow thinking in human decision-making—we introduce Robotics with Fast and Slow Thinking (RFST), a framework that mimics human cognitive architecture to classify tasks and makes decisions on two systems based on instruction types. Our RFST consists of two key components: 1) an instruction discriminator to determine which system should be activated based on the current user’s instruction, and 2) a slow-thinking system that is comprised of a fine-tuned vision-language model aligned with the policy networks, which allow the robot to recognize user’s intention or perform reasoning tasks. To assess our methodology, we built a dataset featuring real-world trajectories, capturing actions ranging from spontaneous impulses to tasks requiring deliberate contemplation. Our results, both in simulation and real-world scenarios, confirm that our approach adeptly manages intricate tasks that demand intent recognition and reasoning. Minjie Zhu, Yichen Zhu 0001, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008 |
ICRA | 7 |
| 2024 | Student-Oriented Teacher Knowledge Refinement for Knowledge DistillationabstractKnowledge distillation has become widely recognized for its ability to transfer knowledge from a large teacher network to a compact and more streamlined student network. Traditional knowledge distillation methods primarily follow a teacher-oriented paradigm that imposes the task of learning the teacher's complex knowledge onto the student network. However, significant disparities in model capacity and architectural design hinder the student's comprehension of the complex knowledge imparted by the teacher, resulting in sub-optimal performance. This paper introduces a novel perspective emphasizing student-oriented and refining the teacher's knowledge to better align with the student's needs, thereby improving knowledge transfer effectiveness. Specifically, we present the Student-Oriented Knowledge Distillation (SoKD), which incorporates a learnable feature augmentation strategy during training to refine the teacher's knowledge of the student dynamically. Furthermore, we deploy the Distinctive Area Detection Module (DAM) to identify areas of mutual interest between the teacher and student, concentrating knowledge transfer within these critical areas to avoid transferring irrelevant information. This customized module ensures a more focused and effective knowledge distillation process. Our approach, functioning as a plug-in, could be integrated with various knowledge distillation methods. Extensive experimental results demonstrate the efficacy and generalizability of our method. Chaomin Shen 0001, Yaomin Huang, Haokun Zhu, Jinsong Fan, Guixu Zhang |
ACM Multimedia | 1 |
| 2024 | AffViT: Fast Affine Medical Image Registration with Convolutional Vision Transformer
Chaomin Shen 0001, Zhongyi Zhou |
PRICAI (3) | 1 |
| 2023 | CP3: Channel Pruning Plug-in for Point-Based NetworksabstractChannel pruning can effectively reduce both computational cost and memory footprint of the original network while keeping a comparable accuracy performance. Though great success has been achieved in channel pruning for 2D image-based convolutional networks (CNNs), existing works seldom extend the channel pruning methods to 3D point-based neural networks (PNNs). Directly implementing the 2D CNN channel pruning methods to PNNs undermine the performance of PNNs because of the different representations of 2D images and 3D point clouds as well as the network architecture disparity. In this paper, we proposed CP3, which is a Channel Pruning Plugin for Point-based network. CP3is elaborately designed to leverage the characteristics of point clouds and PNNs in order to enable 2D channel pruning methods for PNNs. Specifically, it presents a coordinate-enhanced channel importance metric to reflect the correlation between dimensional information and individual channel features, and it recycles the discarded points in PNN's sampling process and reconsiders their potentially-exclusive information to enhance the robustness of channel pruning. Experiments on various PNN architectures show that CP3constantly improves state-of-the-art 2D CNN pruning approaches on different point cloud tasks. For instance, our compressed PointNeXt-S on ScanObjectNN achieves an accuracy of 88.52% with a pruning rate of 57.8%, outperforming the baseline pruning methods with an accuracy gain of 1.94%. Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Guixu Zhang, Xinmei Liu, Feifei Feng, Jian Tang 0008 |
CVPR | 5 |
| 2023 | CMG-Net: An End-to-End Contact-based Multi-Finger Dexterous Grasping NetworkabstractIn this paper, we propose a novel representation for grasping using contacts between multi-finger robotic hands and objects to be manipulated. This representation significantly reduces the prediction dimensions and accelerates the learning process. We present an effective end-to-end network, CMG-Net, for grasping unknown objects in a cluttered environment by efficiently predicting multi-finger grasp poses and hand configurations from a single-shot point cloud. Moreover, we create a synthetic grasp dataset that consists of five thousand cluttered scenes, 80 object categories, and 20 million annotations. We perform a comprehensive empirical study and demonstrate the effectiveness of our grasping representation and CMG-Net. Our work significantly outperforms the state-of-the-art for three-finger robotic hands. We also demonstrate that the model trained using synthetic data perform very well for real robots. Mingze Wei, Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Feifei Feng, Chun Shan, Jian Tang 0008 |
ICRA | 7 |
| 2022 | Label-Guided Auxiliary Training Improves 3D Object Detector
Yaomin Huang, Xinmei Liu, Yichen Zhu 0001, Chaomin Shen 0001, Zhengping Che, Guixu Zhang, Yaxin Peng, Feifei Feng, Jian Tang 0008 |
ECCV (9) | 5 |
| 2022 | Scale robust point matching-Net: End-to-end scale point matching using Lie groupabstractAbstract Point cloud matching is an important procedure in a variety of computer vision tasks. Traditional point cloud matching methods have made great progress, while neural network‐based approaches are becoming a trend, powered by their strong capabilities of feature extraction. Existing point matching neural networks, however, mainly focus on the rigid transformation. More complex transformations should also be considered in many scenarios. In this regard, the authors extend the rigid registration to non‐rigid cases and propose a network called the Scale Robust Point Matching (SRPM)‐Net for scale point matching. This robust structure‐preserving network is implemented by incorporating Lie group parametrisation. It is conducted by Lie group linearisation representation with the constraints of parameters under the corresponding basis of Lie algebra. SRPM‐Net preserves the structure of the solution and avoids degeneration. The contributions of this paper lie in two aspects: Most importantly, SRPM‐Net provides an extendable framework for handling complicated transformations. Secondly, it introduces a new feature learning module, which better preserves the shape structure by aggregating the high‐dimensional feature and calculating the normal vector of point cloud surface automatically. Experimental results show that SRPM‐Net is more robust and accurate than existing traditional and recent deep learning methods under various situations. Xin Wang 0084, Guangwei Zhao, Yaxin Peng, Chaomin Shen 0001 |
IET Comput. Vis. | 5 |
| 2021 | VAN: Voting and Attention Based Network for Unsupervised Medical Image Registration
Zhiang Zu, Guixu Zhang, Yaxin Peng, Chaomin Shen 0001 |
PRICAI (1) | 5 |
| 2020 | Label Smoothing Technique for Ordinal Classification in Cloud AssessmentabstractSatellite image classification is a challenging task if the input labels are not sufficiently accurate. The automatic cloud cover assessment (ACCA), for example, aims to classify the cloud covers of satellite images as alphabetical categories from A to E showing the escalating levels of clouds; however, those labels for training are often obtained by a subjective qualitative assessment, i.e., they may be not accurate. Therefore, this paper studies how to conduct ACCA under this circumstance. We propose a label smoothing approach and improve the accuracy around 3 percentage points (e.g., from 75.9% to 78.4% for ResNet network) without changing other network structures and parameters. Yuxuan Wei, Qixuan Liu, Guixu Zhang, Yaxin Peng, Chaomin Shen 0001 |
IGARSS | 5 |
| 2020 | Rate-distortion model for grayscale-invariance reversible data hiding
Siyan Zhou, Weiming Zhang 0001, Chaomin Shen 0001 |
Signal Process. | 3 |
| 2019 | The Adversarial Attack and Detection under the Fisher Information MetricabstractMany deep learning models are vulnerable to the adversarial attack, i.e., imperceptible but intentionally-designed perturbations to the input can cause incorrect output of the networks. In this paper, using information geometry, we provide a reasonable explanation for the vulnerability of deep learning models. By considering the data space as a non-linear space with the Fisher information metric induced from a neural network, we first propose an adversarial attack algorithm termed one-step spectral attack (OSSA). The method is described by a constrained quadratic form of the Fisher information matrix, where the optimal adversarial perturbation is given by the first eigenvector, and the vulnerability is reflected by the eigenvalues. The larger an eigenvalue is, the more vulnerable the model is to be attacked by the corresponding eigenvector. Taking advantage of the property, we also propose an adversarial detection method with the eigenvalues serving as characteristics. Both our attack and detection algorithms are numerically optimized to work efficiently on large datasets. Our evaluations show superior performance compared with other methods, implying that the Fisher information is a promising approach to investigate the adversarial attacks and defenses. Chenxiao Zhao, P. Thomas Fletcher, Mixue Yu, Yaxin Peng, Guixu Zhang, Chaomin Shen 0001 |
AAAI | 6 |
| 2019 | Deep Ordinal Classification for Automatic Cloud AssessmentabstractWe develop a deep neural network for ordinal classification and use this technique for the automatic cloud cover assessment (ACCA) of satellite images. We adopt a VGG+ResNet approach with a novel loss function to solve this problem. Results using Quicklook images from Centre for Remote Imaging, Sensing and Processing (CRISP) are promising, as approximately 97.3% of all sub-scenes are correctly labelled, potentially even higher when one ordinal error is tolerated. Qixuan Liu, Jinsong Fan, Chaomin Shen 0001, Yaxin Peng |
IGARSS | 3 |
| 2019 | Strict Subspace and Label-Space Structure for Domain Adaptation
Ying Li 0028, Yaxin Peng, Chaomin Shen 0001 |
KSEM (1) | 5 |
| 2019 | Partial Alignment of Data Sets Based on Fast Intrinsic Feature Match
Yaxin Peng, Naiwu Wen, Xiaohuang Zhu, Chaomin Shen 0001 |
KSEM (1) | 4 |
| 2018 | Parallel Hashing Using Representative Points in HyperoctantsabstractThe goal of hashing is to learn a low-dimensional binary representation of high-dimensional information, leading to a tremendous reduction of computational cost. Previous studies usually achieved this goal by applying projection or quantization methods. However, the projection method fails to capture the intrinsic data structures, and the quantization method cannot make full use of complete information by its strategy of partitioning original space. To combine their advantages and avoid their drawbacks, we propose a novel algorithm, termed as representative points quantization (RPQ), by using the representative points defined as the barycenters of points in the hyperoctants. To settle the problem of exponential time complexity with the growth of the coding length, for long hashing codes, we further propose a parallel RPQ (PRPQ) algorithm, by separating a long code into several short codes, re-coding the short codes in different low dimensional subspaces, and then concatenating them to a long code. Experiments on image retrieval tasks demonstrate that RPQ and PRPQ can well capture the main topology structure of data, showing that our algorithm achieves better performance than state-of-the-art methods. Chaomin Shen 0001, Mixue Yu, Chenxiao Zhao, Yaxin Peng, Guixu Zhang |
CIKM | 1 |
| 2018 | Cloud Cover Assessment in Satellite Images Via Deep Ordinal ClassificationabstractThe percentage of cloud cover is one of the key indices for satellite data products. To date, cloud cover assessment is performed manually in most groundstations. To facilitate the process, this paper proposes a deep learning approach for cloud cover assessment in quicklook satellite images. The quicklook images from Centre for Remote Imaging, Sensing and Processing (CRISP) are used for demonstration. Same as the manual operation, given a quicklook image, the algorithm returns 8 labels ranging from A to E and *, indicating the cloud percentages in different areas of the image. This is achieved by constructing 8 improved VGG-16 models, where parameters such as the loss function, learning rate and dropout are tailored for better performance. Results indicate that approach is promising, as around 85% of sub-scenes are correctly labelled, and the accuracy is even higher if one ordinal error is accepted. This paper demonstrates a new application in remote sensing using state-of-the-art deep learning techniques. Chaomin Shen 0001, Chenxiao Zhao, Mixue Yu, Yaxin Peng |
IGARSS | 1 |
| 2018 | Global Nonlinear Metric Learning by Gluing Local Linear MetricsabstractWe address the nonlinear metric learning by constructing a smooth nonlinear metric from the data. First, we locally define an initial linear metric on each cluster by principal component analysis. Second, we glue such local linear metrics to form a smooth nonlinear metric by a partition of unity on the sample space, and further learn the global nonlinear metric. Third, we conduct the intrinsic steepest descent algorithm on matrix manifolds for implementation. Finally, we compare our approach with several state-of-the-art methods on a variety of datasets. The results validate that the robustness and accuracy of classification are both improved under our nonlinear metric. The novelty of our global smooth nonlinear metric learning model lies in that it has completely overcome drawbacks of local metric learning methods: the partition coefficients obtained by the partition of unity is smooth, while the metric at any point on the manifold can be directly defined. Yaxin Peng, Lingfang Hu, Shihui Ying, Chaomin Shen 0001 |
SDM | 4 |
| 2018 | Performance Analysis for SVM Combining with Metric Learning
Lingfang Hu, Chaomin Shen 0001, Yaxin Peng |
Neural Process. Lett. | 4 |
| 2016 | Joint distribution adaptation based TSK Fuzzy logic system for epileptic EEG signal identificationabstractTransfer learning based method, which utilizes plenty labeled data in the source domain to build an accuracy classifier for the target domain, serves as an effective means in the epileptic detection by using electroencephalogram (EEG) signals. Among existing approaches, Fuzzy logic system (FLS) based on transductive transfer learning is an efficient method due to its superior interpretability and strong learning abilities. However, this kind of method cannot simultaneously reduce the differences in both marginal distributions and conditional distributions between the training and test datasets of EEG signals. To overcome this problem, in this paper, we construct a Takagi-Sugeno-Kang (TSK) FLS based on the joint distribution adaptation (JDA), which refers to TSK-JDA-FLS. It aims to match both marginal and conditional distributions, and we extend the algorithm to perform a multi-class classification for identifying epileptic EEG signals. Extensive experiments verify that TSK-JDA-FLS significantly outperforms competitive non-transfer learning and transfer learning methods in the epileptic EEG datasets. Yaxin Peng, Guixu Zhang, Chaomin Shen 0001 |
BIBM | 4 |
| 2016 | Feature Selection in Click-Through Rate Prediction Based on Gradient Boosting
Qingsong Yu, Chaomin Shen 0001, Wenxin Hu |
IDEAL | 3 |
| 2016 | Single Image Super-Resolution Based on Nonlocal Sparse and Low-Rank Regularization
Chunhong Liu, Faming Fang, Chaomin Shen 0001 |
PRICAI | 4 |
| 2015 | Change Detection Using L 0 Smoothing and Superpixel TechniquesabstractWe propose an unsupervised change detection method for satellite images using $$L_0$$ smoothing, superpixel techniques and k-means. First, we produce the difference image according to image types (synthetic aperture radar or optical images). Second, we use $$L_0$$ smoothing, an image editing method that can simultaneously sharpen major edges and smooth low-amplitude structures, to generate two difference images with distinct smooth levels. Third, k-means algorithm with $$k=2$$ is applied on one smoothed difference image to cluster all pixels into changed or unchanged classes. Fourth, the Voronoi-Cells (VCells) algorithm is applied on the other difference image to obtain roughly uniform superpixels while preserving local image boundaries. Finally, we calculate the change degree for each superpixel, and the change detection map is produced by using k-means again. The novelties of this paper are that we use the $$L_0$$ smoothing to reduce noise and preserve edges, and utilize the spatial information with the help of superpixel. Experimental results on synthetic aperture radar and optical images show the effectiveness of our approach. Xiaoliang Shi, Guixu Zhang, Chaomin Shen 0001 |
KSEM | 4 |
| 2014 | Spectral unmixing using Lasso screening rulesabstractLasso (Least Absolute Shrinkage and Selection Operator) is a technique for selecting a sparse combination of given features. Spectral unmixing is a Lasso problem, regarding that the spectrum of every pixel is a linear combination of a small number of spectrums from a possibly very large spectral library. In this paper we apply a technique called screening to speedup the Lasso process for spectral unmixing. Our contribution is two-fold: we make use of the theoretical results to practical remote sensing problems; more importantly, we develop a tailored Lasso algorithm coupled with screening, as the unmixing requires that the fractions should be positive and sum to one. We also solve the high mutual coherence problem in the library by ticking out the spectrums with high mutual coherence. We use the AVIRIS data over Cuprite, Nevada (250 lines by 191 columns) to demonstrate the idea. Experiments demonstrate the effectiveness of the screening method. Chaomin Shen 0001, Xiaoliang Shi, Yaxin Peng |
IGARSS | 1 |
| 2014 | Framelet based pan-sharpening via a variational method
Faming Fang, Guixu Zhang, Fang Li 0004, Chaomin Shen 0001 |
Neurocomputing | 4 |
| 2013 | A Variational Approach for Pan-SharpeningabstractPan-sharpening is a process of acquiring a high resolution multispectral (MS) image by combining a low resolution MS image with a corresponding high resolution panchromatic (PAN) image. In this paper, we propose a new variational pan-sharpening method based on three basic assumptions: 1) the gradient of PAN image could be a linear combination of those of the pan-sharpened image bands; 2) the upsampled low resolution MS image could be a degraded form of the pan-sharpened image; and 3) the gradient in the spectrum direction of pan-sharpened image should be approximated to those of the upsampled low resolution MS image. An energy functional, whose minimizer is related to the best pan-sharpened result, is built based on these assumptions. We discuss the existence of minimizer of our energy and describe the numerical procedure based on the split Bregman algorithm. To verify the effectiveness of our method, we qualitatively and quantitatively compare it with some state-of-the-art schemes using QuickBird and IKONOS data. Particularly, we classify the existing quantitative measures into four categories and choose two representatives in each category for more reasonable quantitative evaluation. The results demonstrate the effectiveness and stability of our method in terms of the related evaluation benchmarks. Besides, the computation efficiency comparison with other variational methods also shows that our method is remarkable. Faming Fang, Fang Li 0004, Chaomin Shen 0001, Guixu Zhang |
IEEE Trans. Image Process. | 3 |
| 2010 | Multiplicative Noise Removal with Spatially Varying Regularization ParametersabstractThe Aubert–Aujol (AA) model is a variational method for multiplicative noise removal. In this paper, we study some basic properties of the regularization parameter in the AA model. We develop a method for automatically choosing the regularization parameter in the multiplicative noise removal process. In particular, we employ spatially varying regularization parameters in the AA model in order to restore more texture details of the denoised image. Experimental results are presented to demonstrate that the spatially varying regularization parameters method can obtain better denoised images than the other tested multiplicative noise removal methods. Fang Li 0004, Michael Kwok-Po Ng, Chaomin Shen 0001 |
SIAM J. Imaging Sci. | 3 |
| 2009 | Variational denoising of partly textured images
Fang Li 0004, Chaomin Shen 0001, Chunli Shen, Guixu Zhang |
J. Vis. Commun. Image Represent. | 2 |
| 2007 | Variational-based speckle noise removal of SAR imageryabstractIn this paper we present a variational method for synthetic aperture radar (SAR) speckle removal. Variational method is a newly developed technique for the removal of SAR's multiplicative noise. For an image, we could define an energy functional. The energy evolves as the original image changes, and the minimum energy corresponds to the speckle reduced result. Partial differential equation (PDE) technique is used to get the minimal solution. Our energy functional makes use of the statistical information of the multiplicative noise since it follows a Gamma law with mean mu = 1 and variance sigma2= 1/M for M-look SAR. Our energy is a regularization term with two constraints. The regularization term is the integral for the norm of image gradient; two constraints are the mean of noise should be 1 and the variance of noise should be 1/M. We use the method of Lagrange multipliers, Euler-Lagrange equation and heat flow method to obtain the minimizer of the energy. ERS Precision Image (PRI) data are to demonstrate our algorithm. Numerical result shows that the speckle reduced image preserves edges and point targets while smoothes homogenous regions in the original image. The algorithm is computationally efficient and easy to implement. Chaomin Shen 0001, Yaxin Peng, Ling Pi |
IGARSS | 1 |
| 2007 | A variational formulation for segmenting desired objects in color images
Ling Pi, Chaomin Shen 0001, Fang Li 0004, Jinsong Fan |
Image Vis. Comput. | 2 |
| 2007 | Image restoration combining a total variational filter and a fourth-order filter
Fang Li 0004, Chaomin Shen 0001, Jingsong Fan, Chunli Shen |
J. Vis. Commun. Image Represent. | 2 |