VLDB 2026 Research / reviewers in the wild / expert
Ke Chen 0004
dblp:47/6529-4
· DBLP profile ↗
77ranked-venue papers
12as first author
39since 2021 · last 2026
0000-0003-0928-5199ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorComputer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly DetectionabstractDeep learning-based 3D anomaly detection methods have demonstrated significant potential in industrial manufacturing. However, many approaches are specifically designed for anomaly detection tasks, which limits their generalizability to other 3D tasks. In contrast, self-supervised point cloud models aim for general representation learning, yet our investigation reveals that these classical models are suboptimal at anomaly detection under the unified fine-tuning paradigm. This motivates us to develop a more generalizable 3D model that can effectively detect anomalies without relying on task-specific designs. Interestingly, we find that using only the curvature of each point as its anomaly score already outperforms several classical self-supervised and dedicated anomaly detection models, highlighting the critical role of curvature in 3D anomaly detection. In this paper, we propose a Curvature-Augmented Self-supervised Learning (CASL) framework based on a reconstruction paradigm. Built upon the classical U-Net architecture, our approach introduces multi-scale curvature prompts to guide the decoder in predicting the coordinates of each point. Without relying on any dedicated anomaly detection mechanisms, it achieves leading detection performance through straightforward anomaly classification fine-tuning. Moreover, the learned representations generalize well to standard 3D understanding tasks such as point cloud classification. Yaohua Zha, Xue Yuerong, Chunlin Fan, Yuansong Wang, Tao Dai 0001, Ke Chen 0004, Shutao Xia |
AAAI | 6 |
| 2026 | UniRM: A Universal Large Model for Multiband 3D Radio Map ConstructionabstractRadio maps play a crucial role in optimizing wireless network performance and configuration, providing insights into the spatial distribution of radio frequency signal power. Existing solutions often face challenges in generalizing and adapting across various environments, frequency bands, and vertical dimensions. To overcome these limitations, we propose UniRM, a universal large model designed for constructing multiband 3D radio maps. UniRM leverages large-scale pre-training and prompt learning techniques to accurately generate radio maps across diverse environments, altitudes, and frequency bands. Specifically, UniRM employs a UNet-based encoder-decoder architecture during pre-training to extract universal latent representations that capture shared features across different environmental conditions. A prompt learning module further enhances this by transforming auxiliary inputs, such as environmental descriptions, frequency bands, and altitudes, into discriminative embeddings, thereby enabling effective cross-domain generalization and ensuring robustness in unseen scenarios. Extensive experiments using a large, diverse dataset covering numerous scenarios demonstrate that UniRM outperforms state-of-the-art baselines by over 10% in key metrics, including mean squared error, normalized mean squared error, root mean squared error, and peak signal-to-noise ratio. Notably, zero-shot evaluations highlight UniRM’s strong ability to generalize to new environments without retraining. The code for UniRM is available at: https://github.com/Shirleyue/UniRM. Tong Li 0013, Zhu Xiao, Ke Chen 0004, Shuai Ma 0002, Zhaocheng Wang 0001, Keqin Li 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2026 | CollaboRadio: A Hybrid Device-Edge-Cloud Collaboration Paradigm for Fine-Grained Radio Map ConstructionabstractRadio map presents communication parameters of interest, e.g., received signal strength, across a geographical region in a specific frequency band. It can be leveraged to improve the efficiency of spectrum utilization. With the rapid development of radio technology and the proliferation of radioenabled devices, there is an increasing demand for finer granularity (higher resolution) in radio maps. However, the problem of fine-grained radio map construction is uniquely challenging, as it requires to utilize an extremely small number of radio samples collected by sparsely distributed sensor devices to infer the map. To address the challenge, we propose a hybrid device-edge-cloud collaboration paradigm called CollaboRadio. CollaboRadio first groups sensor devices into multiple clusters, with cluster locations optimized based on principles of radio propagation. Next, it leverages a small AI model on each edge server to generate a local radio map for the respective cluster region from radio samples collected by intra-cluster sensors. Finally, it employs a large AI model in the cloud to construct a global radio map for the entire region from the local maps produced by edge servers of different clusters. For the implementation of CollaboRadio, we develop an UNet-based small model for the edge server and a Transformerbased large model for the cloud. Extensive simulations show that CollaboRadio is capable of constructing fine-grained radio maps from an ultra-low sampling rate of 0.1%, and significantly outperforms state-of-the-art. Shuai Shao 0014, Ke Chen 0004, Shuhang Zhang, Hongliang Zhang 0001, Lingyang Song |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | Topology-Aware Modeling for Unsupervised Simulation-to-Reality Point Cloud RecognitionabstractLearning semantic representations from point sets of 3D object shapes is often challenged by significant geometric variations, primarily due to differences in data acquisition methods. Typically, training data is generated using point simulators, while testing data is collected with distinct 3D sensors, leading to a simulation-to-reality (Sim2Real) domain gap that limits the generalization ability of point classifiers. Current unsupervised domain adaptation (UDA) techniques struggle with this gap, as they often lack robust, domain-insensitive descriptors capable of capturing global topological information, resulting in overfitting to the limited semantic patterns of the source domain. To address this issue, we introduce a novel Topology-Aware Modeling (TAM) framework for Sim2Real UDA on object point clouds. Our approach mitigates the domain gap by leveraging global spatial topology, characterized by low-level, high-frequency 3D structures, and by modeling the topological relations of local geometric features through a novel self-supervised learning task. Additionally, we propose an advanced self-training strategy that combines cross-domain contrastive learning with self-training, effectively reducing the impact of noisy pseudo-labels and enhancing the robustness of the adaptation process. Experimental results on three public Sim2Real benchmarks validate the effectiveness of our TAM framework, showing consistent improvements over state-of-the-art methods across all evaluated tasks. The source code of this work will be available athttps://github.com/zou-longkun/TAG.git. Longkun Zou, Kangjun Liu, Ke Chen 0004, Kailing Guo, Kui Jia, Yaowei Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Unsupervised Domain Adaptation on Point Cloud Classification via Imposing Structural Manifolds into Representation Space
Hongchao Zhong, Li Yu 0004, Longkun Zou, Ke Chen 0004 |
CVM (3) | 4 |
| 2025 | Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionabstractPoint cloud video understanding is becoming increasingly important in fields such as robotics, autonomous driving, and augmented reality, as they can accurately represent object motion and environmental changes. Despite the progress made in self-supervised learning methods for point cloud video understanding, the limited availability of 4D data and the high computational cost of training 4D-specific models remain significant obstacles. In this paper, we investigate the potential of transferring pre-trained static 3D point cloud models to the 4D domain, pointing out the limitations of static models that capture only spatial information while neglecting temporal dynamics. To address this, we propose a novel Cross-frame Spatio-temporal Adaptation (CSA) strategy by introducing the Point Tube Adapter as the embedding layer and the Geometric Constraint Temporal Adapter (GCTA) to enforce temporal consistency across frames. This strategy extracts both short-term and long-term temporal dynamics, effectively integrating them with spatial features and enriching the model’s understanding of temporal changes in point cloud videos. Extensive experiments on 3D action and gesture recognition tasks demonstrate that our method achieves state-of-the-art performance, establishing its effectiveness for point cloud video understanding. Code is available at: https://github.com/LvBaixuan/Point-CSA. Baixuan Lv, Yaohua Zha, Tao Dai 0001, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 5 |
| 2025 | PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba AdapterabstractApplying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementary information in the intermediate layer, thereby failing to fully unlock the potential of pre-trained models. To overcome this limitation, we propose an orthogonal solution: Point Mamba Adapter (PMA), which constructs an ordered feature sequence from all layers of the pre-trained model and leverages Mamba to fuse all complementary semantics, thereby promoting comprehensive point cloud understanding. Constructing this ordered sequence is non-trivial due to the inherent isotropy of 3D space. Therefore, we further propose a geometry-constrained gate prompt generator (G2PG) shared across different layers, which applies shared geometric constraints to the output gates of the Mamba and dynamically optimizes the spatial order, thus enabling more effective integration of multi-layer information. Extensive experiments conducted on challenging point cloud datasets across various tasks demonstrate that our PMA elevates the capability for point cloud understanding to a new level by fusing diverse complementary intermediate features. Code is available at https://github.com/zyh16143998882/PMA. Yaohua Zha, Yanzi Wang, Hang Guo 0002, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhihao Ouyang, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 9 |
| 2025 | ALLGCD: Leveraging All Unlabeled Data for Generalized Category Discovery
Xinzi Cao, Ke Chen 0004, Feidiao Yang, Xiawu Zheng, Yonghong Tian 0001, Yutong Lu |
ICCV | 2 |
| 2025 | T-Dreamer: Topology-Aware Text-to-3D GenerationabstractThe problem of generating 3D assets from text prompt (i.e., Text-to-3D generation) can be addressed by using optimization-based 2D diffusion models (or termed as 2D lifting algorithms), but 3D shape generated by these algorithms can suffer from the Janus problem caused by topological inconsistencies under different viewpoints. In light of this, we propose a novel topology-aware 2D diffusion method (namely, a T-Dreamer) to incorporate pose-specific topological information into the 2D lifting process. On one hand, our method can regularize the 2D diffusion model with favoring consistent topological structure with high-fidelity in a manner of introducing an extra loss on pose-specific high-order geometric priors (e.g., the 1st-order normal map, the 2nd-order curvature map) with low frequencies. On the other hand, the 2D lifting backbone is adapted with multi-granularity perspective information (directional text, continuous poses) to reduce rendering variations of 3D assets from diverse viewpoints. Extensive experiments demonstrate that the proposed T-Dreamer significantly alleviates topological inconsistencies of 3D assets, outperforming the state-of-the-art 3D shape generators. Source codes: https://github.com/Alisaxxwhg/T-Dreamer Xiaoxuan Wu, Qiulu Li, Ke Chen 0004 |
ICME | 5 |
| 2025 | Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised LearningabstractPoint clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on learning domain-specific 3D representations within a single domain, neglecting the complementary nature of cross-domain knowledge, which limits the learning of 3D representations. In this paper, we propose to learn a comprehensive Point cloud Mixture-of-Domain-Experts model (Point-MoDE) via a block-to-scene pre-training strategy. Specifically, We first propose a mixture-of-domain-expert model consisting of scene domain experts and multiple shared object domain experts. Furthermore, we propose a block-to-scene pretraining strategy, which leverages the features of point blocks in the object domain to regress their initial positions in the scene domain through object-level block mask reconstruction and scene-level block position regression. By integrating the complementary knowledge between object and scene, this strategy simultaneously facilitates the learning of both object-domain and scene-domain representations, leading to a more comprehensive 3D representation. Extensive experiments in downstream tasks demonstrate the superiority of our model. Yaohua Zha, Tao Dai 0001, Hang Guo 0002, Yanzi Wang, Bin Chen 0011, Ke Chen 0004, Shutao Xia |
IJCAI | 6 |
| 2025 | Optimal Distance-Constrained Path Planning for Sparse Radio Map RecoveryabstractIn urban environments, numerous IoT devices rely on accurate radio map information to enable critical applications such as localization, trajectory planning, and communication. The effectiveness of these applications hinges on the availability of high-fidelity and spatially comprehensive radio maps. To address this, we propose an optimal path planning approach with distance constraints for efficient sampling in urban radio map reconstruction. Unlike traditional random sampling strategies, our method restricts sampling points to lie along a feasible road network and imposes a fixed distance constraint on the sampling path. We design a structure-aware genetic algorithm (GA_s) to optimize the path for mobile sampling, incorporating an unsupervised population evaluation metric to assess the fitness of candidate solutions. Experimental results show that, under equal sampling distance conditions, GA_s achieves a Root Mean Square Error (RMSE) of 0.0626 in radio map recovery—outperforming random sampling (0.0796), A* algorithm (0.0748), and a conventional genetic algorithm (GA_c). These results demonstrate the effectiveness of our method in enabling efficient and accurate radio map reconstruction for mobile sampling platforms in real-world urban settings. Kangjun Liu, Longkun Zou, Ke Chen 0004 |
VTC2025-Fall | 5 |
| 2025 | LocVMunet: A Vision Mamba-Based Method for RSS-Driven Outdoor LocalizationabstractGlobal Navigation Satellite Systems (GNSS) and base station (BS)-based wireless positioning are the predominant technologies for outdoor user equipment (UE) localization. However, their performance deteriorates significantly in dense urban environments with complex architectural structures. This degradation arises because obstructions disrupt line-of-sight (LoS) conditions between the UE and satellites or base stations, reducing positioning accuracy. To address this challenge, this paper introduces LocVMunet, a novel deep neural network designed for high-precision localization based on received signal strength (RSS) from a limited number of base stations. The proposed LocVMunet seamlessly integrates the Cross-Scan Module (CSM) with signal distribution modeling, enhancing spatial feature extraction and improving localization accuracy. Unlike traditional methods, this neural network-based approach inherently mitigates the impact of non-line-of-sight (NLOS) conditions, enabling real-time localization across any region covered by the radio map. Experimental evaluations demonstrate that the LocVMunet establishes a new benchmark for RSS-based urban localization, significantly outperforming state-of-the-art neural network-based methods. Notably, it achieves a 24.8% positioning error reduction over LocUNet and a 22.7% improvement over LocSwinUnet in the five-base-station setup, highlighting its robustness and effectiveness in urban positioning. Chunyan Qiu, Kangjun Liu, Longkun Zou, Xinhuai Wang, Ke Chen 0004 |
VTC2025-Fall | 5 |
| 2025 | Fine-Grained Radio Map Construction from Ultra-Sparse Sampling: An Edge-Cloud Model Collaboration ParadigmabstractRadio map presents communication parameters of interest, e.g., received signal strength, across a geographical region. It can be leveraged to improve the efficiency of spectrum utilization. With the rapid development of radio technology and the proliferation of radio-enabled devices, there is an increasing demand for finer granularity (higher resolution) in radio maps. However, the problem of fine-grained radio map construction is uniquely challenging, as it requires to utilize an extremely small number of radio samples collected by sparsely distributed sensors to infer the map. To address the problem, we propose a novel edge-cloud model collaboration paradigm. This paradigm first groups sensors into multiple clusters, with cluster locations optimized based on principles of radio propagation. Next, a small model on each edge device generates a local radio map for the respective cluster region using intra-cluster radio samples. Finally, a large model in the cloud constructs a global radio map for the entire region from the local maps produced by edge devices of different clusters. To evaluate the proposed paradigm, we also develop an UNet-based small model for the edge device and a Transformer-based large model for the cloud. Extensive simulations show that our paradigm is capable of constructing fine-grained radio maps from an ultra-low sampling rate of 0.1%, and significantly outperforms state-of-the-art. Shuai Shao 0014, Ke Chen 0004, Shuhang Zhang, Lingyang Song |
VTC2025-Fall | 3 |
| 2025 | Nonlinearly activated gradient neuronet with fixed-time convergence for online equality-constrained quadratic programming
Xuanjiao Lv, Ke Chen 0004, Junying Yuan |
Neurocomputing | 2 |
| 2025 | Large Models for Aerial Edges: An Edge-Cloud Model Evolution and Communication ParadigmabstractThe future sixth-generation (6G) of wireless networks is expected to surpass its predecessors by offering ubiquitous coverage through integrated air-ground deployments in both communication and computing domains. In such networks, aerial platforms, such as unmanned aerial vehicles (UAVs), conduct artificial intelligence (AI) computations based on multi-modal data to support diverse applications including surveillance and environment construction. However, these multi-domain inference and content generation tasks require large AI models, demanding powerful computing capabilities and finely tuned inference models trained on rich datasets, thus posing significant challenges for UAVs. To tackle this problem, we propose an integrated air-ground edge-cloud model framework, in which UAVs serve as edge nodes for data collection and small model computation. Through wireless channels, UAVs collaborate with ground cloud servers providing large model computation and model updating for edge UAVs. With limited wireless communication bandwidth, the proposed framework faces the challenge of information exchange scheduling between the edge UAVs and the cloud server. To tackle this, we present joint task allocation, transmission resource allocation, transmission data quantization design, and edge model update design to enhance the inference accuracy of the integrated air-ground edge-cloud model evolution framework by mean average precision (mAP) maximization. A closed-form lower bound on the mAP of the proposed framework is derived based on the mAP of the edge model and mAP of the cloud model, and the solution to the mAP maximization problem is optimized accordingly. Simulations, based on results from vision-based classification experiments, consistently demonstrate that the mAP of the proposed integrated air-ground edge-cloud model evolution framework outperforms both a centralized cloud model framework and a distributed edge model framework across various communication bandwidths and data sizes. Shuhang Zhang, Ke Chen 0004, Boya Di, Hongliang Zhang 0001, Wenhan Yang, Dusit Niyato, Zhu Han 0001, H. Vincent Poor |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Improving deep representation learning via auxiliary learnable target coding
Kangjun Liu, Ke Chen 0004, Kui Jia, Yaowei Wang 0001 |
Pattern Recognit. | 2 |
| 2025 | Bridging Domain Gap of Point Cloud Representations via Self-Supervised Geometric AugmentationabstractRecent progress of semantic point clouds analysis is largely driven by synthetic data (e.g., the ModelNet and the ShapeNet), which are typically complete, well-aligned and noisy-free. Therefore, representations of those ideal synthetic point clouds have limited variations in the geometric perspective and can gain good performance on a number of 3D vision tasks such as point cloud classification. In the context of unsupervised domain adaptation (UDA), representation learning designed for synthetic point clouds can hardly capture domain invariant geometric patterns from incomplete and noisy point clouds. To address such a problem, we introduce a novel scheme for induced geometric invariance of point cloud representations across domains, via regularizing representation learning with two self-supervised geometric augmentation tasks. On one hand, a novel pretext task of predicting translation distances of augmented samples is proposed to alleviate centroid shift of point clouds due to occlusion and noises. On the other hand, we pioneer an integration of the self-supervised relational learning on geometrically-augmented point clouds in a cascade manner, utilizing the intrinsic relationship of augmented variants and other samples as extra constraints of cross-domain geometric features. Experiments on the PointDA-10 dataset demonstrate the effectiveness of the proposed method, achieving the state-of-the-art performance. Li Yu 0004, Hongchao Zhong, Longkun Zou, Ke Chen 0004, Pan Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Adversarial Geometric Transformations of Point Clouds for Physical Attack
Jingyu Xiang, Xuanxiang Lin, Ke Chen 0004, Kui Jia |
CVM (1) | 3 |
| 2024 | RobUNet: A Radio Map Construction Method with A Strong Generalization CapabilityabstractRadio map presents communication parameters of interest, e.g., radio power, across a geographical region in a specific frequency band. It can be leveraged to improve the efficiency of spectrum utilization. The problem of constructing a radio map involves utilizing measurements from sparsely distributed sensors to infer the parameter of interest at every point in the region. One major challenge of solving this problem is ensuring a strong generalization capability for the solution, i.e., ensuring that the solution is able to make accurate inferences even in the presence of unknown system variables that can affect radio propagation, such as obstacle layouts and weather conditions. This paper focuses on analyzing and optimizing the generalization capability of radio map construction, an aspect that has been neglected in prior research as far as we know. We design RobUNet—a UNet-based algorithm that captures multi-scale features of radio maps. It further leverages mechanisms of residual connection, channel attention, and pixel attention to enhance its generalization capability. Simulations based on real-world dataset show that (i) the radio maps constructed by RobUNet are much more accurate than those constructed by many existing solutions in the presence of unknown system variables, and (ii) the accuracy of the radio maps constructed by RobUNet when system variables are unknown is close to that of the radio maps constructed by RobUNet when system variables are known. As a result, the generalization capability of RobUNet is much stronger than that of state-of-the-art. Shuai Shao 0014, Kangjun Liu, Shuhang Zhang, Ke Chen 0004, Lingyang Song |
GLOBECOM | 5 |
| 2024 | TransWild: Enhancing 3D interacting hands recovery in the wild with IoU-guided Transformer
Wanru Zhu, Ke Chen 0004, Lihua Guo |
Image Vis. Comput. | 3 |
| 2024 | Boosting Cross-Domain Point Classification via Distilling Relational Priors From 2D TransformersabstractSemantic pattern of an object point cloud is determined by its topological configuration of local geometries. Learning discriminative representations can be challenging due to large shape variations of point sets in local regions and incomplete surface in a global perspective, which can be made even more severe in the context of unsupervised domain adaptation (UDA). In specific, traditional 3D networks mainly focus on local geometric details and ignore the topological structure between local geometries, which greatly limits their cross-domain generalization. Recently, the transformer-based models have achieved impressive performance gain in a range of image-based tasks, benefiting from its strong generalization capability and scalability stemming from capturing long range correlation across local patches. Inspired by such successes of visual transformers, we propose a novel Relational Priors Distillation (RPD) method to extract relational priors from the well-trained transformers on massive images, which can significantly empower cross-domain representations with consistent topological priors of objects. To this end, we establish a parameter-frozen pre-trained transformer module shared between 2D teacher and 3D student models, complemented by an online knowledge distillation strategy for semantically regularizing the 3D student model. Furthermore, we introduce a novel self-supervised task centered on reconstructing masked point cloud patches using corresponding masked multi-view image features, thereby empowering the model with incorporating 3D geometric information. Experiments on the PointDA-10 and the Sim-to-Real datasets verify that the proposed method consistently achieves the state-of-the-art performance of UDA for point cloud classification. The source code of this work is available athttps://github.com/zou-longkun/RPD.git. Longkun Zou, Wanru Zhu, Ke Chen 0004, Lihua Guo, Kailing Guo, Kui Jia, Yaowei Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Quality-Aware Self-Training on Differentiable Synthesis of Rare Relational DataabstractData scarcity is a very common real-world problem that poses a major challenge to data-driven analytics. Although a lot of data-balancing approaches have been proposed to mitigate this problem, they may drop some useful information or fall into the overfitting problem. Generative Adversarial Network (GAN) based data synthesis methods can alleviate such a problem but lack of quality control over the generated samples. Moreover, the latent associations between the attribute set and the class labels in a relational data cannot be easily captured by a vanilla GAN. In light of this, we introduce an end-to-end self-training scheme (namely, Quality-Aware Self-Training) for rare relational data synthesis, which generates labeled synthetic data via pseudo labeling on GAN-based synthesis. We design a semantic pseudo labeling module to first control the quality of the generated features/samples, then calibrate their semantic labels via a classifier committee consisting of multiple pre-trained shallow classifiers. The high-confident generated samples with calibrated pseudo labels are then fed into a semantic classification network as augmented samples for self-training. We conduct extensive experiments on 20 benchmark datasets of different domains, including 14 industrial datasets. The results show that our method significantly outperforms state-of-the-art methods, including two recent GAN-based data synthesis schemes. Codes are available at https://github.com/yaxinhou/QAST. Chongsheng Zhang, Yaxin Hou, Ke Chen 0004, Shuang Cao, Gaojuan Fan, Ji Liu 0003 |
AAAI | 3 |
| 2023 | Manifold-Aware Self-Training for Unsupervised Domain Adaptation on Regressing 6D Object PoseabstractDomain gap between synthetic and real data in visual regression (e.g., 6D pose estimation) is bridged in this paper via global feature alignment and local refinement on the coarse classification of discretized anchor classes in target space, which imposes a piece-wise target manifold regularization into domain-invariant representation learning. Specifically, our method incorporates an explicit self-supervised manifold regularization, revealing consistent cumulative target dependency across domains, to a self-training scheme (e.g., the popular Self-Paced Self-Training) to encourage more discriminative transferable representations of regression tasks. Moreover, learning unified implicit neural functions to estimate relative direction and distance of targets to their nearest class bins aims to refine target classification predictions, which can gain robust performance against inconsistent feature scaling sensitive to UDA regressors. Experiment results on three public benchmarks of the challenging 6D pose estimation task can verify the effectiveness of our method, consistently achieving superior performance to the state-of-the-art for UDA on 6D pose estimation. Codes and pre-trained models are available https://github.com/Gorilla-Lab-SCUT/MAST. Jiehong Lin, Ke Chen 0004, Zelin Xu 0002, Yaowei Wang 0001, Kui Jia |
IJCAI | 3 |
| 2023 | Classification of single-view object point clouds
Zelin Xu 0002, Kangjun Liu, Ke Chen 0004, Changxing Ding, Yaowei Wang 0001, Kui Jia |
Pattern Recognit. | 3 |
| 2022 | Fine-Grained Object Classification via Self-Supervised Pose AlignmentabstractSemantic patterns offine-grained objects are determined by subtle appearance difference of local parts, which thus inspires a number of part-based methods. However, due to uncontrollable object poses in images, distinctive de-tails carried by local regions can be spatially distributed or even self-occluded, leading to a large variation on ob-ject representation. For discounting pose variations, this paper proposes to learn a novel graph based object rep-resentation to reveal a global configuration of local parts for self-supervised pose alignment across classes, which is employed as an auxiliary feature regularization on a deep representation learning network. Moreover, a coarse-to-fine supervision together with the proposed pose-insensitive constraint on shallow-to-deep sub-networks encourages discriminative features in a curriculum learning manner. We evaluate our method on three popular fine-grained ob-ject classification benchmarks, consistently achieving the state-of-the-art performance. Source codes are available at https://github.com/yangxhll/P2P-Net. Xuhui Yang, Yaowei Wang 0001, Ke Chen 0004, Yong Xu 0007, Yonghong Tian 0001 |
CVPR | 3 |
| 2022 | Quasi-Balanced Self-Training on Noise-Aware Synthesis of Object Point Clouds for Closing Domain Gap
Yongwei Chen, Longkun Zou, Ke Chen 0004, Kui Jia |
ECCV (33) | 4 |
| 2022 | BiCo-Net: Regress Globally, Match Locally for Robust 6D Pose EstimationabstractThe challenges of learning a robust 6D pose function lie in 1) severe occlusion and 2) systematic noises in depth images. Inspired by the success of point-pair features, the goal of this paper is to recover the 6D pose of an object instance segmented from RGB-D images by locally matching pairs of oriented points between the model and camera space. To this end, we propose a novel Bi-directional Correspondence Mapping Network (BiCo-Net) to first generate point clouds guided by a typical pose regression, which can thus incorporate pose-sensitive information to optimize generation of local coordinates and their normal vectors. As pose predictions via geometric computation only rely on one single pair of local oriented points, our BiCo-Net can achieve robustness against sparse and occluded point clouds. An ensemble of redundant pose predictions from locally matching and direct pose regression further refines final pose output against noisy observations. Experimental results on three popularly benchmarking datasets can verify that our method can achieve state-of-the-art performance, especially for the more challenging severe occluded scenes. Source codes are available at https://github.com/Gorilla-Lab-SCUT/BiCo-Net. Zelin Xu 0002, Ke Chen 0004, Kui Jia |
IJCAI | 3 |
| 2022 | Data-Driven Oracle Bone Rejoining: A Dataset and Practical Self-Supervised Learning SchemeabstractOracle Bone Inscriptions (OBI) is one of the oldest scripts in the world. The rejoining of Oracle Bone (OB) fragments is of vital importance to the research of ancient scripts and history. Although significant progress has been achieved in the past decades, the rejoining work still heavily relies on domain knowledge and manual work, thus remains a low efficient and time-consuming process Therefore, an automatic and practical algorithm/system for OB rejoining is of great value to the OBI community. To this end, we collect a real-world dataset for rejoining Oracle Bone fragments, namely OB-Rejoin, which consists of 998 OB rubbing images that suffer from low quality image problems, due to intrinsic underground eroding over time and extrinsic imaging conditions in the past. Moreover, a practical Self-Supervised Splicing Network, S3-Net, is proposed to rejoin the OB fragments based on shape similarity of their borderlines. Specifically, we first transform the manually annotated borderline strokes of OB images into times series style shape representations, which are fed as input to a Generative Adversarial Network for augmenting positive pairs of rejoinable OBs for each OB fragment that does not have rejoinable counterparts. A Siamese network is trained on such augmented data in a contrastive learning manner to retrieve the matching OB fragments of an unseen query from an OB fragment gallery. Experiments on the OB-Rejoin benchmark show that our data-driven approach outperforms two recent methods for time-series analysis. In order to demonstrate its practical potential, we deploy the proposed S3-Net method in real tests and ultimately discover dozens of new rejoinings missed by domain experts for decades. Chongsheng Zhang, Bin Wang 0063, Ke Chen 0004, Ruixing Zong, Bofeng Mo, Yi Men, George Almpanidis, Shanxiong Chen, Xiangliang Zhang 0001 |
KDD | 3 |
| 2022 | Towards Uncovering the Intrinsic Data Structures for Unsupervised Domain Adaptation Using Structurally Regularized Deep ClusteringabstractUnsupervised domain adaptation (UDA) is to learn classification models that make predictions for unlabeled data on a target domain, given labeled data on a source domain whose distribution diverges from the target one. Mainstream UDA methods strive to learn domain-aligned features such that classifiers trained on the source features can be readily applied to the target ones. Although impressive results have been achieved, these methods have a potential risk of damaging the intrinsic data structures of target discrimination, raising an issue of generalization particularly for UDA tasks in an inductive setting. To address this issue, we are motivated by a UDA assumption of structural similarity across domains, and propose to directly uncover the intrinsic target discrimination via constrained clustering, where we constrain the clustering solutions using structural source regularization that hinges on the very same assumption. Technically, we propose a hybrid model of Structurally Regularized Deep Clustering, which integrates the regularized discriminative clustering of target data with a generative one, and we thus term our method as H-SRDC. Our hybrid model is based on a deep clustering framework that minimizes the Kullback-Leibler divergence between the distribution of network prediction and an auxiliary one, where we impose structural regularization by learning domain-shared classifier and cluster centroids. By enriching the structural similarity assumption, we are able to extend H-SRDC for a pixel-level UDA task of semantic segmentation. We conduct extensive experiments on seven UDA benchmarks of image classification and semantic segmentation. With no explicit feature alignment, our proposed H-SRDC outperforms all the existing methods under both the inductive and transductive settings. We make our implementation codes publicly available at https://github.com/huitangtang/H-SRDC. Xiatian Zhu, Ke Chen 0004, Kui Jia, C. L. Philip Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Geometry-Aware Generation of Adversarial Point CloudsabstractMachine learning models have been shown to be vulnerable to adversarial examples. While most of the existing methods for adversarial attack and defense work on the 2D image domain, a few recent attempts have been made to extend them to 3D point cloud data. However, adversarial results obtained by these methods typically contain point outliers, which are both noticeable and easy to defend against using the simple techniques of outlier removal. Motivated by the different mechanisms by which humans perceive 2D images and 3D shapes, in this paper we propose the new design ofgeometry-aware objectives, whose solutions favor (the discrete versions of) the desired surface properties of smoothness and fairness. To generate adversarial point clouds, we use a targeted attack misclassification loss that supports continuous pursuit of increasingly malicious signals. Regularizing the targeted attack loss with our proposed geometry-aware objectives results in our proposed method,Geometry-Aware Adversarial Attack ($GeoA^3$GeoA3). The results of$GeoA^3$tend to be more harmful, arguably harder to defend against, and of the key adversarial characterization of being imperceptible to humans. While the main focus of this paper is to learn to generate adversarial point clouds, we also present a simple but effective algorithm termed$Geo_{+}A^3$-IterNormPro, with Iterative Normal Projection (IterNorPro) that solves a new objective function$Geo_{+}A^3$, towards surface-level adversarial attacks via generation of adversarial point clouds. We quantitatively evaluate our methods on both synthetic and physical objects in terms of attack success rate and geometric regularity. For a qualitative evaluation, we conduct subjective studies by collecting human preferences from Amazon Mechanical Turk. Comparative results in comprehensive experiments confirm the advantages of our proposed methods. Our source codes are publicly available athttps://github.com/Yuxin-Wen/GeoA3. Yuxin Wen, Jiehong Lin, Ke Chen 0004, C. L. Philip Chen, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Soft pseudo-Label shrinkage for unsupervised domain adaptive person re-identification
Dingyuan Zheng, Jimin Xiao, Ke Chen 0004, Xiaowei Huang 0001, Yao Zhao 0001 |
Pattern Recognit. | 3 |
| 2022 | Improving Semantic Analysis on Point Clouds via Auxiliary Supervision of Local Geometric PriorsabstractExisting deep learning algorithms for point cloud analysis mainly concern discovering semantic patterns from the global configuration of local geometries in a supervised learning manner. However, very few explore geometric properties revealing local surface manifolds embedded in 3-D Euclidean space to discriminate semantic classes or object parts as additional supervision signals. This article is the first attempt to propose a unique multitask geometric learning network to improve semantic analysis by auxiliary geometric learning with local shape properties, which can be either generated via physical computation from point clouds themselves as self-supervision signals or provided as privileged information. Owing to explicitly encoding local shape manifolds in favor of semantic analysis, the proposed geometric self-supervised and privileged learning algorithms can achieve superior performance to their backbone baselines and other state-of-the-art methods, which are verified in the experiments on the popular benchmarks. Lulu Tang, Ke Chen 0004, Chaozheng Wu, Kui Jia, Zhi-Xin Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Convolutional Fine-Grained Classification With Self-Supervised Target Relation RegularizationabstractFine-grained visual classification can be addressed by deep representation learning under supervision of manually pre-defined targets (e.g., one-hot or the Hadamard codes). Such target coding schemes are less flexible to model inter-class correlation and are sensitive to sparse and imbalanced data distribution as well. In light of this, this paper introduces a novel target coding scheme - dynamic target relation graphs (DTRG), which, as an auxiliary feature regularization, is a self-generated structural output to be mapped from input images. Specifically, online computation of class-level feature centers is designed to generate cross-category distance in the representation space, which can thus be depicted by a dynamic graph in a non-parametric manner. Explicitly minimizing intra-class feature variations anchored on those class-level centers can encourage learning of discriminative features. Moreover, owing to exploiting inter-class dependency, the proposed target graphs can alleviate data sparsity and imbalanceness in representation learning. Inspired by recent success of the mixup style data augmentation, this paper introduces randomness into soft construction of dynamic target relation graphs to further explore relation diversity of target classes. Experimental results can demonstrate the effectiveness of our method on a number of diverse benchmarks of multiple visual classification, especially achieving the state-of-the-art performance on three popular fine-grained object benchmarks and superior robustness against sparse and imbalanced data. Source codes are made publicly available at https://github.com/AkonLau/DTRG. Kangjun Liu, Ke Chen 0004, Kui Jia |
IEEE Trans. Image Process. | 2 |
| 2022 | Attribute-Aware Feature Encoding for Object Recognition and SegmentationabstractExisting multi-task models for object recognition and segmentation have verified the effectiveness of joint optimization of two semantic tasks. However, learning discriminative representations with insufficient training data and redundant contextual information from the background remains challenging. Semantic attributes are designed as powerful and informative mid-level features that 1) share information across categories to model the interclass correlation and that 2) can be localized in the object region to benefit foreground extraction. This paper introduces a novel attribute-aware feature encoding (AFE) module to a multi-task network for object recognition and segmentation with the aim of improving both semantic tasks by regularizing feature encoding with auxiliary attribute learning. Intuitively, attribute learning in our method not only provides extra supervision signals to capture interclass correlation in object classification but also refines the output of object segmentation via weakly supervised attribute localization. The experimental results on two public benchmarks show that our method yields remarkable improvement in both semantic tasks and auxiliary attribute estimation over existing methods. Shu Yang 0007, Yaowei Wang 0001, Ke Chen 0004, Wei Zeng 0006, Zesong Fei |
IEEE Trans. Multim. | 3 |
| 2021 | 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingabstractThe ability to understand the ways to interact with objects from visual cues, a.k.a. visual affordance, is essential to vision-guided robotic research. This involves categorizing, segmenting and reasoning of visual affordance. Relevant studies in 2D and 2.5D image domains have been made previously, however, a truly functional understanding of object affordance requires learning and prediction in the 3D physical domain, which is still absent in the community. In this work, we present a 3D AffordanceNet dataset, a bench-mark of 23k shapes from 23 semantic object categories, annotated with 18 visual affordance categories. Based on this dataset, we provide three benchmarking tasks for evaluating visual affordance understanding, including full-shape, partial-view and rotation-invariant affordance estimations. Three state-of-the-art point cloud deep learning networks are evaluated on all tasks. In addition we also investigate a semi-supervised learning setup to explore the possibility to benefit from unlabeled data. Comprehensive results on our contributed dataset show the promise of visual affordance understanding as a valuable yet challenging benchmark. Shengheng Deng, Xun Xu 0002, Chaozheng Wu, Ke Chen 0004, Kui Jia |
CVPR | 4 |
| 2021 | Geometry-Aware Self-Training for Unsupervised Domain Adaptation on Object Point CloudsabstractThe point cloud representation of an object can have a large geometric variation in view of inconsistent data acquisition procedure, which thus leads to domain discrepancy due to diverse and uncontrollable shape representation cross datasets. To improve discrimination on unseen distribution of point-based geometries in a practical and feasible perspective, this paper proposes a new method of geometry-aware self-training (GAST) for unsupervised domain adaptation of object point cloud classification. Specifically, this paper aims to learn a domain-shared representation of semantic categories, via two novel self-supervised geometric learning tasks as feature regularization. On one hand, the representation learning is empowered by a linear mixup of point cloud samples with their self-generated rotation labels, to capture a global topological configuration of local geometries. On the other hand, a diverse point distribution across datasets can be normalized with a novel curvature-aware distortion localization. Experiments on the PointDA-10 dataset show that our GAST method can significantly outperform the state-of-the-art methods. Source codes and pre-trained models are available at https://github.com/zou-longkun/GAST. Longkun Zou, Ke Chen 0004, Kui Jia |
ICCV | 3 |
| 2021 | Object Point Cloud Classification via Poly-Convolutional Architecture SearchabstractExisting point cloud classifiers concern on handling irregular data structures to discover a global and discriminative configuration of local geometries. These classification methods design a number of effective permutation-invariant feature encoding kernels, but still suffer from the intrinsic challenge of large geometric feature variations caused by inconsistent point distributions along object surface. In this paper, point cloud classification can be addressed via deep graph representation learning on aggregating multiple convolutional feature kernels (namely, a poly convolutional operation) anchored on each point with its local neighbours. Inspired by recent success of neural architecture search, we introduce a novel concept of poly-convolutional architecture search (PolyConv search in short) to model local geometric patterns in a more flexible manner. Xuanxiang Lin, Ke Chen 0004, Kui Jia |
ACM Multimedia | 2 |
| 2021 | Sparse Steerable Convolutions: An Efficient Learning of SE(3)-Equivariant Features for Estimation and Tracking of Object Poses in 3D SpaceabstractAs a basic component of SE(3)-equivariant deep feature learning, steerable convolution has recently demonstrated its advantages for 3D semantic analysis. The advantages are, however, brought by expensive computations on dense, volumetric data, which prevent its practical use for efficient processing of 3D data that are inherently sparse. In this paper, we propose a novel design of Sparse Steerable Convolution (SS-Conv) to address the shortcoming; SS-Conv greatly accelerates steerable convolution with sparse tensors, while strictly preserving the property of SE(3)-equivariance. Based on SS-Conv, we propose a general pipeline for precise estimation of object poses, wherein a key design is a Feature-Steering module that takes the full advantage of SE(3)-equivariance and is able to conduct an efficient pose refinement. To verify our designs, we conduct thorough experiments on three tasks of 3D object semantic analysis, including instance-level 6D pose estimation, category-level 6D pose and size estimation, and category-level 6D pose tracking. Our proposed pipeline based on SS-Conv outperforms existing methods on almost all the metrics evaluated by the three tasks. Ablation studies also show the superiority of our SS-Conv over alternative convolutions in terms of both accuracy and efficiency. Our code is released publicly at https://github.com/Gorilla-Lab-SCUT/SS-Conv. Jiehong Lin, Ke Chen 0004, Jiangbo Lu, Kui Jia |
NeurIPS | 3 |
| 2021 | New Noise-Tolerant Neural Algorithms for Future Dynamic Nonlinear Optimization With Estimation on Hessian Matrix InversionabstractNonlinear optimization problems with dynamical parameters are widely arising in many practical scientific and engineering applications, and various computational models are presented for solving them under the hypothesis of short-time invariance. To eliminate the large lagging error in the solution of the inherently dynamic nonlinear optimization problem, the only way is to estimate the future unknown information by using the present and previous data during the solving process, which is termed the future dynamic nonlinear optimization (FDNO) problem. In this paper, to suppress noises and improve the accuracy in solving FDNO problems, a novel noise-tolerant neural (NTN) algorithm based on zeroing neural dynamics is proposed and investigated. In addition, for reducing algorithm complexity, the quasi-Newton Broyden-Fletcher-Goldfarb-Shanno (BFGS) method is employed to eliminate the intensively computational burden for matrix inversion, termed NTN-BFGS algorithm. Moreover, theoretical analyses are conducted, which show that the proposed algorithms are able to globally converge to a tiny error bound with or without the pollution of noises. Finally, numerical experiments are conducted to validate the superiority of the proposed NTN and NTN-BFGS algorithms for the online solution of FDNO problems. Long Jin 0001, Chenguang Yang 0001, Ke Chen 0004, Weibing Li |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Cascading Convolutional Color ConstancyabstractRegressing the illumination of a scene from the representations of object appearances is popularly adopted in computational color constancy. However, it's still challenging due to intrinsic appearance and label ambiguities caused by unknown illuminants, diverse reflection properties of materials and extrinsic imaging factors (such as different camera sensors). In this paper, we introduce a novel algorithm – Cascading Convolutional Color Constancy (in short, C4) to improve robustness of regression learning and achieve stable generalization capability across datasets (different cameras and scenes) in a unique framework. The proposed C4 method ensembles a series of dependent illumination hypotheses from each cascade stage via introducing a weighted multiply-accumulate loss function, which can inherently capture different modes of illuminations and explicitly enforce coarse-to-fine network optimization. Experimental results on the public Color Checker and NUS 8-Camera benchmarks demonstrate superior performance of the proposed algorithm in comparison with the state-of-the-art methods, especially for more difficult scenes. Huanglin Yu, Ke Chen 0004, Kaiqi Wang, Yanlin Qian, Zhaoxiang Zhang 0001, Kui Jia |
AAAI | 2 |
| 2020 | Unsupervised Domain Adaptation via Structurally Regularized Deep ClusteringabstractUnsupervised domain adaptation (UDA) is to make predictions for unlabeled data on a target domain, given labeled data on a source domain whose distribution shifts from the target one. Mainstream UDA methods learn aligned features between the two domains, such that a classifier trained on the source features can be readily applied to the target ones. However, such a transferring strategy has a potential risk of damaging the intrinsic discrimination of target data. To alleviate this risk, we are motivated by the assumption of structural domain similarity, and propose to directly uncover the intrinsic target discrimination via discriminative clustering of target data. We constrain the clustering solutions using structural source regularization that hinges on our assumed structural domain similarity. Technically, we use a flexible framework of deep network based discriminative clustering that minimizes the KL divergence between predictive label distribution of the network and an introduced auxiliary one; replacing the auxiliary distribution with that formed by ground-truth labels of source data implements the structural source regularization via a simple strategy of joint network training. We term our proposed method as Structurally Regularized Deep Clustering (SRDC), where we also enhance target discrimination with clustering of intermediate network features, and enhance structural regularization with soft selection of less divergent source examples. Careful ablation studies show the efficacy of our proposed SRDC. Notably, with no explicit domain alignment, SRDC outperforms all existing methods on three UDA benchmarks. Ke Chen 0004, Kui Jia |
CVPR | 2 |
| 2020 | Compositional Few-Shot Recognition with Primitive Discovery and EnhancingabstractFew-shot learning (FSL) aims at recognizing novel classes given only few training samples, which still remains a great challenge for deep learning. However, humans can easily recognize novel classes with only few samples. A key component of such ability is the compositional recognition that human can perform, which has been well studied in cognitive science but is not well explored in FSL. Inspired by such capability of humans, to imitate humans' ability of learning visual primitives and composing primitives to recognize novel classes, we propose an approach to FSL to learn a feature representation composed of important primitives, which is jointly trained with two parts, i.e. primitive discovery and primitive enhancing. In primitive discovery, we focus on learning primitives related to object parts by self-supervision from the order of image splits, avoiding extra laborious annotations and alleviating the effect of semantic gaps. In primitive enhancing, inspired by current studies on the interpretability of deep networks, we provide our composition view for the FSL baseline model. To modify this model for effective composition, inspired by both mathematical deduction and biological studies (the Hebbian Learning rule and the Winner-Take-All mechanism), we propose a soft composition mechanism by enlarging the activation of important primitives while reducing that of others, so as to enhance the influence of important primitives and better utilize these primitives to compose novel classes. Extensive experiments on public benchmarks are conducted on both the few-shot image classification and video recognition tasks. Our method achieves the state-of-the-art performance on all these datasets and shows better interpretability. Yixiong Zou, Shanghang Zhang, Ke Chen 0004, Yonghong Tian 0001, Yaowei Wang 0001, José M. F. Moura |
ACM Multimedia | 3 |
| 2020 | On the investigation of activation functions in gradient neural network for online solving linear matrix equation
Zhiguo Tan, Yueming Hu 0002, Ke Chen 0004 |
Neurocomputing | 3 |
| 2019 | Deep Cascade Generation on Point SetsabstractThis paper proposes a deep cascade network to generate 3D geometry of an object on a point cloud, consisting of a set of permutation-insensitive points. Such a surface representation is easy to learn from, but inhibits exploiting rich low-dimensional topological manifolds of the object shape due to lack of geometric connectivity. For benefiting from its simple structure yet utilizing rich neighborhood information across points, this paper proposes a two-stage cascade model on point sets. Specifically, our method adopts the state-of-the-art point set autoencoder to generate a sparsely coarse shape first, and then locally refines it by encoding neighborhood connectivity on a graph representation. An ensemble of sparse refined surface is designed to alleviate the suffering from local minima caused by modeling complex geometric manifolds. Moreover, our model develops a dynamically-weighted loss function for jointly penalizing the generation output of cascade levels at different training stages in a coarse-to-fine manner. Comparative evaluation on the publicly benchmarking ShapeNet dataset demonstrates superior performance of the proposed model to the state-of-the-art methods on both single-view shape reconstruction and shape autoencoding applications. Kaiqi Wang, Ke Chen 0004, Kui Jia |
IJCAI | 2 |
| 2019 | Nonlinear gradient neural network for solving system of linear equations
Lin Xiao 0002, Kenli Li 0001, Zhiguo Tan, Zhijun Zhang 0003, Bolin Liao, Ke Chen 0004, Long Jin 0001, Shuai Li 0002 |
Inf. Process. Lett. | 6 |
| 2019 | A new noise-tolerant and predefined-time ZNN model for time-dependent matrix inversion
Lin Xiao 0002, Jianhua Dai 0003, Ke Chen 0004, Weibing Li, Bolin Liao, Lei Ding 0007, Jichun Li 0002 |
Neural Networks | 4 |
| 2019 | Cumulative attribute space regression for head pose estimation and color constancyabstractTwo-stage Cumulative Attribute (CA) regression has been found effective in regression problems of computer vision such as facial age and crowd density estimation. The first stage regression maps input features to cumulative attributes that encode correlations between target values. The previous works have dealt with single output regression. In this work, we propose cumulative attribute spaces for 2- and 3-output (multivariate) regression. We show how the original CA space can be generalized to multiple output by the Cartesian product (CartCA). However, for target spaces with more than two outputs the CartCA becomes computationally infeasible and therefore we propose an approximate solution - multi-view CA (MvCA) - where CartCA is applied to output pairs. We experimentally verify improved performance of the CartCA and MvCA spaces in 2D and 3D face pose estimation and three-output (RGB) illuminant estimation for color constancy. Ke Chen 0004, Kui Jia, Heikki Huttunen, Jiri Matas, Joni-Kristian Kämäräinen |
Pattern Recognit. | 1 |
| 2019 | Convolutional low-resolution fine-grained classification
Dingding Cai, Ke Chen 0004, Yanlin Qian, Joni-Kristian Kämäräinen |
Pattern Recognit. Lett. | 2 |
| 2018 | Nonlinear recurrent neural networks for finite-time solution of general time-varying linear matrix equations
Lin Xiao 0002, Bolin Liao, Shuai Li 0002, Ke Chen 0004 |
Neural Networks | 4 |
| 2018 | Generalized Multi-View Embedding for Visual Recognition and Cross-Modal RetrievalabstractIn this paper, the problem of multi-view embedding from different visual cues and modalities is considered. We propose a unified solution for subspace learning methods using the Rayleigh quotient, which is extensible for multiple views, supervised learning, and nonlinear embeddings. Numerous methods including canonical correlation analysis, partial least square regression, and linear discriminant analysis are studied using specific intrinsic and penalty graphs within the same framework. Nonlinear extensions based on kernels and (deep) neural networks are derived, achieving better performance than the linear ones. Moreover, a novel multi-view modular discriminant analysis is proposed by taking the view difference into consideration. We demonstrate the effectiveness of the proposed multi-view embedding methods on visual object recognition and cross-modal image retrieval, and obtain superior results in both applications compared to related methods. Guanqun Cao, Alexandros Iosifidis, Ke Chen 0004, Moncef Gabbouj |
IEEE Trans. Cybern. | 3 |
| 2018 | Hierarchical Sliding Slice Regression for Vehicle Viewing Angle EstimationabstractWe propose a novel hierarchical sliding slice regression which in a coarse-to-fine manner represents global circular target space with a number of ordinally localized and overlapping subspaces. Our method is particularly suitable for visual regression problems where the regression target is circular (e.g., car viewing angle) and visual similarity inconsistent over the target space (e.g., repetitive appearance). A good application example is the camera-based car viewing angle estimation problem, where visual similarity of different views is highly inconsistent-front and back views and left and right side views are pair-wise similar, but appear at the far ends of the circular view angle space. In practice, the problem is even more complicated due to large visual variation of objects (e.g., different car models). We perform extensive experiments on the Lausanne Federal of Institute of Technology Multi-view Car and KITTI Data Sets as well as the Technische Universitat Darmstadt Multi-view Pedestrians Data Set and achieve superior performance as compared to the state-of-the-art algorithms. Dan Yang 0012, Yanlin Qian, Ke Chen 0004, Eleni Berki, Joni-Kristian Kämäräinen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Recurrent Color ConstancyabstractWe introduce a novel formulation of temporal color constancy which considers multiple frames preceding the frame for which illumination is estimated. We propose an end-to-end trainable recurrent color constancy network – the RCC-Net – which exploits convolutional LSTMs and a simulated sequence to learn compositional representations in space and time. We use a standard single frame color constancy benchmark, the SFU Gray Ball Dataset, which can be adapted to a temporal setting. Extensive experiments show that the proposed method consistently outperforms single-frame state-of-the-art methods and their temporal variants. Yanlin Qian, Ke Chen 0004, Jarno Nikkanen, Joni-Kristian Kämäräinen, Jiri Matas |
ICCV | 2 |
| 2017 | Cross-Granularity Graph Inference for Semantic Video Object SegmentationabstractWe address semantic video object segmentation via a novel cross-granularity hierarchical graphical model to integrate tracklet and object proposal reasoning with superpixel labeling. Tracklet characterizes varying spatial-temporal relations of video object which, however, quite often suffers from sporadic local outliers. In order to acquire high-quality tracklets, we propose a transductive inference model which is capable of calibrating short-range noisy object tracklets with respect to long-range dependencies and high-level context cues. In the center of this work lies a new paradigm of semantic video object segmentation beyond modeling appearance and motion of objects locally, where the semantic label is inferred by jointly exploiting multi-scale contextual information and spatial-temporal relations of video object. We evaluate our method on two popular semantic video object segmentation benchmarks and demonstrate that it advances the state-of-the-art by achieving superior accuracy performance than other leading methods. Tinghuai Wang, Ke Chen 0004, Joni-Kristian Kämäräinen |
IJCAI | 3 |
| 2017 | Spectral attribute learning for visual regression
Ke Chen 0004, Kui Jia, Zhaoxiang Zhang 0001, Joni-Kristian Kämäräinen |
Pattern Recognit. | 1 |
| 2017 | Learning to Classify Fine-Grained Categories with Privileged Visual-Semantic MisalignmentabstractImage categorisation is an active yet challenging research topic in computer vision, which is to classify the images according to their semantic content. Recently, fine-grained object categorisation has attracted wide attention and remains difficult due to feature inconsistency caused by smaller inter-class and larger intra-class variation as well as large varying poses. Most of the existing frameworks focused on exploiting a more discriminative imagery representation or developing a more robust classification framework to mitigate the suffering. The concern has recently been paid to discovering the dependency across fine-grained class labels based on Convolutional Neural Networks. Encouraged by the success of semantic label embedding to discover the fine-grained class labels' correlation, this paper exploits the misalignment between visual feature space and semantic label embedding space and incorporates it as a privileged information into a cost-sensitive learning framework. Owing to capturing both the variation of imagery feature representation and also the label correlation in the semantic label embedding space, such a visual-semantic misalignment can be employed to reflect the importance of instances, which is more informative that conventional cost-sensitivities. Experiment results demonstrate the effectiveness of the proposed framework on public fine-grained benchmarks with achieving superior performance to state-of-the-arts. Ke Chen 0004, Zhaoxiang Zhang 0001 |
IEEE Trans. Big Data | 1 |
| 2017 | Pedestrian Counting With Back-Propagated Information and Target Drift RemedyabstractPedestrian density is one of the important factors in designing visual surveillance and intelligent transportation systems, but it is challenging to obtain accurate and robust estimates because of both inconsistent crowd patterns in the scenes and target drift caused by imbalanced data distribution. Most of existing global regression frameworks focus on the former challenge to improve the robustness of regression learning, but very few work concerns on mitigating the suffering from the latter one. This paper proposes a novel counting-by-regression framework to utilize the importance of training samples to improve the robustness against inconsistent feature-target relationship based on a recently-proposed learning paradigm-learning with privileged information. To this end, the concept of back-propagation is for the first time considered to select more informative samples contributed to robust fitting performance. Moreover, the direction of target drift along the continuously-changing target dimension is discovered by learning local classifiers under different situation of pedestrian density, which can thus be exploited in our algorithm to further boost the performance. Experimental evaluation on the public UCSD and shopping Mall benchmarks verifies that our approach significantly beats the state-of-the-art counting-by-regression frameworks. Ke Chen 0004, Zhaoxiang Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | Learning with Ambiguous Label Distribution for Apparent Age Estimation
Ke Chen 0004, Joni-Kristian Kämäräinen |
ACCV (3) | 1 |
| 2016 | Deep structured-output regression learning for computational color constancyabstractThe color constancy problem is addressed by structured-output regression on the values of the fully-connected layers of a convolutional neural network. The AlexNet and the VGG are considered and VGG slightly outperformed AlexNet. Best results were obtained with the first fully-connected “fc6” layer and with multi-output support vector regression. Experiments on the SFU Color Checker and Indoor Dataset benchmarks demonstrate that our method achieves competitive performance, outperforming the state of the art on the SFU indoor benchmark. Yanlin Qian, Ke Chen 0004, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas |
ICPR | 2 |
| 2016 | Car type recognition with Deep Neural NetworksabstractIn this paper we study automatic recognition of cars of four types: Bus, Truck, Van and Small car. For this problem we consider two data driven frameworks: a deep neural network and a support vector machine using SIFT features. The accuracy of the methods is validated with a database of over 6500 images, and the resulting prediction accuracy is over 97 %. This clearly exceeds the accuracies of earlier studies that use manually engineered feature extraction pipelines. Heikki Huttunen, Fatemeh Shokrollahi Yancheshmeh, Ke Chen 0004 |
Intelligent Vehicles Symposium | 3 |
| 2016 | Facial Age Estimation Using Robust Label DistributionabstractFacial age estimation, to predict the persons' exact ages given facial images, usually encounters the data sparsity problem due to the difficulties in data annotation. To mitigate the suffering from sparse data, a recent label distribution learning (LDL) algorithm attempts to embed label correlation into a classification based framework. However, the conventional label distribution learning framework only considers correlations across the neighbouring variables (ages), which omits the intrinsic complexity of age classes during different ageing periods (age groups). In the light of this, we introduce a novel concept of robust label distribution for scalar-valued labels, which is designed to encode the age scalars into label distribution matrices, i.e. two-dimensional Gaussian distributions along age classes and age groups respectively. Overcoming the limitations of conventional hard group boundaries in age grouping and capturing intrinsic inter-group dependency, our framework achieves robust and competitive performance over the conventional algorithms on two popular benchmarks for human age estimation. Ke Chen 0004, Joni-Kristian Kämäräinen, Zhaoxiang Zhang 0001 |
ACM Multimedia | 1 |
| 2016 | A parameter-free label propagation algorithm for person identification in stereo videos
Chongsheng Zhang, Jingjun Bi, Changchang Liu, Ke Chen 0004 |
Neurocomputing | 4 |
| 2016 | Improved neural dynamics for online Sylvester equations solving
Ke Chen 0004 |
Inf. Process. Lett. | 1 |
| 2016 | Pedestrian Density Analysis in Public Scenes With Spatiotemporal Tensor FeaturesabstractPedestrian density estimation is one of the key problems in intelligent transportation systems and has been widely applied to a number of applications in other fields of engineering. Counting-by-regression methods are more favorable for coping with such a problem owing to their robustness against interperson occlusion and relaxing the impractical requirement of a high video frame rate, compared to counting-by-detection and counting-by-clustering methods. However, imagery features in the existing counting-by-regression approaches are extracted from the whole region or spatially localized cells/pixels of each single video frame, which omits the unique motion patterns of the same pedestrians across the neighboring frames. In the light of this, this paper exploits a novel tensor-formed spatiotemporal feature representation and applies it in a multilinear regression learning framework, which can capture spatially distributed dynamic crowd patterns by discovering the latent multidimensional structural correlations of tensor features along both spatial (i.e., horizontal and vertical) and temporal dimensions. Extensive evaluation with the public UCSD and Shopping Mall benchmarks demonstrate superior performance of our approach to the state-of-the-art counting methods even when the surveillance data has a low frame rate. Ke Chen 0004, Joni-Kristian Kämäräinen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Unsupervised visual alignment with similarity graphsabstractAlignment of semantically meaningful visual patterns, such as object classes, is an important pre-processing step for a number of applications such as object detection and image categorization. Considering the expensive manpower spent on the annotation for supervised alignment methods, unsupervised alignment techniques are more favorable especially for large-scale problems. Fine adjustment can be effectively and efficiently achieved with image congealing methods, but they require moderately good initialization which is largely invalid in practice. Alignment of visual class examples with large view point changes remains as an open problem. Feature-based methods can solve the problem to some degree, but require manual selection of a good seed image and omit the fact that examples of a semantic class can be visually very different (e.g., Harley-Davidsons and Scooters in “motorbikes”). In this work, we overcome the aforementioned drawbacks by defining visual similarity under the generalized assignment problem which is solved by fast approximation and non-linear optimization. From pair-wise image similarities we construct an image graph which is used to step-wise align, “morph”, an image to another by graph traveling. We automatically find a suitable seed by novel centrality measure which identifies “similarity hubs” in the graph. The proposed approach in the unsupervised manner outperforms the state-of-the-art methods with classes from the popular benchmark datasets. Fatemeh Shokrollahi Yancheshmeh, Ke Chen 0004, Joni-Kristian Kämäräinen |
CVPR | 2 |
| 2014 | Learning to Count with Back-propagated InformationabstractError back-propagation is one of the principled learning strategies widely used in pattern recognition and machine learning, e.g. neural networks. The existing frameworks employed back-propagated error as a performance criteria (or termed, object function) aiming for supervising model-learning. Inspired by the recent success achieved by learning with the privileged information (LPI), we propose a novel regression-based framework by extending the concept of back-propagation in supervised learning methods to high-level guiding the model learning, so the proposed model is able to mine the importance of samples contributed to the fitting performance, which is missed in the existing regression techniques. To verify the effectiveness of the proposed learning paradigm, both low-level imagery features and intermediary semantic attributes are adopted in this paper. Extensive evaluations on pedestrian counting with public UCSD and Mall benchmarks demonstrate that the effectiveness of the proposed framework. Ke Chen 0004, Joni-Kristian Kämäräinen |
ICPR | 1 |
| 2014 | Learning Generative Models of Object Parts from a Few Positive ExamplesabstractA number of computer vision problems such as object detection, pose estimation, and face recognition utilise local parts to represent objects, which include the distinguished information of objects. In this work, we introduce a novel probabilistic framework which automatically learns class-specific object parts (landmarks) in generative-learning manner. Encouraged by the success in learning and detecting facial landmarks, we employ bio-inspired multi-resolution Gabor features in the proposed framework. Specifically, complex-valued Gabor filter responses are first transformed to landmark specific likelihoods using Gaussian Mixture Models (GMM), and then efficient response matrix shift operations provide detection over orientations and scales. We avoid the undesirable characteristic of generative learning, a large number of training instances, with the novel concept of randomised Gaussian mixture model. Extensive experiments with public benchmarking Caltech-101 and BioID datasets demonstrate the effectiveness of our proposed method for localising object landmarks. Ekaterina Riabchenko, Joni-Kristian Kämäräinen, Ke Chen 0004 |
ICPR | 3 |
| 2014 | Density-Aware Part-Based Object Detection with Positive ExamplesabstractPart-based models have become the mainstream approach for visual object classification and detection. The key tools adopted by the most methods are interest point detectors and descriptors, shared codes for object parts (visual codebook) and discriminative learning using positive and negative class examples. Distinction of our method from the existing part-based methods for object detection is the use of sparse class-specific landmarks with semantic meaning. The landmarks are the additional distinguished information of object location in the proposed framework. Additionally, localising semantic and discriminative landmarks (object parts) is significant in other related applications of computer vision, such as facial expression recognition and pose/orientation estimation of objects. Therefore, we propose a model which deviates from the mainstream by the fact that the object parts' appearance and spatial variation, constellation, are explicitly modelled in a generative probabilistic manner. With using only positive examples our method can achieve object detection accuracy comparable to state-of-the-art discriminative method. Ekaterina Riabchenko, Joni-Kristian Kämäräinen, Ke Chen 0004 |
ICPR | 3 |
| 2013 | Cumulative Attribute Space for Age and Crowd Density EstimationabstractA number of computer vision problems such as human age estimation, crowd density estimation and body/face pose (view angle) estimation can be formulated as a regression problem by learning a mapping function between a high dimensional vector-formed feature input and a scalar-valued output. Such a learning problem is made difficult due to sparse and imbalanced training data and large feature variations caused by both uncertain viewing conditions and intrinsic ambiguities between observable visual features and the scalar values to be estimated. Encouraged by the recent success in using attributes for solving classification problems with sparse training data, this paper introduces a novel cumulative attribute concept for learning a regression model when only sparse and imbalanced data are available. More precisely, low-level visual features extracted from sparse and imbalanced image samples are mapped onto a cumulative attribute space where each dimension has clearly defined semantic interpretation (a label) that captures how the scalar output value (e.g. age, people count) changes continuously and cumulatively. Extensive experiments show that our cumulative attribute framework gains notable advantage on accuracy for both age estimation and crowd counting when compared against conventional regression models, especially when the labelled training data is sparse with imbalanced sampling. Ke Chen 0004, Shaogang Gong, Tao Xiang 0002, Chen Change Loy |
CVPR | 1 |
| 2012 | Feature Mining for Localised Crowd CountingabstractThis paper presents a multi-output regression model for crowd counting in public scenes. Existing counting by regression methods either learn a single model for global counting, or train a large number of separate regressors for localised density estimation. In contrast, our single regression model based approach is able to estimate people count in spatially localised regions and is more scalable without the need for training a large number of regressors proportional to the number of local regions. In particular, the proposed model automatically learns the functional mapping between interdependent low-level features and multi-dimensional structured outputs. The model is able to discover the inherent importance of different features for people counting at different spatial locations. Extensive evaluations on an existing crowd analysis benchmark dataset and a new more challenging dataset demonstrate the effectiveness of our approach. 1 Ke Chen 0004, Chen Change Loy, Shaogang Gong, Tony Xiang |
BMVC | 1 |
| 2009 | MATLAB Simulink modeling and simulation of LVI-based primal-dual neural network for solving linear and quadratic programs
Yunong Zhang, Weimu Ma, Xiaodong Li 0011, Hongzhou Tan, Ke Chen 0004 |
Neurocomputing | 5 |
| 2008 | A weights-directly-determined simple neural network for nonlinear system identificationabstractBased on polynomial interpolation and approximation theory, a special feed-forward neural network using power activation functions is constructed in this paper. The neural model employs a three-layer structure with the hidden-layer neurons activated by a group of order-increasing power functions (while other layers’ neurons use linear activation functions). In addition, the weights-updating formula for such a neural network could be derived from the standard BP training method. A pseudoinverse-based method (or termed, weights-direct-determination / one-step-weights-determination method) is then established to determine immediately the neural-network weights without lengthy iterative BP-training. It is shown that such a power-activated feed-forward neural network could perform effectively and efficiently for nonlinear system identification. Computer-simulation results further substantiate the benefits of its weights-direct-determination method. Yunong Zhang, Chenfu Yi, Ke Chen 0004 |
FUZZ-IEEE | 4 |
| 2008 | MATLAB Simulation and Comparison of Zhang Neural Network and Gradient Neural Network for Online Solution of Linear Time-Varying Matrix Equation AXB-C=0
Ke Chen 0004, Shuai Yue, Yunong Zhang |
ICIC (2) | 1 |
| 2008 | Zhang Neural Network Versus Gradient Neural Network for Online Time-Varying Quadratic Function Minimization
Yunong Zhang, Chenfu Yi, Ke Chen 0004 |
ICIC (2) | 4 |
| 2008 | Zhang neural network without using time-derivative information for constant and time-varying matrix inversionabstractTo obtain the inverses of time-varying matrices in real time, a special kind of recurrent neural networks has recently been proposed by Zhang et al. It is proved that such a Zhang neural network (ZNN) could globally exponentially converge to the exact inverse of a given time-varying matrix. To find out the effect of time-derivative term on global convergence as well as for easier hardware-implementation purposes, the ZNN model without exploiting time-derivative information is investigated in this paper for inverting online matrices. Theoretical results of both constant matrix inversion case and time-varying matrix inversion case are presented for comparative and illustrative purposes. In order to substantiate the presented theoretical results, computer-simulation results are shown, which demonstrate the importance of time derivative term of given matrices on the exact convergence of ZNN model to time-varying matrix inverses. Yunong Zhang, Zenghai Chen, Ke Chen 0004, Binghuang Cai |
IJCNN | 3 |
| 2008 | A simplified LVI-based primal-dual neural network for repetitive motion planning of PA10 robot manipulator starting from different initial statesabstractThis paper presents a simplified primal-dual neural network based on linear variational inequalities (LVI) for online repetitive motion planning of PA10 robot manipulator. To do this, a drift-free criterion is exploited in the form of a quadratic function. In addition, the repetitive-motion-planning scheme could incorporate the joint limits and joint velocity limits simultaneously. Such a scheme is finally reformulated as a time-varying quadratic program (QP). As a QP real-time solver, the simplified LVI-based primal-dual neural network (LVI-PDNN) is designed based on the QP-LVI conversion and Karush-Kuhn-Tucker (KKT) conditions. It has a simple piecewise-linear dynamics and could globally exponentially converge to the optimal solution of strictly-convex quadratic-programs. The simplified LVI-PDNN model is simulated based on PA10 robot arm, and simulation results show the effective remedy of the joint angle drift problem of PA10 robot. Yunong Zhang, Zhiguo Tan, Zhi Yang 0004, Xuanjiao Lv, Ke Chen 0004 |
IJCNN | 5 |
| 2008 | MATLAB Simulation and Comparison of Zhang Neural Network and Gradient Neural Network for Time-Varying Lyapunov Equation Solving
Yunong Zhang, Shuai Yue, Ke Chen 0004, Chenfu Yi |
ISNN (1) | 3 |
| 2007 | MATLAB Simulation of Gradient-Based Neural Network for Online Matrix Inversion
Yunong Zhang, Ke Chen 0004, Weimu Ma, Xiaodong Li 0011 |
ICIC (2) | 2 |