EDBT 2026 Demo / reviewers in the wild / expert
Ke Lu 0002
dblp:33/1254-2 · also Ke Lv 0002
· DBLP profile ↗
197ranked-venue papers
6as first author
88since 2021 · last 2026
0000-0003-0176-3088ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 121 · 3 first-author · 58 since 2021Artificial intelligence and machine learning · 65 · 2 first-author · 29 since 2021Databases, data management, data science and information retrieval · 18 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 12 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Heterophily-aware Contrastive Learning for Heterophilic HypergraphsabstractHypergraph neural networks (HNNs) have emerged as powerful tools for modeling high-order relationships in complex systems. However, most existing HNNs are designed under the assumption of homophily, which does not hold in many real-world scenarios where connected nodes often exhibit diverse semantics, i.e., heterophily. This inconsistency leads to suboptimal aggregation and degraded performance, especially in low-label regimes. While a few recent methods have attempted to enhance heterophilic hypergraph learning, they often rely heavily on label supervision and overlook the potential of self-supervised techniques. In this paper, we propose HeroCL, a heterophily-aware contrastive learning framework that improves hypergraph representation under both structural heterogeneity and label scarcity. Specifically, HeroCL integrates a multi-hop neighbor encoding module to capture informative higher-order context and incorporates two complementary contrastive objectives, label-aware and structure-aware, to guide representation learning from both semantic and relational perspectives. A multi-granularity contrastive strategy is introduced to exploit latent signals across multiple neighborhood levels. Extensive experiments on several benchmark datasets against 11 existing baselines demonstrate that HeroCL achieves consistent and significant performance gains, particularly under strong heterophily and limited supervision, validating its robustness and effectiveness. Ming Li 0065, Yongqi Li 0015, Feilong Cao, Ke Lu 0002 |
AAAI | 5 |
| 2026 | Self-Supervised Hypergraph Learning with Substructure Awareness for Hyperedge PredictionabstractHyperedge prediction plays a central role in hypergraph learning, enabling the inference of high-order relations among multiple entities. However, existing methods often rely on a simplistic flat set assumption, treating candidate hyperedges as unstructured collections of nodes and neglecting their potential internal compositionality. Furthermore, the severe scarcity of observed hyperedges poses a challenge for effective supervision. In this work, we propose S3Hyper, a Substructure-contextualized Self-Supervised framework for Hyperedge prediction, which jointly addresses these two challenges. Specifically, we design a substructure-contextualized hyperedge aggregator that models the internal hierarchy of candidate hyperedges by leveraging sub-hyperedge information. In parallel, we introduce an adaptive tri-directional contrastive learning module that incorporates node-level, hyperedge-level, and cross-level alignment objectives, supported by temperature-adaptive mechanisms. Experimental results on four public datasets demonstrate that S3Hyper consistently outperforms strong baselines, with ablation studies verifying the effectiveness of each component. Ming Li 0065, Huiting Wang, Lu Bai 0001, Lixin Cui, Feilong Cao, Ke Lu 0002 |
AAAI | 7 |
| 2026 | Multi-Granular Graph Learning with Fine-Grained Behavioral Pattern Awareness for Session-Based RecommendationabstractSession-based recommendation aims to predict users’ next actions by modeling their ongoing interaction sequences, particularly in scenarios where long-term user profiles are unavailable. While existing methods have achieved promising results by leveraging sequential and graph-based structures, they often rely on global aggregation strategies that emphasize dominant user interests while overlooking the transient and fine-grained behavior patterns embedded in sessions. In practice, user intent evolves across sessions and is reflected through diverse behavioral patterns, ranging from immediate preferences to segmented co-occurrence interests and long-range goals. To address these limitations, we propose GraphFine, a novel multi-granular graph learning framework that achieves fine-grained behavioral pattern awareness for session-based recommendation. Our approach models user behavior at different temporal and semantic granularities through a combination of graph and hypergraph neural networks. Specifically, we employ a position-aware graph to capture short-term item transitions, and construct segmented co-occurrence hypergraphs to uncover high-order semantic relations among co-occurred items. To preserve diverse user intents, we further introduce a multi-view intent readout mechanism that extracts and adaptively integrates intent signals from short-term actions, segmented co-occurrence patterns, and entire sessions. Extensive experiments on benchmark datasets demonstrate that GraphFine consistently outperforms existing state-of-the-art methods, confirming its effectiveness in capturing fine-grained and dynamic user preferences for more accurate recommendation. Ming Li 0065, Zihao Yan, Lixin Cui, Lu Bai 0001, Feilong Cao, Ke Lu 0002, Zhao Li 0007 |
AAAI | 7 |
| 2026 | HyperNoRA: Hyperedge Prediction via Node-Level Relation-Aware Self-Supervised Hypergraph LearningabstractHyperedge prediction plays a critical role in high-order relational modeling with hypergraphs, yet most existing methods primarily focus on sampling strategies or local aggregation within candidate hyperedges. These approaches often overlook global structural dependencies that are essential for learning expressive node and hyperedge representations. In this paper, we propose HyperNoRA, a novel self-supervised hypergraph learning framework that integrates global node-level relation awareness with contrastive learning. Specifically, we construct a global node relation graph that captures both direct and indirect structural correlations, which guides a structure-aware aggregator to enhance node representations with informative global context. To prevent over-smoothing and maintain discriminability, a contrastive learning module is introduced to align representations across graph augmentations while separating semantically dissimilar nodes. Extensive experiments on several benchmark datasets demonstrate that HyperNoRA consistently outperforms state-of-the-art baselines, and ablation studies verify the effectiveness of its key components. Ming Li 0065, Zhanle Zhu, Lu Bai 0001, Lixin Cui, Feilong Cao, Ke Lu 0002 |
AAAI | 7 |
| 2026 | SCG-SSC: Semantic Scene Completion via Self-and-Cross Gated Fusion of Depth Maps and Semantic Priors
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
ICPR (9) | 5 |
| 2026 | SAP-DQR: Joining Spatial-Adaptive Pyramid and Adaptive Query Reorganization for Speed-Accuracy Instance Segmentation
Jiahao Zou, Congxuan Zhang, Liyue Ge, Jiawen Yang, Zhen Chen 0004, Ke Lu 0002 |
MMM (1) | 7 |
| 2026 | Adaptive semi-supervised 3D object detection network with multi-level constraints
Yang Zhao 0028, Ke Lu 0002 |
Neurocomputing | 5 |
| 2026 | ADI-SAM: Adapting segment anything model for degraded images
Yang Zhao 0028, Zhaoxiang Liu, Yibing Nan, Ke Lu 0002, Shiguo Lian |
Neurocomputing | 5 |
| 2026 | Decoupled Hierarchical Distillation for Multimodal Emotion RecognitionabstractHuman multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1.3%/2.4% (ACC$_{7}$7), 1.3%/1.9% (ACC$_{2}$2) and 1.9%/1.8% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces. Yong Li 0032, Yuanzhi Wang, Yi Ding 0012, Shiqing Zhang, Ke Lu 0002, Cuntai Guan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Mambafusion: State-space model-driven object-scene fusion for multi-modal 3D object detection
Tong Ning, Ke Lu 0002, Jian Xue 0002 |
Pattern Recognit. | 2 |
| 2026 | SGTP-Net: Semantic Guidance and Texture Priors-Based Dual-Branch Segmentation Network for Surface Defect DetectionabstractDeep learning-based surface defect segmentation approaches have shown promising performance in recent years. However, segmenting defects with complex shapes, large variations in size, and weakly textured defects with indistinct characteristics still poses significant challenges. In this article, a novel semantic guidance and texture priors based dual-branch surface defect segmentation network (SGTP-Net) is proposed for those issues. Firstly, we construct a feature extraction network combines semantic and texture branches. The semantic branch establishes global contextual relationships, while the texture branch captures local features of defects, this dual-branch ensured the network to extract features from various complex defects. Secondly, we design a feature fusion strategy based on semantic guidance and texture priors. The semantic information is used to guides the output of texture branch. After that, the guided texture information provides valuable edge texture priors for each layers output in semantic branch. The two branches mutually guide each other for improving ability of weak textures feature extraction. Finally, we run our method on the NEU-Seg, MT-Defect and MSD datasets to conduct a comprehensive comparison with some state-of-the-art general object segmentation models and specialized surface defect segmentation methods. The experimental results show that our SGTP-Net performs well in surface defect detection, offering excellent semantic segmentation accuracy and exhibiting good stability and robustness in detecting various surface defects. Leqi Jiang, Liyue Ge, Chengzhong Wu, Yaonan Wang 0001, Ke Lu 0002, Congxuan Zhang |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | CLSS: A Plug-and-Play Closed-Loop Semantic Supervisor for Actor-Critic Reinforcement Learning
Jian Xue 0002, Ke Lu 0002 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Iter3DDet: Depth-Guided Iterative Fusion and Refinement for Monocular 3D Object DetectionabstractMonocular 3D object detection offers significant potential for autonomous systems due to its inherent cost-effectiveness and scalability. While DETR-based architectures excel in 2D vision tasks, critical limitations persist in extending them effectively to monocular 3D detection, as evidenced in existing frameworks like MonoDETR and MonoDGP. These methods typically suffer from inefficient serial fusion of multimodal features and lack iterative refinement mechanisms, limiting their performance, especially for mid-to-long range targets. To overcome these shortcomings, we propose Iter3DDet, a novel depth-guided iterative refinement framework that integrates fine-grained feature fusion to significantly enhance detection performance. The core novelty of our approach lies in two key innovations: (1) A hybrid feature encoder combining MonoDGP’s region segmentation head with MonoDETR’s visual backbone, augmented by a multi-scale context attention module that dynamically aggregates structural and semantic cues across pyramid levels, eliminating heuristic fusion rules; (2) A depth-guided adaptive cross-modal decoder that iteratively fuses depth and context features through prioritized attention mechanisms, coupled with a novel iterative refinement training strategy that progressively refines 3D detection hypotheses, substantially improving accuracy across targets of varying difficulty levels. Extensive experiments on the KITTI, nuScenes, and Waymo benchmarks demonstrate Iter3DDet’s state-of-the-art performance, validating the effectiveness of our iterative refinement paradigm. The code will be open-sourced at https://github.com/PCwenyue. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | DreamAssemble: Complex Multi-Object Text-to-3D Generation via Multi-Density Neural Fields
Bin Huang 0016, Jinbao Wang 0001, Dongmei Jiang, Hongjuan Pei, Qiulu Li, Jian Xue 0002, Ke Lu 0002 |
IEEE Trans. Image Process. | 7 |
| 2026 | DRDFNet: A Degradation-Aware Restoration and Detail-Preserving Fusion Network for Infrared and Visible ImageabstractMulti-source image fusion combines infrared and visible information to improve scene perception in applications such as drone reconnaissance and autonomous driving. However, most existing infrared-visible image fusion methods are developed under ideal imaging assumptions. In adverse environments, visible images often lose structural and textural details, whereas infrared images are affected by noise, stripe artifacts, and low contrast, leading to degraded fusion quality and weakened downstream perception performance. To address these limitations, we propose a unified Degradation-aware Restoration and Detail-preserving Fusion Network (DRDFNet), which consists of a Degradation-Aware Restoration Transformer and a Detail-Preserving Fusion Mamba. The restoration branch uses a Compound Degradation Restoration Module (CDRM) to remove complex degradations, while the fusion branch employs a Dynamic Feature Fusion Module (DFFM) to integrate local complementary cues and global correlations across modalities. A two-stage training strategy is further introduced to reduce the optimization conflict between restoration and fusion. In addition, we construct DIVIF, a large-scale degraded IVIF benchmark generated by a physics-based imaging simulator. Experiments on the DIVIF and AWMM-100k benchmarks demonstrate that DRDFNet achieves robust and competitive performance compared with SOTA methods. Both the dataset and source code will be made publicly available at https://github.com/Liupeng97/DRDFNet. Peng Liu 0024, An Wei, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002 |
IEEE Trans. Image Process. | 8 |
| 2026 | Improving Unsupervised Ultrasonic Image Anomaly Detection via Frequency-Spatial Feature Filtering and Gaussian Mixture ModelingabstractUltrasonic image anomaly detection faces significant challenges due to limited labeled data, strong structural and random noise, and highly diverse defect manifestations. To overcome these obstacles, we introduce UltraChip, a new large-scale C-scan benchmark containing about 8,000 real-world images from various chip packaging types, each meticulously annotated with pixel-level masks for cracks, holes, and layers. Building on this resource, we present FSGM-Net, a fully unsupervised framework tailored for anomaly detection. FSGM-Net leverages an adaptive Frequency-Spatial feature filtering mechanism: a learnable FFT-Spatial patch filter first suppresses noise and dynamically assigns normality weights to Vision Transformer (ViT) patch features. Subsequently, an Adaptive Gaussian Mixture Model (Ada-GMM) captures the distribution of normal features and guides a deep-shallow multi-scale interaction decoder for accurate, pixel-level anomaly inference. In addition, we propose a filter loss that enforces encoder-filter consistency and entropy-based sparse gating, together with a distributional loss that encourages both feature reconstruction and confident Gaussian mixture modeling. Extensive experiments demonstrate that FSGM-Net not only achieves state-of-the-art results on UltraChip but also exhibits superior cross-domain generalization to MVTec-AD and VisA, while supporting real-time inference on a single GPU. Together, the dataset and framework advance robust, annotation-free ultrasonic NDT in practical applications. The UltraChip dataset can be obtained via https://iiplab.net/ultrachip/. Ke Lu 0002, Jinbao Wang 0001, Can Gao, Jian Xue 0002 |
IEEE Trans. Image Process. | 2 |
| 2026 | MotionFlow: Efficient Motion Generation With Latent Flow MatchingabstractIn the field of human centric multimedia, text-driven human motion generation is a significant pursuit with wide-ranging applications across diverse scenarios. Despite substantial advancements, existing methods often suffer from a trade-off between inference latency and high-quality generation. To overcome this gap, we propose the Motion Latent Flow Matching model (MotionFlow), a novel and powerful framework for motion generation. It introduces flow matching algorithm in the latent space, which can achieve superior performance with just one-step inference. In addition to the text-driven task, we further extend our method to controllable motion generation. Specifically, we integrate a control encoder into the latent space and further decode the predicted latent code into motion space to support explicit supervision, ensuring the synthesized motion can tightly align with the input signals. Extensive experiments demonstrate that our MotionFlow not only outperforms current leading approaches for the text-driven task, but also delivers remarkable capabilities in controllable motion generation. Kun Dong 0001, Jian Xue 0002, Xing Lan, Qingyuan Liu 0001, Ke Lu 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | Learning Hierarchical Continuous Dynamics for Facial Action Unit Intensity EstimationabstractDynamic facial action recognition is key to understanding human emotions and behaviors, yet estimating facial action units (AUs) intensities in videos is difficult due to subtle muscle motions and complex spatial-temporal dependencies. Existing methods often use fixed or coarse graphs, limiting the ability to capture intricate AU relations and long-range dynamics. This paper presents a hierarchical framework CDAU to effectively capture Continuous Dynamics for AU intensity estimation task. Our approach dynamically constructs multiscale graphs for fine-grained spatiotemporal AU interactions and adaptively fusing information across levels. A bidirectional state-space module further captures long-range temporal dependencies. Extensive experiments on FEAFA and DISFA show that CD-AU outperforms existing methods in both ICC and MAE metrics, validating its generalization capability and stability across subjects and expressions. Ke Lu 0002, Yan Li 0121, Menghao Hu, Guohong Hu, Dongmei Jiang, Jian Xue 0002 |
BIBM | 2 |
| 2025 | MotionFlow: Joint Motion Priors and Appearance Enhancement for High-Accuracy Optical Flow EstimationabstractAlthough optical flow estimation has improved significantly in recent years, large displacements and occlusions remain challenging for current methods due to motion discontinuities that may hinder accurate feature correspondences in these regions, leading to degraded performance. To address this challenge, we propose a novel method named MotionFlow for high-accuracy optical flow estimation. In the encoding stage, we integrate multi-scale features enhance motion and context appearance information via cross- and inter-enhancement module. Subsequently, cross-frame features are utilized to establish motion priors, thereby providing essential prior knowledge for motion estimation. During the decoding stage, we align appearance features of the target frame with the reference frame through warping to retrieve missing context crucial for motion decoding. Experimental results demonstrate the efficacy of our approach, achieving state-of-the-art performance, particularly outperforming online benchmarks on Sintel Final pass and KITTI-2015 datasets. Congxuan Zhang, Zhen Chen 0004, Hongye Chen, Liyue Ge, Ke Lu 0002 |
ICASSP | 6 |
| 2025 | Text-to-Any-Skeleton Motion Generation Without Retargeting
Qingyuan Liu 0001, Ke Lu 0002, Kun Dong 0001, Jian Xue 0002, Zehai Niu, Jinbao Wang 0001 |
ICCV | 2 |
| 2025 | One General Plug-In for Facial Heatmap-based Keypoint DetectionabstractIn this paper, we systematically investigate the error distribution in predicted heatmaps for face alignment, and point out that previous works are unreliable in following the rule that decodes coordinates by locating the maximum-score pixel. Our research reveals that the majority of ground-truth positions do not match that pixel but rather lie within a range of a few pixels. Building on this phenomenon, we transform the model’s objective from predicting inaccurate landmarks to identifying precise proposals with that range. We propose a simple but effective module, termed the Response Aware Module (RAM), leveraging response scores in the proposal to regress the proposal offset, which can be used as a plug-and-play layer integrated into public models. Furthermore, we present a novel Heatmap RCNN framework to exploit the distribution of multi-scale heatmaps. Extensive experiments have demonstrated that the trained RAM can be integrated seamlessly as a ready-to-use plugin with the model, yielding impressive improvements. Meanwhile, Heatmap RCNN performs far superior to SOTA results, with 3.82 NME on WFLW, 3.09 on COFW, and 2.90 on 300W. Hanyu Jiang 0004, Jian Xue 0002, Xing Lan, Ke Lu 0002 |
ICME | 4 |
| 2025 | EduLLM: Leveraging Large Language Models and Framelet-Based Signed Hypergraph Neural Networks for Student Performance PredictionabstractThe growing demand for personalized learning underscores the importance of accurately predicting students’ future performance to support tailored education and optimize instructional strategies. Traditional approaches predominantly focus on temporal modeling using historical response records and learning trajectories. While effective, these methods often fall short in capturing the intricate interactions between students and learning content, as well as the subtle semantics of these interactions. To address these gaps, we present EduLLM, the first framework to leverage large language models in combination with hypergraph learning for student performance prediction. The framework incorporates FraS-HNN ($\underline{\mbox{Fra}}$melet-based $\underline{\mbox{S}}$igned $\underline{\mbox{H}}$ypergraph $\underline{\mbox{N}}$eural $\underline{\mbox{N}}$etworks), a novel spectral-based model for signed hypergraph learning, designed to model interactions between students and multiple-choice questions. In this setup, students and questions are represented as nodes, while response records are encoded as positive and negative signed hyperedges, effectively capturing both structural and semantic intricacies of personalized learning behaviors. FraS-HNN employs framelet-based low-pass and high-pass filters to extract multi-frequency features. EduLLM integrates fine-grained semantic features derived from LLMs, synergizing with signed hypergraph representations to enhance prediction accuracy. Extensive experiments conducted on multiple educational datasets demonstrate that EduLLM significantly outperforms state-of-the-art baselines, validating the novel integration of LLMs with FraS-HNN for signed hypergraph learning. Ming Li 0065, Yukang Cheng, Lu Bai 0001, Feilong Cao, Ke Lu 0002, Jiye Liang, Pietro Liò |
ICML | 5 |
| 2025 | MATCH: Modality-Calibrated Hypergraph Fusion Network for Conversational Emotion RecognitionabstractMultimodal emotion recognition aims to identify emotions by integrating multimodal features derived from spoken utterances. However, existing work often neglects the calibration of conversational entities, focusing mainly on extracting potential intra- or cross-modal information. This leads to the underutilization of utterance information that is essential for accurately characterizing emotion. Additionally, the lack of effective modeling of conversational patterns limits the ability to capture emotional pathways across contexts, modalities and speakers, impacting the overall emotional understanding. In this study, we propose the modality-calibrated hypergraph fusion network (MATCH), which leverages multimodal fusion and hypergraph learning techniques to address these challenges. In particular, we introduce an entity calibration strategy that refines the representations of conversational entities both at the modality and context levels, allowing for deeper insights into emotion-related cues. Furthermore, we present an emotion-aligned hypergraph fusion method that incorporates a line graph to explore conversational patterns, facilitating flexible knowledge transfer across modalities through hyperedge-level and graph-level alignments. Experiments demonstrate that MATCH outperforms state-of-the-art approaches on two benchmark datasets. Jiandong Shi, Ming Li 0065, Lu Bai 0001, Feilong Cao, Ke Lu 0002, Jiye Liang |
IJCAI | 5 |
| 2025 | DC-BEV: Depth-Completed Bird's Eye View Representation for Multi-Modal 3D Object DetectionabstractThe bird’s eye view (BEV) representation is essential for accurate 3D perception tasks (e.g., 3D object detection) in autonomous driving for its precise localization, scale consistency, and modality independence. However, traditional depth-prediction-based methods, which rely on predicted depth distributions derived from semantic image features for BEV transformation, face challenges such as ambiguous depth-prior information and dependence on predefined depth distribution types. To address these limitations, we propose a novel multimodal 3D object detection method, Depth Completed BEV (DC-BEV), which leverages ground-truth sparse LiDAR depth to guide BEV transformations, significantly enhancing depth estimation accuracy. Specifically, we introduce a Multimodal-Depth Completion (MDC) mechanism, which enriches sparse LiDAR depth into dense depth maps by integrating semantic and geometric cues from images. Additionally, to mitigate gradient instability caused by inadequate implicit supervision, we present an Explicit Depth Supervision (EDS) mechanism that directly supervises depth predictions using a dedicated depth loss. Comprehensive experiments conducted on the nuScenes dataset demonstrate that DC-BEV achieves superior performance, notably improving detection accuracy through enhanced depth estimation quality and robust BEV representation. Tong Ning, Ke Lu 0002, Jian Xue 0002 |
SMC | 2 |
| 2025 | Adaptive Multi-Layer Prioritized Fictitious Self-play in Multi-Agent Reinforcement LearningabstractIn the field of Multi-Agent Reinforcement Learning (MARL), strategy selection and optimization are key challenges in improving agent performance. The Policy Space Response Oracles (PSRO) algorithm is widely used in MARL, and the meta-solver is one of its cores. However, existing meta-solvers may exhibit limitations in complex MARL environments, such as significant consumption of computing resources, poor convergence, instability, and so on. Therefore, an adaptive multi-layer Prioritized Fictitious Self-Play (PFSP) is proposed in this paper as a meta-solver method for the PSRO algorithm, which further improves the effectiveness of strategy optimization by utilizing the game results of the meta-game in a more efficient and reasonable way. Adaptive multi-layer PFSP can flexibly make strategy choices based on payoff at different layers, thus overcoming the shortcomings of traditional meta-solvers in dealing with complex strategy spaces. The experimental results show that the proposed method significantly improves the convergence speed and performance of strategies in complex MARL environments, especially in the training and testing of the Google Research Football environment, demonstrating its potential and advantages in practical applications. Jian Xue 0002, Ke Lu 0002 |
SMC | 6 |
| 2025 | SRSA-Depth: shape and region similarity awareness for outdoor monocular depth estimation
Liyue Ge, Congxuan Zhang, Zhen Chen 0004, Ke Lu 0002 |
Multim. Syst. | 4 |
| 2025 | FrameERC: Framelet Transform Based Multimodal Graph Neural Networks for Emotion Recognition in Conversation
Ming Li 0065, Jiandong Shi, Lu Bai 0001, Changqin Huang, Yunliang Jiang, Ke Lu 0002, Shijin Wang 0001, Edwin R. Hancock |
Pattern Recognit. | 6 |
| 2025 | AVES: An Audio-Visual Emotion Stream Dataset for Temporal Emotion DetectionabstractHuman emotions vary over time, which can be vividly described as a stream of emotions. Observing the emotion stream in daily life provides valuable insights into an individual's mental state. However, existing research in emotion understanding has mainly focused on classification tasks, assigning an emotion category to a well-trimmed segment or each frame within a continuous signal. In contrast, the task of temporal emotion detection, which involveslocatingthe boundaries of emotion segments andrecognizingtheir categories in untrimmed signals, has not been fully explored. To advance research in this area, this paper introduces an in-the-wild Audio-Visual Emotion Stream (AVES) dataset, which is reliably annotated with the time boundaries and emotion category for each emotion segment in the videos. Thus, AVES can serve as a solid benchmark for temporal emotion detection tasks. Moreover, considering the flexible boundaries and varying durations of emotion segments, we propose a Boundary Combination Network (BoCoNet) for temporal emotion detection, which leverages short-term temporal context information to first predict the boundaries of emotion segments and then locate the entire emotion segments. Extensive experiments conducted on various representative unimodal and multimodal representations demonstrate that BoCoNet achieves state-of-the-art results. The AVES dataset will be released to the research community. We expect that this paper can advance the research on emotion stream and temporal emotion detection. Yan Li 0121, Ke Lu 0002, Dongmei Jiang, Ramesh Jain 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | DinoQuery: Promoting Small 3D Object Detection With Textual PromptabstractQuery-based 3D object detection has gained significant success in the application of autonomous driving due to its ability to achieve good performance while maintaining low computational cost. However, it still struggles with the reliable detection of small objects such as bicycles and pedestrians. To address this challenge, this paper introduces a novel sparse query-based approach, termed DinoQuery. This approach utilizes Grounding-DINO with textual prompts to select small-sized objects and generate 2D category-aware queries. These 2D category-aware queries combined with 2D global queries are then lifted to 3D queries by associating each sampled query with its respective 3D position, orientation, and size. The validity of these 3D queries, along with the 2D queries, is verified by the Comprehensive Contrastive Learning (CCL) mechanism. This is achieved by aligning all 2D and 3D queries with their respective 2D and 3D ground truth labels, and computing similarity to select true positive and false positive queries. Then a contrastive loss is introduced to enhance true positive queries and weaken false positive ones based on geometric and semantic similarity. The DinoQuery was tested on the nuScenes dataset and demonstrated excellent performance. Notably, the largest increase of our method is 3.2% on NDS and 3.1% on mAP. Tong Ning, Ke Lu 0002, Hongjuan Pei, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Self-Supervised Monocular Depth Estimation With Dual-Path Encoders and Offset Field InterpolationabstractAlthough self-supervised learning approaches have demonstrated tremendous potential in multi-frame depth estimation scenarios, existing methods struggle to perform well in cases involving dynamic targets and static ego-camera conditions. To address this issue, we propose a self-supervised monocular depth estimation method featuring dual-path encoders and learnable offset interpolation (LOI). First, we construct a dual-path encoding scheme that utilizes residual and transformer blocks to extract both single- and multi-frame features from the input frames. We design a contrastive learning strategy to effectively decouple single- and multi-frame features, enabling weighted fusion guided by a confidence map. Next, we explore two distinct decoding heads for simultaneously generating low-resolution predictions and offset fields. We then design an LOI module to directly upsample a low-resolution depth map to a full-resolution map. This one-step decoding framework enables accurate and efficient depth prediction. Finally, we evaluate our proposed method on the KITTI and Cityscapes benchmarks, conducting a comprehensive comparison with state-of-the-art approaches. The experimental results demonstrate that our DualDepth method achieves competitive performance in terms of both estimation accuracy and efficiency. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
IEEE Trans. Image Process. | 5 |
| 2025 | ExpLLM: Towards Chain of Thought for Facial Expression RecognitionabstractFacial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is essential for accurately recognizing them. Current approaches, such as those based on facial action units (AUs), typically provide AU names and intensities but lack insight into the interactions and relationships between AUs and the overall expression. In this paper, we propose a novel method called ExpLLM, which leverages large language models to generate an accurate chain of thought (CoT) for facial expression recognition. Specifically, we have designed the CoT mechanism from three key perspectives: key observations, overall emotional interpretation, and conclusion. The key observations describe the AU's name, intensity, and associated emotions. The overall emotional interpretation provides an analysis based on multiple AUs and their interactions, identifying the dominant emotions and their relationships. Finally, the conclusion presents the final expression label derived from the preceding analysis. Furthermore, we also introduce the Exp-CoT Engine, designed to construct this expression CoT and generate instruction-description data for training our ExpLLM. Extensive experiments on the RAF-DB and AffectNet datasets demonstrate that ExpLLM outperforms current state-of-the-art FER methods. ExpLLM also surpasses the latest GPT-4o in expression CoT generation, particularly in recognizing micro-expressions where GPT-4o frequently fails. Xing Lan, Jian Xue 0002, Ji Qi 0003, Dongmei Jiang, Ke Lu 0002, Tat-Seng Chua |
IEEE Trans. Multim. | 5 |
| 2025 | Beyond 3D: Generic IoU for 3D Object DetectionabstractObject detection from point clouds is a fundamental task for 3D scene understanding and has a wide range of applications in the field of multimedia data processing and analysis, such as autonomous driving and virtual interaction. The IoU evaluates the overlap between the two bounding boxes to ensure consistency across network optimization and testing, becoming a recognized regression loss in the field of 3D object detection. However, there is a kind of error coupling between the IoU and the angle, i.e., the IoU does not decrease as the angle error increases and vice versa. This problem leads to sub-optimal solutions for the neural network model, which severely hampers the improvement of 3D object detection accuracy. In this paper, a novel 4DIoU method is introduced for detecting 3D objects from point clouds, which provides a comprehensive rethinking of IoU computation by integrating angular information as an additional dimension. 4DIoU not only solves the problem of error coupling between IoU and angular but also facilitates neural network optimization using angle information. Furthermore, to solve the different impacts of various object shapes on IoU variations, a special 4DIoU called TV4DIoU is proposed to fuse shape information based on three orthogonal projection views, which can adaptively learn the information of objects with different shapes. In addition, to enhance the generalization of the 4DIoU method, a high-flexibility anchor encoding method and a cyclic consistent computation formula for angular errors are designed to make 4DIoU a plug-and-play module for both anchor-based and anchor-free frameworks. Extensive evaluations conducted on the nuScenes, Waymo, and KITTI datasets have confirmed the effectiveness of the proposed method. Hengsheng Lun, Ke Lu 0002, Liping Hou, Jian Xue 0002 |
IEEE Trans. Multim. | 2 |
| 2025 | DA-Net: Density-Aware 3D Object Detection Network for Point Cloudsabstract3D object detection is an important but demanding task, which has become an active research topic in the field of multimedia. Much recent research has been devoted to exploiting end-to-end trainable object detection networks with point clouds. However, most state-of-the-art methods have bottlenecks in detecting occluded objects and small objects, because the sparseness of point clouds is exacerbated on these objects. In this paper, a Density-Aware 3D object detection network (DA-Net) is proposed to improve the perception performance for detecting occluded and small objects, which contains four components: a backbone module with an inverse density scoring module (IDM) and a point-wise attention module (PAM), a 3D intersection over union Estimation Module (3DEM), a Consistent Label Assignment (CLA) method and an Adaptive-Soft-NMS method. The proposed backbone module makes the network concentrate on low-density points of occluded objects, and suppresses outliers and background points. Then, the 3DEM is introduced to evaluate the localization quality of the prediction boxes. Furthermore, the proposed CLA method can more accurately select positive and negative samples for small objects. Finally, Adaptive-Soft-NMS is proposed in our method to reduce the number of false detections during inference and thereby improve detection performance substantially. Extensive experiments demonstrated that the proposed method achieves state-of-the-art performance on two large-scale datasets, SUN RGB-D (62.1% in terms of [email protected]) and ScanNetV2 (67.1% in terms of [email protected]), and in particular, the detection accuracy of small objects and occluded objects are extremely improved. Ke Lu 0002, Jian Xue 0002, Yang Zhao 0028 |
IEEE Trans. Multim. | 2 |
| 2025 | GPDF-Net: geometric prior-guided stereo matching with disparity fusion refinement
Congxuan Zhang, Zhibo Rao, Zhen Chen 0004, Zige Wang, Ke Lu 0002 |
Vis. Comput. | 6 |
| 2024 | SRA-YOLO: Spatial Resolution Adaptive YOLO for Semi-supervised Cross-Domain Aerial Object Detection
Jian Xue 0002, Yuqiu Li, Ke Lu 0002 |
ICANN (2) | 5 |
| 2024 | 3Dlaneformer: Rethinking Learning Views for 3D Lane DetectionabstractAccurate 3D lane detection from monocular images is crucial for autonomous driving. Recent advances leverage either front-view (FV) or bird’s-eye-view (BEV) features for prediction, inevitably limiting their ability to perceive driving environments precisely and resulting in suboptimal performance. To overcome the limitations of using features from a single view, we design a novel dual-view cross-attention mechanism, which leverages features from FV and BEV simultaneously. Based on this mechanism, we propose 3DLaneFormer, a powerful framework for 3D lane detection. It outperforms the latest BEV-based or FV-based approaches through extensive experiments on challenging benchmarks and thus verifies the necessity and benefits of utilizing features in both views. Kun Dong 0001, Jian Xue 0002, Xing Lan, Ke Lu 0002 |
ICIP | 4 |
| 2024 | From 3D to 4D: Fixing the Erroneous Coupling between IoU and Angle for Optimizing 3D Object DetectionabstractThe IoU metric directly measures the overlap between two boxes, maintaining consistency in model optimization and testing stages. It has emerged as a highly regarded regression loss in the field of 3D object detection. However, the optimization of IoU often leads to an increased angular error. This erroneous coupling phenomenon renders the model susceptible to settling into sub-optimal solutions, which have not been extensively analyzed and addressed, significantly impeding further advancements in the accuracy of 3D object detection. In this paper, a novel concept "4DIoU" is introduced for 3D object detection, where the angle information is integrated as an additional dimension in the IoU calculation, and a new formula for measuring angle correlation is proposed. The 4DIoU not only resolves the erroneous coupling between IoU and angles but also capitalizes on angle information to enhance network optimization. Furthermore, a new encoding and decoding paradigm is proposed, which is more compatible with 4DIoU for object detection in point clouds. Extensive experiments on nuScenes, Waymo and KITTI datasets demonstrate the effectiveness of our method. The plug-and-play design of our approach proves to be highly versatile. Hengsheng Lun, Ke Lu 0002, Liping Hou, Jian Xue 0002 |
ICME | 2 |
| 2024 | VS3D: A Vote-Based Semi-Supervised 3D Object Detection Framework for Point CloudsabstractIn recent years, the 3D object detection method has undergone rapid evolution, heavily relying on substantial amounts of high-quality labeled data. However, the process of annotating 3D data is both time-consuming and costly. In response to this challenge, we propose a vote-based semi-supervised 3D object detection framework called VS3D. First, a data augmentation technique named Random Grid Deleting (RGD) is proposed to detect occluded objects and small objects more robustly. Then, an auxiliary branch with Voting Consistency Learning (VCL) is added to predict object centers more accurately. Additionally, a Teacher-Student Matching (TSM) module with stricter consistency constraints is designed to accelerate network convergence and improve detection performance. Our method can integrate any vote-based fully supervised network seamlessly. Extensive experiments on SUN RGB-D and ScanNet V2 datasets demonstrate that the proposed method outperforms the state-of-the-art fully supervised model when using only 70% labeled data. Ke Lu 0002, Yang Zhao 0028, Hengsheng Lun, Zehai Niu, Jian Xue 0002 |
ICME | 2 |
| 2024 | MISTA: A Large-Scale Dataset for Multi-Modal Instruction Tuning on Aerial ImagesabstractThis paper introduces MISTA, a novel dataset for visual instruction tuning on aerial imagery, designed to enhance large multi-modal model applications in remote sensing. Originating from the renowned DOTA-v2.0 aerial object detection benchmark, MISTA uniformly processes high-resolution images into 2048×2048 pixels, creating a detailed and complex dataset tailored for remote sensing analysis. To craft this dataset, we design an automated annotation pipeline, employing advanced language models such as GPT-4 and LLaVA-1.5, to generate diverse and specialized instruction-following data. The annotations include various instruction types like multi-turn conversation, detailed description, and complex reasoning, each reflecting the intricacies inherent in remote sensing tasks. The innovative approach of subdividing aerial images into individually annotated sub-patches significantly enhances the richness of the dataset and allows for a more granular analysis of visual content. As a robust foundation for multi-modal model development in remote sensing, MISTA represents a significant advancement, setting the stage for future research and further applications in the field. Ke Lu 0002, Yuqiu Li, Jian Xue 0002 |
ICME | 2 |
| 2024 | MA-Mamba: Multi-Agent Reinforcement Learning with State Space Model
Jian Xue 0002, Ke Lu 0002 |
ICONIP (3) | 5 |
| 2024 | Real-time Integration of Fine-tuned Large Language Model for Improved Decision-Making in Reinforcement LearningabstractIn this paper, we investigate a novel and efficacious methodology for training reinforcement learning agents. We also further reveal the potential of the Large Language Model (LLM) in intricate decision-making environments. Reinforcement learning, one main approach in training decision-making capabilities for intelligent agents, suffers from the challenges of sparse rewards and inefficient exploration. Considering the current phenomenal performance of pre-trained LLMs, certain studies have adopted the LLMs to shepherd the actions of intelligent agents. However, there are some limitations in the application of LLMs, such as the lack of domain knowledge of LLMs for specific fields, and LLMs generally access a slower response time but consume a higher economic cost. Embarking from these constraints, this paper takes on the formidable challenge of the Unmanned Air Vehicle (UAV) air combat simulations environment, where decision-making is notably circumscribed by temporal limitations. We initially fine-tuned the LLMs to a domain-specific expertise while concurrently constructing a knowledge base. Hence, upon gaining profound insight into the field of air combat, it evolves into an adept model capable of effectively guiding UAV decisions. Further addressing the difficulties of real-time accessing the LLM, we propose a novel approach named Reward Shaping with Large Language Model (LLM-RS) to augment the autonomous decision-making competency of UAVs within the context of air combat simulations. We compared the agents trained by conventional reinforcement learning techniques and those trained by LLM-RS. The experimental results reveal that, under the same equipment training conditions, the LLM-RS technique grounded in the fine-tuning of the LLM and the knowledge base substantially improves the performance of UAVs in air combat simulations, while concurrently diminishing the requisite training duration. Xiancai Xiang, Jian Xue 0002, Ke Lu 0002 |
IJCNN | 6 |
| 2024 | Realistic Full-Body Motion Generation from Sparse Tracking with State Space ModelabstractIn the domain of generative multimedia and interactive experiences, generating realistic and accurate full-body poses from sparse tracking is crucial for many real-world applications, while achieving sequence modeling and efficient motion generation remains challenging. Recently, state space models (SSMs) with efficient hardware-aware designs (i.e., Mamba) have shown great potential for sequence modeling, particularly in temporal contexts. However, processing motion data is still challenging for SSMs. Specifically, the sparsity of input conditions makes motion generation an ill-posed problem. Moreover, the complex structure of the human body further complicates this task. To address these issues, we present Motion Mamba Diffusion (MMD), a novel conditional diffusion model, which effectively utilizes the sequence modeling capability of SSMs and the robust generation ability of diffusion models to track full-body poses accurately. In particular, we design a bidirectional Temporal Mamba Module (TMM) to model motion sequence. Additionally, a Spatial Mamba Module (SMM) is further proposed for feature enhancement within a single frame. Extensive experiments on the large motion capture dataset (AMASS) demonstrate that our proposed approach outperforms the latest methods in terms of accuracy and smoothness, thus providing a crucial advancement for creating realistic virtual avatars in various applications. Kun Dong 0001, Jian Xue 0002, Zehai Niu, Xing Lan, Ke Lu 0002, Qingyuan Liu 0001, Xiaoyu Qin 0001 |
ACM Multimedia | 5 |
| 2024 | A Robust and Real-Time RGB-D SLAM Method with Dynamic Point Recognition and Depth Segmentation Optimization
Shuaixin Chen, Baolin Gan, Congxuan Zhang, Zhen Chen 0004, Ke Lu 0002 |
PRCV (9) | 5 |
| 2024 | Skeleton Cluster Tracking for robust multi-view multi-person 3D human pose estimation
Zehai Niu, Ke Lu 0002, Jian Xue 0002, Jinbao Wang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2024 | Dynamic spatial-temporal topology graph network for skeleton-based action recognition
Lian Chen, Ke Lu 0002, Zehai Niu, Runchen Wei, Jian Xue 0002 |
Multim. Syst. | 2 |
| 2024 | From Methods to Applications: A Review of Deep 3D Human Motion CaptureabstractMotion capture technology is crucial in various applications like animation, virtual reality and sports analysis. With the development of deep learning methods, significant progress has been experienced in this field, producing cost-effective and user-friendly solutions for various applications. This paper provides a comprehensive review of deep learning-based human motion capture techniques. Our review aims to bridge the gap between academic research and practical applications, providing valuable insights and guidance for researchers and practitioners in deep learning-based human motion capture. Our study puts forth a new application-oriented taxonomy that comprehensively summarises five fundamental routes of motion capture technology. In addition to that, we also delve into the research priorities linked with each route, following the structure of “hardware requirements - technical routes - datasets - evaluation metrics” and extending the necessary criteria for transferring traditional motion capture systems to deep learning-based ones. Meanwhile, for the motion capture technology, the current state of the art is reviewed, the challenges are identified, and the future directions of the research are outlined. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Xiaoyu Qin 0001, Jinbao Wang 0001, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | ACR-Net: Learning High-Accuracy Optical Flow via Adaptive-Aware Correlation Recurrent NetworkabstractAlthough recurrent network-based optical flow estimation methods have shown great success in recent years, most of these methods have difficulty handling large displacements and occlusions because the existing recurrent networks are usually restricted to coarse-resolution single-scale models while ignoring the multiscale features brought by hierarchical concepts in previous coarse-to-fine approaches. In this paper, we propose an adaptive-aware correlation recurrent network for optical flow estimation, named ACR-Net, which preserves fine motion features with a single-scale resolution recurrent framework and adaptively incorporates multiscale features at different stages to achieve high-accuracy optical flow estimation. First, our proposed self-adaptation scale-aware correlation module can incorporate the adaptive correlation of multiscale inter- and intra-motion features, which makes the features more discriminative for capturing long-range dependencies between pixels. Second, our presented adaptive-aware motion module can effectively extract the required features of different kinds of motion from multilevel correspondence. Third, our introduced cross-guide motion and fusion modules can accurately guide the propagation of reliable pixels towards unreliable pixels and dynamically determine the most suitable expression to address the occlusion challenges. Comprehensive experiments demonstrate that ACR-Net outperforms existing two-view models, striking a good balance between speed and accuracy and achieving the best performance on the MPI-Sintel final pass and KITTI-2015 test datasets. The code will be made publicly available. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge, Zige Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Facial Action Unit Representation Based on Self-Supervised Learning With Ensembled Priori ConstraintsabstractFacial action units (AUs) focus on a comprehensive set of atomic facial muscle movements for human expression understanding. Based on supervised learning, discriminative AU representation can be achieved from local patches where the AUs are located. Unfortunately, accurate AU localization and characterization are challenged by the tremendous manual annotations, which limits the performance of AU recognition in realistic scenarios. In this study, we propose an end-to-end self-supervised AU representation learning model (SsupAU) to learn AU representations from unlabeled facial videos. Specifically, the input face is decomposed into six components using auto-encoders: five photo-geometric meaningful components, together with 2D flow field AUs. By constructing the canonical neutral face, posed neutral face, and posed expressional face gradually, these components can be disentangled without supervision, therefore the AU representations can be learned. To construct the canonical neutral face without manually labeled ground truth of emotion state or AU intensity, two priori knowledge based assumptions are proposed: 1) identity consistency, which explores the identical albedos and depths of different frames in a face video, and helps to learn the camera color mode as an extra cue for canonical neutral face recovery. 2) average face, which enables the model to discover a 'neutral facial expression' of the canonical neutral face and decouple the AUs in representation learning. To the best of our knowledge, this is the first attempt to design self-supervised AU representation learning method based on the definition of AUs. Substantial experiments on benchmark datasets have demonstrated the superior performance of the proposed work in comparison to other state-of-the-art approaches, as well as an outstanding capability of decomposing input face into meaningful factors for its reconstruction. The code is made available at https://github.com/Sunner4nwpu/SsupAU. Peng Zhang 0005, Chujia Guo, Ke Lu 0002, Dongmei Jiang |
IEEE Trans. Image Process. | 4 |
| 2024 | Gated Multi-Modal Edge Refinement Network for Light Field Salient Object DetectionabstractLight field can be decoded into multiple representations and provides valuable focus and depth information. This breakthrough overcomes the limitations of traditional 2D and 3D saliency detection methods, opening up new possibilities for more accurate and comprehensive analysis of visual scenes. To tackle the challenges of inaccurate edge prediction and effectively leverage the rich multi-modal light field information, we propose a gated multi-modal edge refinement network (GMERNet). It first obtains the preliminary position and structure information of the salient object and then gradually refines the object edge. This involves two modules: gated multi-modal feature complement (GMFC) module and progressive edge refinement (PER) module. The GMFC module captures dependencies across the all-in-focus image and its corresponding focal stack and depth map, effectively aggregating multiple features through gate mechanisms. The PER module progressively refines edges by combining salient object features with edge features through a cascaded structure. Experimental results demonstrate that GMERNet achieves state-of-the-art performance on five benchmark datasets and shows significant advantages in extracting salient objects with complex edges. Yefan Li, Fuqing Duan, Ke Lu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | BiUNet: Towards More Effective UNet with Bi-Level Routing Attention
Kun Dong 0001, Jian Xue 0002, Xing Lan, Ke Lu 0002 |
BMVC | 4 |
| 2023 | Semi-Supervised Contrastive Learning of Global and Local Representation for 3d Medical Image SegmentationabstractAlthough the application of supervised deep learning in medical image analysis is still very successful, it mainly depends on the quantity and quality of labeled data; and it is time-consuming and labor-intensive to obtain 3D medical image annotation. Recently, contrastive learning has shown its remarkable ability for self-supervised learning and has achieved impressive results on many downstream tasks. In this study, we extend the popular contrastive learning medical image segmentation framework to 3D and design extra reconstruction loss for volumetric medical images to improve the performance of global contrastive learning. We evaluate our method on two public 3D medical image datasets of different modalities. Our proposed method achieves competitive results compared to other methods for different proportions of labeled data. Chuang Jia, Jian Xue 0002, Ke Lu 0002, Zhongqi Wu |
ICIP | 3 |
| 2023 | A Novel Cross-Fusion Method of Different Types of Features for Image CaptioningabstractMulti-modal tasks are receiving more and more attention, including image captioning. Based on X-Linear attention, we simultaneously introduce grid features and region features extracted by Faster RCNN. We obtain a global feature vector of each type of original features through mean pooling. The two types of features are encoded by two parallel encoders. Each encoder has two inputs: a set of feature vectors (region/grid) and the corresponding global feature vector. Each encoding layer outputs an encoded global feature vector and a set of encoded feature vectors. We cross-fuse the global feature vector output by each encoding layer for region features and the set of encoded feature vectors for grid features. In the same way, we cross-fuse another pair of the global feature (grid) and the set of encoded feature vectors (region). Finally, we fuse the two global feature vectors output by the two encoders as the final global features, and the two sets of encoded feature vectors output by the two encoders as the final visual features. Experimental results on the COCO dataset show that our model achieves a new SOTA performance of BLEU-1 81.5%, BLEU-4 40.5%, METEOR 29.6%, and ROUGE 59.5% on the Karpathy test split. Liangshan Lou, Ke Lu 0002, Jian Xue 0002 |
IJCNN | 2 |
| 2023 | TIVA-KG: A Multimodal Knowledge Graph with Text, Image, Video and AudioabstractKnowledge graphs serve as a powerful tool to boost model performances for various applications covering computer vision, natural language processing, multimedia data mining, etc. The process of knowledge acquisition for human is multimodal in essence, covering text, image, video and audio modalities. However, existing multimodal knowledge graphs fail to cover all these four elements simultaneously, severely limiting their expressive powers in performance improvement for downstream tasks. In this paper, we propose TIVA-KG, a multimodal Knowledge Graph covering Text, Image, Video and Audio, which can benefit various downstream tasks. Our proposed TIVA-KG has two significant advantages over existing knowledge graphs in i) coverage of up to four modalities including text, image, video, audio, and ii) capability of triplet grounding which grounds multimodal relations to triples instead of entities. We further design a Quadruple Embedding Baseline (QEB) model to validate the necessity and efficacy of considering four modalities in KG. We conduct extensive experiments to test the proposed TIVA-KG with various knowledge graph representation approaches over link prediction task, demonstrating the benefits and necessity of introducing multiple modalities and triplet grounding. TIVA-KG is expected to promote further research on mining multimodal knowledge graph as well as the relevant downstream tasks in the community. TIVA-KG is now available at our website: http://mn.cs.tsinghua.edu.cn/tivakg. Xin Wang 0019, Benyuan Meng, Hong Chen 0011, Ke Lu 0002, Wenwu Zhu 0001 |
ACM Multimedia | 5 |
| 2023 | Attention Weighted Local DescriptorsabstractLocal features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc. Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Continuous cross-modal hashing
Hao Zheng 0008, Jinbao Wang 0001, Xiantong Zhen, Jingkuan Song, Feng Zheng 0001, Ke Lu 0002, Guo-Jun Qi |
Pattern Recognit. | 6 |
| 2023 | Region Attentive Action Unit Intensity Estimation With Uncertainty Weighted Multi-Task LearningabstractFacial action units (AUs) refer to a comprehensive set of atomic facial muscle movements. Recent works have focused on exploring complementary information by learning the relationships among AUs. Most existing approaches process AU co-occurrence and enhance AU recognition by learning the dependencies among AUs from labels, however, the complementary information among features of different AUs are ignored. Moreover, ground truth annotations suffer from a large intra-class variance and their associated intensity levels may vary depending on the annotators’ experience. In this paper, we propose the Region Attentive AU intensity estimation method with Uncertainty Weighted Multi-task Learning (RA-UWML). A RoI-Net is first used to extract features from the pre-defined facial patches where the AUs locate. Then, we use the co-occurrence of AUs using both within patch and between patches representation learning. Within a given patch, we propose sharing representation learning in a multi-task manner. To achieve complementarity and avoid redundancy between different image patches, we propose to use a multi-head self-attention mechanism to adaptively and attentively encode each patch specific representation. Moreover, the AU intensity is represented as a Gaussian distribution, instead of a single value, where the mean value indicates the most likely AU intensity and the variance indicates the uncertainty of the estimated AU intensity. The estimated variances are leveraged to automatically weight the loss of each AU in the multitask learning model. In extensive experiments on the Disfa, Fera2015 and Feafa benchmarks, it is shown that the proposed AU intensity estimation model achieves better results compared to the state-of-the-art models. Dongmei Jiang, Xiaoyong Wei, Ke Lu 0002, Hichem Sahli |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Decoupling Multimodal Transformers for Referring Video Object SegmentationabstractReferring Video Object Segmentation (RVOS) aims to segment the text-depicted object from video sequences. With excellent capabilities in long-range modelling and information interaction, transformers have been increasingly applied in existing RVOS architectures. To better leverage multimodal data, most efforts focus on the interaction between visual and textual features. However, they ignore the syntactic structures of the text during the interaction, where all textual components are intertwined, resulting in ambiguous vision-language alignment. In this paper, we improve the multimodal interaction by DECOUPLING the interweave. Specifically, we train a lightweight subject perceptron, which extracts the subject part from the input text. Then, the subject and text features are fed into two parallel branches to interact with visual features. This enables us to perform subject-aware and context-aware interactions, respectively, thus encouraging more explicit and discriminative feature embedding and alignment. Moreover, we find the decoupled architecture also facilitates incorporating the vision-language pre-trained alignment into RVOS, further improving the segmentation performance. Experimental results on all RVOS benchmark datasets demonstrate the superiority of our proposed method over the state-of-the-arts. The code of our method is available at:https://github.com/gaomingqi/dmformer. Mingqi Gao 0003, Jungong Han, Ke Lu 0002, Feng Zheng 0001, Giovanni Montana |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Cancelable Fingerprint Template Construction Using Vector Permutation and Shift-OrderingabstractThe need for cancelable biometric techniques has seen a progressive rise due to the rapid deployment of biometric authentication systems. These techniques prevent compromising biometric data by generating and using their corresponding cancelable templates for user authentication. However, the non-invertible distance preserving transformation methods employed in various schemes are often vulnerable to information leakage since matching is performed in the transform domain. This paper proposed a non-invertible distance preserving scheme based on vector permutation and shift-order process. First, the dimension of feature vectors is reduced using kernelized principal component analysis before randomly permuting the extracted vector features. A shift-order process is then applied to the generated features to achieve non-invertibility and combat similarity correlation-based attacks. The generated hash codes are resilient to various security and privacy attacks such as ARM, masquerade, and brute-force preimage. Experimental evaluations conducted on eight fingerprint datasets from FVC2002, FVC2004, and FVC2006 reveal a high matching performance of the proposed method with better recognition accuracy than other existing state-of-the-art. The scheme also fulfills the revocability and unlinkability requirements of cancelable biometrics. Sani M. Abdullahi, Ke Lu 0002, Shuifa Sun, Hongxia Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | Boosting Semantic Segmentation of Aerial Images via Decoupled and Multilevel Compaction and DispersionabstractSemantic segmentation is a valuable task in practical applications for aerial images. Nevertheless, the segmentation performance is unsatisfactory due to aerial images’ huge intra-class variance and inter-class similarity. To solve this problem, we propose an approach to increase the distinction between classes and compact the features of the same class. Specifically, since a single aerial image contains only a small number of categories, which is fatal for previous contrastive learning, we discard InfoNCE loss in contrastive learning and use the simple Mean Square Error (MSE) loss that does not require negative samples to decouple the dispersion and compaction operations. Besides, we set up more representative prototypes for classes and extend the prototypes to the whole dataset level, which we call image- and dataset-level prototypes. Based on the calculated prototypes, we propose Multi-level intra-class Feature Compaction (MFC) and Multi-level inter-class Feature Dispersion (MFD) to compact the features of the same class and disperse the features of different classes in the latent feature space. More importantly, some measures are proposed to ensure the two do not conflict. MFC and MFD can be applied to any existing segmentation network to improve performance significantly without increasing computational complexity during inference. Moreover, we feed the calculated multi-level prototypes directly into the classifier, thus keeping the feature extraction and classifier consistent. Results on four challenging datasets, Deepglobe, iSAID, Potsdam, and Vaihingen, demonstrate the significant effect of our method, and sufficient ablation studies verify the role of each module. Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Viewpoint-Adaptive Representation Disentanglement Network for Change CaptioningabstractChange captioning is to describe the fine-grained change between a pair of images. The pseudo changes caused by viewpoint changes are the most typical distractors in this task, because they lead to the feature perturbation and shift for the same objects and thus overwhelm the real change representation. In this paper, we propose a viewpoint-adaptive representation disentanglement network to distinguish real and pseudo changes, and explicitly capture the features of change to generate accurate captions. Concretely, a position-embedded representation learning is devised to facilitate the model in adapting to viewpoint changes via mining the intrinsic properties of two image representations and modeling their position information. To learn a reliable change representation for decoding into a natural language sentence, an unchanged representation disentanglement is designed to identify and disentangle the unchanged features between the two position-embedded representations. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the four public datasets. The code is available at https://github.com/tuyunbin/VARD. Yunbin Tu, Liang Li 0003, Li Su 0003, Junping Du 0001, Ke Lu 0002, Qingming Huang |
IEEE Trans. Image Process. | 5 |
| 2023 | Semi-Supervised Medical Image Segmentation With Voxel Stability and Reliability ConstraintsabstractSemi-supervised learning is becoming an effective solution in medical image segmentation because annotations are costly and tedious to acquire. Methods based on the teacher-student model use consistency regularization and uncertainty estimation and have shown good potential in dealing with limited annotated data. Nevertheless, the existing teacher-student model is seriously limited by the exponential moving average algorithm, which leads to the optimization trap. Moreover, the classic uncertainty estimation method calculates the global uncertainty for images but does not consider local region-level uncertainty, which is unsuitable for medical images with blurry regions. In this article, the Voxel Stability and Reliability Constraint (VSRC) model is proposed to address these issues. Specifically, the Voxel Stability Constraint (VSC) strategy is introduced to optimize parameters and exchange effective knowledge between two independent initialized models, which can break through the performance bottleneck and avoid model collapse. Moreover, a new uncertainty estimation strategy, the Voxel Reliability Constraint (VRC), is proposed for use in our semi-supervised model to consider the uncertainty at the local region level. We further extend our model to auxiliary tasks and propose a task-level consistency regularization with uncertainty estimation. Extensive experiments on two 3D medical image datasets demonstrate that our method outperforms other state-of-the-art semi-supervised medical image segmentation methods under limited supervision. Yang Zhao 0028, Ke Lu 0002, Jian Xue 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Show, Tell and Rephrase: Diverse Video Captioning via Two-Stage Progressive TrainingabstractDescribing a video using natural language is an inherently one-to-many translation task. To generate diverse captions, existing VAE-based generative models typically learn factorized latent codes via one-stage training merely from stand-alone video-caption pairs. However, such a paradigm neglects set-level relationships among captions from the same video, not fully capturing the underlying multimodality of the generative process. To overcome this shortcoming, we leverage neighbouring descriptions for the same video that are articulated with noticeable topics and language variations (i.e., paraphrases). To this end, we propose a novel progressive training method by decomposing the learning of latent variables into two stages that are topic-oriented and paraphrase-oriented, respectively. Specifically, the model learns from divergent topic sentences obtained by semantic-based clustering in the first stage. It is then trained again through paraphrases with a cluster-aware adaptive regularization, allowing more intra-cluster variations. Furthermore, we introduce an overall metric DAUM, aDiversity-AccuracyUnifiedMetric to consider both the precision of the generated caption set and its coverage on the reference set, which has proved to have a higher correlation with human judgment than previous precision-only metrics. Extensive experiments on three large-scale video datasets show that the proposed training strategy can achieve superior performance in terms of accuracy, diversity, and DAUM over several baselines. Zhu Liu 0005, Teng Wang 0007, Feng Zheng 0001, Ke Lu 0002 |
IEEE Trans. Multim. | 6 |
| 2023 | Neighborhood Contrastive Transformer for Change CaptioningabstractChange captioning is to describe the semantic change between a pair of similar images in natural language. It is more challenging than general image captioning, because it requires capturing fine-grained change information while being immune to irrelevant viewpoint changes, and solving syntax ambiguity in change descriptions. In this paper, we propose a neighborhood contrastive transformer to improve the model's perceiving ability for various changes under different scenes and cognition ability for complex syntax structure. Concretely, we first design a neighboring feature aggregating to integrate neighboring context into each feature, which helps quickly locate the inconspicuous changes under the guidance of conspicuous referents. Then, we devise a common feature distilling to compare two images at neighborhood level and extract common properties from each image, so as to learn effective contrastive information between them. Finally, we introduce the explicit dependencies between words to calibrate the transformer decoder, which helps better understand complex syntax structure during training. Extensive experimental results demonstrate that the proposed method achieves the state-of-the-art performance on three public datasets with different change scenarios. The code is available athttps://github.com/tuyunbin/NCT. Yunbin Tu, Liang Li 0003, Li Su 0003, Ke Lu 0002, Qingming Huang |
IEEE Trans. Multim. | 4 |
| 2023 | A Facial Landmark Detection Method Based on Deep Knowledge TransferabstractFacial landmark detection is a crucial preprocessing step in many applications that process facial images. Deep-learning-based methods have become mainstream and achieved outstanding performance in facial landmark detection. However, accurate models typically have a large number of parameters, which results in high computational complexity and execution time. A simple but effective facial landmark detection model that achieves a balance between accuracy and speed is crucial. To achieve this, a lightweight, efficient, and effective model is proposed called the efficient face alignment network (EfficientFAN) in this article. EfficientFAN adopts the encoder-decoder structure, with a simple backbone EfficientNet-B0 as the encoder and three upsampling layers and convolutional layers as the decoder. Moreover, deep dark knowledge is extracted through feature-aligned distillation and patch similarity distillation on the teacher network, which contains pixel distribution information in the feature space and multiscale structural information in the affinity space of feature maps. The accuracy of EfficientFAN is further improved after it absorbs dark knowledge. Extensive experimental results on public datasets, including 300 Faces in the Wild (300W), Wider Facial Landmarks in the Wild (WFLW), and Caltech Occluded Faces in the Wild (COFW), demonstrate the superiority of EfficientFAN over state-of-the-art methods. Ke Lu 0002, Jian Xue 0002, Jiayi Lyu, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | TransZero: Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which are strong prior for localization of object attribute for representing discriminative region features enabling significant visual-semantic interaction. Although few attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network to learn the attribute localization for discriminative visual-semantic embedding representations in ZSL, termed TransZero. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks and improve the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the most relevant image regions to each attributes from a given image under the guidance of attribute semantic information. Then, the locality-augmented visual features and semantic vectors are used for conducting effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves a new state-of-the-art on three ZSL benchmarks. The codes are available at: https://github.com/shiming-chen/TransZero. Shiming Chen 0002, Ziming Hong, Yang Liu 0069, Guosen Xie, Baigui Sun, Hao Li 0030, Qinmu Peng, Ke Lu 0002, Xinge You |
AAAI | 8 |
| 2022 | Shape-Adaptive Selection and Measurement for Oriented Object DetectionabstractThe development of detection methods for oriented object detection remains a challenging task. A considerable obstacle is the wide variation in the shape (e.g., aspect ratio) of objects. Sample selection in general object detection has been widely studied as it plays a crucial role in the performance of the detection method and has achieved great progress. However, existing sample selection strategies still overlook some issues: (1) most of them ignore the object shape information; (2) they do not make a potential distinction between selected positive samples; and (3) some of them can only be applied to either anchor-free or anchor-based methods and cannot be used for both of them simultaneously. In this paper, we propose novel flexible shape-adaptive selection (SA-S) and shape-adaptive measurement (SA-M) strategies for oriented object detection, which comprise an SA-S strategy for sample selection and SA-M strategy for the quality estimation of positive samples. Specifically, the SA-S strategy dynamically selects samples according to the shape information and characteristics distribution of objects. The SA-M strategy measures the localization potential and adds quality information on the selected positive samples. The experimental results on both anchor-free and anchor-based baselines and four publicly available oriented datasets (DOTA, HRSC2016, UCAS-AOD, and ICDAR2015) demonstrate the effectiveness of the proposed method. Liping Hou, Ke Lu 0002, Jian Xue 0002, Yuqiu Li |
AAAI | 2 |
| 2022 | Exploring Visual Context for Weakly Supervised Person SearchabstractPerson search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are available. However, annotating identities is labor-intensive, limiting the practicability and scalability of current frameworks. This paper inventively considers weakly supervised person search with only bounding box annotations. We propose to address this novel task by investigating three levels of context clues (i.e., detection, memory and scene) in unconstrained natural images. The first two are employed to promote local and global discriminative capabilities, while the latter enhances clustering accuracy. Despite its simple design, our CGPS boosts the baseline model by 8.8% in mAP on CUHK-SYSU. Surprisingly, it even achieves comparable performance with several supervised person search models. Our code is available at https://github. com/ljpadam/CGPS. Yichao Yan, Jinpeng Li 0004, Shengcai Liao, Jie Qin 0004, Bingbing Ni, Ke Lu 0002, Xiaokang Yang 0001 |
AAAI | 6 |
| 2022 | Vote-Based Multi-Level Context Attention Network for 3D Point Cloud Object Detectionabstract3D object detection is a challenging task because point clouds are characterized by sparsity and irregularity. Most state-of-the-art detectors recognize objects individually without considering the rich context relationships of objects at different levels. In this paper, we propose an end-to-end vote-based multi-level context attention network. Specifically, a Patch-Context-Module is designed to extract multi-level context features among point patches. Meanwhile, because low-level features contain fine location description information, a Spatial-Context-Module is adopted to combine low-level spatial and semantic features. Furthermore, a Fusion Sampling and Aggregation module is proposed to consider additional semantic information of each vote point, thereby increasing the ratio of positive points and improving detection performance. Finally, the Class-IoU-Guide NMS with an adaptive threshold is implemented to suppress false detection at the inference time. Experiments on the ScanNetV2 and SUN RGB-D datasets demonstrated that our proposed method out-performs current state-of-the-art approaches. Ke Lu 0002, Jian Xue 0002, Liping Hou, Hengsheng Lun |
ICME | 2 |
| 2022 | Improved Transformer with Parallel Encoders for Image CaptioningabstractImage captioning is currently one of the most important multimodal tasks. With Transformer proposed, many Transformer-based models have achieved good performance in image captioning. However, substantial work is still required to improve the performance in the field of image captioning. We propose an improved model that uses the meshed-memory Transformer as its backbone. We propose the use of region features and grid features together. In addition, we use two identical parallel encoders to process region features and grid features separately, and fuse the outputs of each layer of the two encoders to form one of the inputs of the decoder. We comprehensively compare the performance of our model with the existing state-of-the-art models on the official COCO dataset. Experiments show that, on the Karpathy test split, our model outperforms the backbone on all evaluation metrics: for example, it increases BLEU-1 from 80.8% to 81.4%, and CIDEr from 131.2% to 133.5%. Liangshan Lou, Ke Lu 0002, Jian Xue 0002 |
ICPR | 2 |
| 2022 | Uncertainty-Aware Semi-Supervised Learning of 3D Face Rigging from Single ImageabstractWe present a method to rig 3D faces via Action Units (AUs), viewpoint and light direction, from single input image. Existing 3D methods for face synthesis and animation rely heavily on 3D morphable model (3DMM), which was built on 3D data and cannot provide intuitive expression parameters, while AU-driven 2D methods cannot handle head pose and lighting effect. We bridge the gap by integrating a recent 3D reconstruction method with 2D AU-driven method in a semi-supervised fashion. Built upon the auto-encoding 3D face reconstruction model that decouples depth, albedo, viewpoint and light without any supervision, we further decouple expression from identity for depth and albedo with a novel conditional feature translation module and pretrained critics for AU intensity estimation and image classification. Novel objective functions are designed using unlabeled in-the-wild images and in-door images with AU labels. We also leverage uncertainty losses to model the probably changing AU region of images as input noise for synthesis, and model the noisy AU intensity labels for intensity estimation of the AU critic. Experiments with face editing and animation on four datasets show that, compared with six state-of-the-art methods, our proposed method is superior and effective on expression consistency, identity similarity and pose similarity. Hichem Sahli, Ke Lu 0002, Dongmei Jiang |
ACM Multimedia | 4 |
| 2022 | Dual-Branch Point Cloud Feature Learning for 3D Object DetectionabstractIncomplete feature information is a key problem that limits 3D point cloud object detection and its applications. Many state-of-the-art detectors address this problem from different perspectives, but a comprehensive solution has not yet been obtained. In this paper, a solution is proposed that consists of two branches, one for channel-wise local feature learning and one for spatial-wise global feature learning. The combination of the local features, global contextual features, channel-wise attention features, and spatial attention features of the 3D point cloud is obtained through the two branches. Specifically, a generic spatial self-attention model is proposed that uses skeleton convolution to enhance the extraction of spatial features and combines it with a self-attention mechanism to improve the learning of global features. Further, the proposed skeleton attention mechanism focuses on object contour and rotation invariance of the point cloud. Through sufficient experimental validation, all module proposed in this paper are shown to have a substantial effect on performance and the visualization results demonstrate the importance of our approach. Hengsheng Lun, Jian Xue 0002, Ke Lu 0002 |
SMC | 3 |
| 2022 | Point cloud decomposition by internal and external critical points
Yinghui Wang 0001, Xiaojuan Ning, Ke Lu 0002 |
Comput. Graph. | 4 |
| 2022 | Human-object interaction detection via interactive visual-semantic graph learning
Tongtong Wu, Fuqing Duan, Liang Chang 0001, Ke Lu 0002 |
Sci. China Inf. Sci. | 4 |
| 2022 | DBRANet: Road Extraction by Dual-Branch Encoder and Regional Attention DecoderabstractAlthough widely exploited in recent decades, road extraction is still a very significant and challenging research in the field of remote sensing image processing due to the complex background and road distribution. Among the existing CNN-based methods, U-shape architectures composed of encoders and decoders have shown their effectiveness. In this letter, we propose an improved encoder–decoder method, named DBRANet, for extracting roads from remote sensing images. In the encoding phase, we present a dual-branch network module (DBNM) to construct more effective features, thus improving the fusion feature maps of different scales. One branch utilizes the residual block, and the other branch utilizes the refined asymmetric block, which effectively increases the feature extraction capability of the backbone. In the decoding phase, considering the sinuous shape and the unbalanced distribution of roads in remote sensing images, we design a novel attention module, named the regional attention network module (RANM), to automatically learn the importance of each channel according to the regional information. Extensive experiments on several public remote sensing road data sets show that our DBRANet achieves higher segmentation [$F1$score and Intersection over Union (IoU)] and connectivity [average path length similarity (APLS)] accuracy, which verifies the effectiveness of our approach. Sibao Chen 0001, Yu-Xin Ji, Jin Tang 0001, Bin Luo 0001, Weiqiang Wang 0001, Ke Lu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Light Attention Embedding for Facial Expression RecognitionabstractFacial expression recognition is important for human–computer interaction and other applications. Several facial expression datasets have been published in recent decades and have enabled improvements in algorithms for classifying emotions. However, recognition of realistic expressions in real-world conditions is still challenging because of uncontrolled conditions, such as lighting, brightness, pose, and occlusion. In this paper, we propose a light attention embedding network based on the spatial attention mechanism (LAENet-SA), which can focus on locations in an image that are relevant to emotion. LAENet-SA allows a small number of attention modules to be embedded and can be constructed from typical convolutional neural networks. The performance of LAENet-SA on facial expression recognition has been validated on three facial expression datasets, including a lab-controlled dataset and two in-the-wild datasets. Experimental results show that LAENet-SA improved the performance on each dataset, compared with state-of-the-art methods, and achieved better generalization when tested on facial images with occlusion. Jian Xue 0002, Ke Lu 0002, Yanfu Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Class-Incremental Learning for Semantic Segmentation in Aerial Imagery via Distillation in All AspectsabstractIncremental learning using neural networks achieves great success in semantic segmentation but still suffers from catastrophic forgetting. In this article, we propose an effective class-incremental segmentation method without storing old data. To alleviate the issue of forgetting, we present two important modules, i.e., the deep feature distillation (DFD) module and the label mixed (LM) module. The DFD module is established to learn a good feature representation of old classes by distilling a new compact feature representation from different layers of networks. The proposed LM module first identifies the examples (pixels) of old classes with high confidences utilizing the output of old models, and then, they are combined with examples of new classes to supervise the training of new models, which can achieve a good balance between learning new classing and avoiding forgetting old ones. Our ablation studies show that the DFD module and the LM module can make the learning network obtain 6.2% and 15% performance gains [mean Intersection over Union (mIOU)], respectively. Furthermore, by introducing the supervision of output distillation loss, we compare our method with several state-of-the-art methods in the extensive experiments, and the experimental results all show that our method is significantly superior to them on the dataset of aerial images. Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Class-Incremental Semantic Segmentation of Aerial Images via Pixel-Level Feature Generation and Task-Wise DistillationabstractDeep neural networks achieve significant progress in semantic segmentation but still suffer from the catastrophic forgetting problem, i.e., networks will forget old classes as they learn new ones. In this article, we propose an effective class-incremental segmentation framework without storing old data. Specifically, to alleviate the issue of catastrophic forgetting, we present two important modules, i.e., the Pixel-level Feature Generation (PFG) module, and the Task-wise Knowledge distillation (TKD) module. The PFG module is designed to constantly generate any number of features of the old classes to keep the old memory. The PFG module is the first attempt to use the generative method in class-incremental segmentation of aerial images, and it abandons the previous image generation approach but to generate pixel-level features, which is more suitable for the segmentation task. Meanwhile, the proposed TKD module is specially designed for class incremental tasks, and it only compares classes in the same learning step (task), thus avoiding the squeezing of new classes to old classes when the output is normalized (softmax), making distillation more effective. Sufficient experiments show that our method is remarkably effective and achieves more than 4.5% gains compared with state-of-the-art methods, and more than 13% compared with baselines, on all learning conditions. The ablation studies show that the PFG module and the TKD module are both indispensables. Besides, the proposed framework can be well combined with any existing class incremental learning method to achieve better performance. Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Spatial-Temporal Pyramid Graph Reasoning for Action RecognitionabstractSpatial-temporal relation reasoning is a significant yet challenging problem for video action recognition. Previous works typically apply local operations like 2D or 3D CNNs to conduct space-time interactions in video sequences, or simply capture space-time long-range relations of a single fixed scale. However, this is inadequate for obtaining a comprehensive action representation. Besides, most models treat all input frames equally for the final classification, without selecting key frames and motion-sensitive regions. This introduces irrelevant video content and hurts the performance of models. In this paper, we propose a generic Spatial-Temporal Pyramid Graph Network (STPG-Net) to adaptively capture long-range spatial-temporal relations in video sequences at multiple scales. Specifically, we design a temporal attention (TA) module and a spatial-temporal attention (STA) module to learn the contribution of each frame and each space-time region to an action at a feature level, respectively. We then apply the selected key information to build spatial-temporal pyramid graphs for long-range relation reasoning and more comprehensive action representation learning. STPG-Net can be flexibly integrated into 2D and 3D backbone networks in a plug-and-play manner. Extensive experiments show that it brings consistent improvements over many challenging baselines on several standard action recognition benchmarks (i.e., Something-Something V1 & V2, and FineGym), demonstrating the effectiveness of our approach. Tiantian Geng, Feng Zheng 0001, Xiaorong Hou, Ke Lu 0002, Guo-Jun Qi, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Refined One-Stage Oriented Object Detection Method for Remote Sensing ImagesabstractMulti-class object detection in remote sensing images plays an important role in many applications but remains a challenging task because of scale imbalance and arbitrary orientations of the objects with extreme aspect ratios. In this paper, the Asymmetric Feature Pyramid Network (AFPN), Dynamic Feature Alignment (DFA) module, and Area-IoU regression loss are proposed on the basis of a one-stage cascaded detection method for the detection of multi-class objects with arbitrary orientations in remote sensing images. The designed asymmetric convolutional block is embedded into the AFPN for handling objects with extreme aspect ratios and improving the space representation with ignorable increases in calculation. The DFA module is proposed to dynamically align mismatched features, which are caused by the deviation between predefined anchors and arbitrarily oriented predicted boxes. The refined Area-IoU regression loss, which reconciles two new regression loss functions, the area-guided regression loss and IoU-guided regression loss, is proposed to simultaneously solve the scale imbalance problem and angle sensitivity problem. Experiments on three publicly available datasets, DOTA, HRSC2016, and ICDAR2015, show the effectiveness of the proposed method. Liping Hou, Ke Lu 0002, Jian Xue 0002 |
IEEE Trans. Image Process. | 2 |
| 2022 | Fine-Grained Categorization From RGB-D ImagesabstractIn the field of computer vision, fine-grained visual categorization has attracted a lot of attention and made great progress due to convolutional neural networks and a large number of publicly available datasets. With next-generation sensing technology, RGB-D cameras can provide high-quality synchronized RGB and depth images for solving many computer vision problems. Although RGB-D cameras have been used in the context of multi-view object category detection and scene understanding, they have not been widely used in fine-grained classification. In this paper, we introduce a multiview RGB-D dataset RGBD-FG for fine-grained categorization. Currently, the dataset contains 93 051 RGB-D images covering 19 super-categories and 50 sub-categories of common vegetables and fruit, and is organized in a hierarchical manner. We provide extensive experimental results to establish state-of-the-art benchmarks for our dataset, illustrating its diversity and scope for improvement through future work. We also propose a novel modality-specific multimodal network called FS-Multimodal network, which can solve two limitations of multimodal networks trained based on fine-tuning techniques: over-fitting and lack of effective depth-specific features. We hope that our study lays the foundations for fine-grained categorization of RGB-D data. Yanhao Tan, Mohammad Muntasir Rahman, Yanfu Yan, Jian Xue 0002, Ling Shao 0001, Ke Lu 0002 |
IEEE Trans. Multim. | 6 |
| 2021 | Spatiotemporal Features and Local Relationship Learning for Facial Action Unit Intensity RegressionabstractThe action units (AU) encoded by the Facial Action Coding System (FACS) have been widely used in the representation of facial expressions. Although work on automatic facial AU detection has achieved quite good results in recent years, there remains much research potential for more accurate AU detection and intensity regression. Moreover, most work only considers the spatial information and ignores the temporal information. In practice, changes in facial AUs involve both spatial and temporal variation. In this paper, by extracting multi-scale spatial features and corresponding temporal features from the faces in the video image sequence, and learning the local relationship of the spatiotemporal features we propose a method that can obtain robust and accurate regression for AU intensity. The proposed method outperforms the baseline system on FEAFA dataset and obtains comparable performance on DISFA dataset. Ke Lu 0002, Jian Xue 0002 |
ICIP | 2 |
| 2021 | Human Pose Estimation based on Attention Multi-resolution NetworkabstractRecently, multi-resolution neural networks, which combine features of different resolutions, have achieved good results in human pose estimation tasks. In this paper, we propose an attention-mechanism-based multi-resolution network, which adds an attention mechanism to the High-Resolution Network (HRNet) to enhance the feature representation of the network. It improves the ability of networks with different resolutions to extract key features from images, and causes the output to contain more effective multi-resolution representation information, so that the corresponding point positions of human joints can be estimated more accurately. Experiments on the MPII and COCO datasets, and verification on the MPII datasets, obtained an average accuracy of 90.3% under the [email protected] evaluation standard, and good results were also achieved on the COCO dataset (with an AP of 76.5). The experimental results show that our network model is effective in improving the accuracy of key point estimation in the human pose estimation task. Qixiang Sun, Xiaojie Yin, Ke Lu 0002 |
ICMR | 5 |
| 2021 | Multi-view 3D Smooth Human Pose Estimation based on Heatmap Filtering and Spatio-temporal InformationabstractThe estimation of 3D human poses from time-synchronized, calibrated multi-view video usually consists of two steps: (1) a 2D detector to locate the 2D coordinate point position of the joint via heatmaps for each frame and (2) a post-processing method such as the recursive pictorial structure model or robust triangulation to obtain 3D coordinate points. However, most existing methods are based on a single frame only. They do not take advantage of the temporal characteristics of the video sequence itself, and must rely on post-processing algorithms. They are also susceptible to human self-occlusion, and the generated sequences suffer from jitter. Therefore, we propose a network model incorporating spatial and temporal features. Using a coarse-to-fine approach, the proposed heatmap temporal network (HTN) generates temporal heatmap information, with an occlusion heatmap filter used to filter low-quality heatmaps before they are sent to the HTN. The heatmap fusion and the triangulation weights are dynamically adjusted, and intermediate supervision is employed to enable better integration of temporal and spatial information. Our network is also end-to-end differentiable. This overcomes the long-standing problem of skeleton jitter being generated and ensures that the sequence is smooth and stable. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Runchen Wei |
ACM Multimedia | 2 |
| 2021 | A Weibull-distribution-based hybrid total variation method for speckle reduction in ultrasound imagesabstractAbstract Speckle reduction is still an intractable task in ultrasound imaging field. Ultrasound speckle is usually described as multiplicative noise with its statistics following a Rayleigh or Gaussian distribution. To employ these two distributions effectively, the authors attempt to describe ultrasound speckle using a Weibull distribution, because it can include the Rayleigh distribution as a special case and also approximate a Gaussian distribution by varying its shape and scale parameters. The authors’ contribution in this paper is to propose a Weibull‐distribution‐based hybrid total variation (WHTV) method to reduce ultrasound speckle. The WHTV energy functional is convex and consists of a new data fidelity term and a new regularization term. The former is derived from the multiplicative Weibull model of ultrasound speckle based on the maximum likelihood criterion. The latter is a new edge‐weighted combination of the first‐ and second‐order total variation, with the advantage of preserving edges while alleviating the staircase effects. The minimization of the WHTV energy functional is implemented by the split Bregman algorithm. Experimental results on synthetic and real ultrasound images have demonstrated not only that the Weibull distribution is a better fitting model for the statistics of ultrasound speckle than other distributions such as Rayleigh, Gaussian, Gamma, and Nakagami, but also that the proposed WHTV method can achieve better despeckling performance than several state‐of‐the‐art variational methods. Wenchao Cui, Liangzhi Shao, Guoqiang Gong, Ke Lu 0002, Shuifa Sun, Yirong Wu, Yiyuan Zhou |
IET Image Process. | 4 |
| 2021 | SHetConv: target keypoint detection based on heterogeneous convolution neural networks
Xiaojie Yin, Ke Lu 0002 |
Multim. Syst. | 4 |
| 2021 | Group-Wise Learning for Aurora Image Classification With Multiple RepresentationsabstractIn conventional aurora image classification methods, it is general to employ only one single feature representation to capture the morphological characteristics of aurora images, which is difficult to describe the complicated morphologies of different aurora categories. Although several studies have proposed to use multiple feature representations, the inherent correlation among these representations are usually neglected. To address this problem, we propose a group-wise learning (GWL) method for the automatic aurora image classification using multiple representations. Specifically, we first extract the multiple feature representations for aurora images, and then construct a graph in each of multiple feature spaces. To model the correlation among different representations, we partition multiple graphs into several groups via a clustering algorithm. We further propose a GWL model to automatically estimate class labels for aurora images and optimal weights for the multiple representations in a data-driven manner. Finally, we develop a label fusion approach to make a final classification decision for new testing samples. The proposed GWL method focuses on the diverse properties of multiple feature representations, by clustering the correlated representations into the same group. We evaluate our method on an aurora image data set that contains 12 682 aurora images from 19 days. The experimental results demonstrate that the proposed GWL method achieves approximately 6% improvement in terms of classification accuracy, compared to the methods using a single feature representation. Jun Zhang 0018, Mingxia Liu 0001, Ke Lu 0002, Yue Gao 0002 |
IEEE Trans. Cybern. | 3 |
| 2021 | Learning Efficient Hash Codes for Fast Graph-Based Data Similarity RetrievalabstractTraditional operations, e.g. graph edit distance (GED), are no longer suitable for processing the massive quantities of graph-structured data now available, due to their irregular structures and high computational complexities. With the advent of graph neural networks (GNNs), the problems of graph representation and graph similarity search have drawn particular attention in the field of computer vision. However, GNNs have been less studied for efficient and fast retrieval after graph representation. To represent graph-based data, and maintain fast retrieval while doing so, we introduce an efficient hash model with graph neural networks (HGNN) for a newly designed task (i.e. fast graph-based data retrieval). Due to its flexibility, HGNN can be implemented in both an unsupervised and supervised manner. Specifically, by adopting a graph neural network and hash learning algorithms, HGNN can effectively learn a similarity-preserving graph representation and compute pair-wise similarity or provide classification via low-dimensional compact hash codes. To the best of our knowledge, our model is the first to address graph hashing representation in the Hamming space. Our experimental results reach comparable prediction accuracy to full-precision methods and can even outperform traditional models in some cases. In real-world applications, using hash codes can greatly benefit systems with smaller memory capacities and accelerate the retrieval speed of graph-structured data. Hence, we believe the proposed HGNN has great potential in further research. Jinbao Wang 0001, Feng Zheng 0001, Ke Lu 0002, Jingkuan Song, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | A Coarse-to-Fine Facial Landmark Detection Method Based on Self-attention MechanismabstractFacial landmark detection in the wild remains a challenging problem in computer vision. Deep learning-based methods currently play a leading role in solving this. However, these approaches generally focus on local feature learning and ignore global relationships. Therefore, in this study, a self-attention mechanism is introduced into facial landmark detection. Specifically, a coarse-to-fine facial landmark detection method is proposed that uses two stacked hourglasses as the backbone, with a new landmark-guided self-attention (LGSA) block inserted between them. The LGSA block learns the global relationships between different positions on the feature map and allows feature learning to focus on the locations of landmarks with the help of a landmark-specific attention map, which is generated in the first-stage hourglass model. A novel attentional consistency loss is also proposed to ensure the generation of an accurate landmark-specific attention map. A new channel transformation block is used as the building block of the hourglass model to improve the model's capacity. The coarse-to-fine strategy is adopted during and between phases to reduce complexity. Extensive experimental results on public datasets demonstrate the superiority of our proposed method against state-of-the-art models. Ke Lu 0002, Jian Xue 0002, Ling Shao 0001, Jiayi Lyu |
IEEE Trans. Multim. | 2 |
| 2020 | Cascade Detector With Feature Fusion For Arbitrary-Oriented Objects In Remote Sensing ImagesabstractDetection of multi-class rotated objects is a challenging task in optical remote sensing images because of large-scale variations, arbitrary orientations and complex backgrounds, etc. Most of the state-of-the-art object detectors for natural images, that use horizontal bounding boxes, are not suitable for oriented objects in remote sensing images. In this paper, we propose an end-to-end cascade detector that can effectively detect rotated objects in complex remote sensing images. Specifically, a feature fusion block is designed to capture features with more details. Meanwhile, a supervised spatial attention mechanism is adopted to improve performance in detecting objects with complex backgrounds by weakening noise and enhancing object regions. Finally, to obtain more accurate object position, a cascade of multi-step detection subnet is implemented to refine anchors. Experiments using a publicly available remote sensing dataset DOTA show that our object detector achieves superior performance over other state-of-the-art approaches. Liping Hou, Ke Lu 0002, Jian Xue 0002 |
ICME | 2 |
| 2020 | Rgbd-Fg: A Large-Scale Rgb-D Dataset For Fine-Grained CategorizationabstractFine-grained visual categorization (FGVC) has received a great deal of attention in recent years. Currently, several public datasets are available for FGVC. However, all these datasets were created with RGB images. RGB-D sensors can provide high-quality synchronized video in terms of both color and depth. In this paper, we introduce a multi-view, large-scale RGB-D dataset called RGBD-FG to establish a novel benchmark for FGVC in RGB-D images. RGBD-FG contains RGB data and the corresponding depth data of vegetables and fruits. Our dataset was captured by a depth sensor and contained 50 categories and a total of 93,051 RGBD images with labels, and organized in a hierarchical manner. Additionally, we used several strategies on our dataset for FGVC, including a multi-modal deep CNN. We present extensive experimental results to create state-of-the-art baselines for the dataset. We hope that this dataset can fill the gap of FGVC among the RGB-D datasets. Yanhao Tan, Ke Lu 0002, Mohammad Muntasir Rahman, Jian Xue 0002 |
ICME | 2 |
| 2020 | Global-Local Attention Network for Semantic Segmentation in Aerial ImagesabstractErrors in semantic segmentation could be classified into two types: the large area misclassification and inaccurate local boundaries. Previously attention-based methods typically capture rich global contextual information, which benefits the large area classification but cannot address the local errors of boundaries. In this paper, we propose a Global-Local Attention Network (GLANet) which can simultaneously consider the global context and local details. Specifically, our GLANet consists of two branches: (1) the global attention branch and (2) local attention branch. Furthermore, three different modules are embedded in GLANet for respectively modelling the semantic interdependencies in spatial, channel and boundary dimension. Lastly, we merge the outputs of different branches to enhance the feature representation further, resulting in more precise segmentation. Overall, the proposed method achieves the competitive segmentation accuracy on two public aerial image datasets, bringing significant improvements over the existing baselines. Minglong Li, Lianlei Shan, Xiaobin Li 0006, Dengji Zhou, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001 |
ICPR | 7 |
| 2020 | UHRSNet: A Semantic Segmentation Network Specifically for Ultra-High-Resolution ImagesabstractSemantic segmentation is a basic task in computer vision, but only limited attention has been devoted to the ultra-high-resolution (UHR) image segmentation. Since UHR images occupy too much memory, they cannot be directly put into GPU for training. Previous methods are cropping images to small patches or downsampling the whole images. Cropping and downsampling cause the loss of contexts and details, which is essential for segmentation accuracy. To solve this problem, we improve and simplify the local and global feature fusion method in previous works. Local features are extracted from patches and global features are from downsampled images. Meanwhile, we propose one new fusion called local feature fusion for the first time, which can make patches get information from surrounding patches. We call the network with these two fusions ultra-high-resolution segmentation network (UHRSNet). These two fusions can effectively and efficiently solve the problem caused by cropping and downsampling. Experiments show a remarkable improvement on Deepglobe dataset [1]. Lianlei Shan, Minglong Li, Xiaobin Li 0006, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001, Weiqiang Wang 0001 |
ICPR | 5 |
| 2020 | EfficientFAN: Deep Knowledge Transfer for Face AlignmentabstractFace alignment plays an important role in many applications that process facial images. At present, deep learning-based methods have achieved excellent results in face alignment. However, these models usually have a large number of parameters, resulting in high computational complexity and execution time. In this paper, a lightweight, efficient, and effective model is proposed and named Efficient Face Alignment Network (EfficientFAN). EfficientFAN adopts the encoder-decoder structure, using a simple backbone Efficient-Net-B0 as the encoder and three deconvolutional layers as the decoder. Compared with state-of-the-art models, it achieves equivalent performance with fewer model parameters, lower computation cost, and higher speed. Moreover, the accuracy of EfficientFAN is further improved by transferring deep knowledge of a complex teacher network through feature-aligned distillation and patch similarity distillation. Extensive experimental results on public data sets demonstrate the superiority of EfficientFAN over state-of-the-art methods. Ke Lu 0002, Jian Xue 0002 |
ICMR | 2 |
| 2020 | A Lightweight Gated Global Module for Global Context Modeling in Neural NetworksabstractGlobal context modeling has been used to achieve better performance in various computer-vision-related tasks, such as classification, detection, segmentation and multimedia retrieval applications. However, most of the existing global mechanisms display problems regarding convergence during training. In this paper, we propose a novel gated global module (GGM) that is lightweight and yet effective in terms of achieving better integration of global information in relation to feature representation. Regarding the original structure of the network as a local block, our module infers global information in parallel with local information, and then a gate function is applied to generate global guidance which is applied to the output of the local module to capture representative information. The proposed GGM can be easily integrated with common CNN architectures and is training friendly. We used a classification task as an example to verify the effectiveness of the proposed GGM, and extensive experiments on ImageNet and CIFAR demonstrated that our method can be widely applied and is conducive to integrating global information into common networks. Liping Hou, Yuantao Song, Ke Lu 0002, Jian Xue 0002 |
ICMR | 4 |
| 2020 | YOLO-mini-tiger: Amur Tiger DetectionabstractIn this paper, we present our solution for tiger detection in the 2019 Computer Vision for Wildlife Conservation Challenge (CVWC2019). We introduce an efficient deep tiger detector, which consists of the convnet channel adaptation method and an improved tiger detection method based on You Only Look Once version 3 (YOLOv3). Considering the limited memory and computing power of tiny embedded devices, we have used EfficientNet-B0 and Darknet-53 as backbone networks for detection and adapted them to balance their depth and width inspired by the channel pruning method and knowledge distillation method. Our results show that after an architecture adjustment of Darknet-53, the floating-point computation decreases by 93%, its model size decreases by 97%, and its accuracy only decreases by 1%; after an architecture adjustment of EfficientNet-B0, the floating-point computation decreases by 66%, its model size decreases by 70% with its accuracy only decreased by 1%. We also compare GIoU loss and MSE loss in the training stage. The GIoU loss has the advantage that it increases the average AP for IoU from 0.5 to 0.95 without affecting training speed and the interface speed, so it is experimentally reasonable for tiger detection in the wild. This proposed method outperforms previous Amur tiger detection methods presented at CVWC2019. Runchen Wei, Ke Lu 0002 |
ICMR | 3 |
| 2020 | Extraction of Multi-class Multi-instance Geometric Primitives from Point Clouds Using Energy Minimization
Liang Wang 0021, Biying Yan, Fuqing Duan, Ke Lu 0002 |
MMM (2) | 4 |
| 2020 | Leveraging 3D blendshape for facial expression recognition using CNN
Sa Wang, Zhengxin Cheng, Xiaoming Deng 0001, Liang Chang 0001, Fuqing Duan, Ke Lu 0002 |
Sci. China Inf. Sci. | 6 |
| 2020 | Energy minimisation-based multi-class multi-instance geometric primitives extraction from 3D point cloudsabstractGeometric primitives contained in three‐dimensional (3D) point clouds can provide the meaningful and concise abstraction of 3D data, which plays a vital role in improving 3D vision‐based intelligent applications. However, how to efficiently and robustly extract multiple geometric primitives from point clouds is still a challenge, especially when multiple instances of multiple classes of geometric primitives are present. In this study, a novel energy minimisation‐based algorithm for multi‐class multi‐instance geometric primitives extraction from the 3D point cloud is proposed. First, an improved sampling strategy is proposed to generate model hypotheses. Then, an improved strategy to establish the neighbourhood is proposed to help construct and optimise an energy function for points labelling. After that, hypotheses and parameters of models are refined. Iterate this process until the energy does not decrease. Finally, models of multi‐class multi‐instance geometric primitives are simultaneously and robustly extracted from the 3D point cloud. In comparison with the state‐of‐the‐art methods, it can automatically determine the classes and numbers of geometric primitives in the 3D point cloud. Experimental results with synthetic and real data validate the proposed algorithm. Liang Wang 0021, Biying Yan, Fuqing Duan, Ke Lu 0002 |
IET Image Process. | 4 |
| 2020 | Editorial: Deep learning for medical image analysis
Ke Lu 0002, Fei Wang 0001, Ling Shao 0001, Weisheng Li 0001 |
Neurocomputing | 1 |
| 2020 | Adaptively weighted nonlocal means and TV minimization for speckle reduction in SAR images
Ruolin Wang, Yixue Wang, Ke Lu 0002 |
Multim. Tools Appl. | 4 |
| 2020 | Rotational-guided optimal cutting-plane extraction from point cloud
Yinghui Wang 0001, Ningna Wang, Xiaojuan Ning, Yanni Zhao, Ke Lu 0002 |
Multim. Tools Appl. | 7 |
| 2020 | SPN: short path network for scene text detection
Yuanqiang Cai, Weiqiang Wang 0001, Haiqing Ren, Ke Lu 0002 |
Neural Comput. Appl. | 4 |
| 2020 | In-air handwritten Chinese text recognition with temporal convolutional recurrent network
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
Pattern Recognit. | 3 |
| 2020 | Compressing the CNN architecture for in-air handwritten Chinese character recognition
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
Pattern Recognit. Lett. | 3 |
| 2019 | An Automated Lung Nodule Segmentation Method Based On Nodule Detection Network and Region GrowingabstractSegmentation of a specific organ or tissue plays an important role in medical image analysis with the rapid development of clinical decision support systems. With medical imaging equipments, segmenting the lung nodules in the images is able to help physicians diagnose lung cancer diseases and formulate proper schemes. Therefore the research of lung nodule segmentation has attracted a lot of attention these years. However, this task faces some challenges, including the intensity similarity between lung nodules and vessel, inaccurate boundaries and presence of noise in most of the images. In this paper, an automated segmentation method is proposed for lung nodules in CT images. At the first stage, a nodule detection network is used to generate region proposals and locate the bounding boxes of nodules, which are employed as the initial input for the following segmentation. Then the nodules are segmented in the bounding boxes at the second stage. Since the image scale for region growing is reduced by locating the nodule in advance, the efficiency of segmentation can be improved. And due to the localization of nodule before segmentation, some tissues with similar intensity can be excluded from the object region. The proposed method is evaluated on a public lung nodule dataset, and the experimental results indicate the effectiveness and efficiency of the proposed method. Yanhao Tan, Ke Lu 0002, Jian Xue 0002 |
MMAsia | 2 |
| 2019 | Dense Attention Network for Facial Expression Recognition in the WildabstractRecognizing facial expression is significant for human-computer interaction system and other applications. A certain number of facial expression datasets have been published in recent decades and helped with the improvements for emotion classification algorithms. However, recognition of the realistic expressions in the wild is still challenging because of uncontrolled lighting, brightness, pose, occlusion, etc. In this paper, we propose an attention mechanism based module which can help the network focus on the emotion-related locations. Furthermore, we produce two network structures named DenseCANet and DenseSANet by using the attention modules based on the backbone of DenseNet. Then these two networks and original DenseNet are trained on wild dataset AffectNet and lab-controlled dataset CK+. Experimental results show that the DenseSANet has improved the performance on both datasets comparing with the state-of-the-art methods. Ke Lu 0002, Jian Xue 0002, Yanfu Yan |
MMAsia | 2 |
| 2019 | A Single-stage Multi-class Object Detection Method for Remote Sensing ImagesabstractImpressive progresses have been achieved in object detection for images by convolution neural networks. However, a robust multi-class object detection method is still one of the great challenges for remote sensing images. Due to the great diversity of scale, orientation, density and background of objects, most advanced object detection algorithms in natural scenes usually suffer a sharp decline in remote sensing images detection. To solve these problems, we proposed a Single-stage Multi-class Object Detection (SMOD) method, aiming at remote sensing images, which can be trained from scratch and detect multi-class objects quickly and precisely. The proposed method introduces a novel Feature Reuse and Attention (FRA) structure as a key module of feature extraction backbone, which combines SE Attention module and dense Feature Reuse connection. Especially, a multiclass detection structure is proposed to learn from multi-scale, multi-level feature map and get effective attention representation for multi-class remote sensing object detection. SMOD can be trained from scratch without pre-trained network stably and converge well simply by employing batch normalization throughout the network. Experiments show that our trainingfrom-scratch method can obtain better performance compared with some state-of-art algorithms on public multi-class remote sensing dataset AIIA2018-6. Liping Hou, Jian Xue 0002, Ke Lu 0002, Mohammad Muntasir Rahman |
VCIP | 3 |
| 2019 | A new perspective: Recognizing online handwritten Chinese characters via 1-dimensional CNN
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
Inf. Sci. | 3 |
| 2019 | 3D object detection: Learning 3D bounding boxes from scaled down 2D bounding boxes in RGB-D images
Mohammad Muntasir Rahman, Yanhao Tan, Jian Xue 0002, Ling Shao 0001, Ke Lu 0002 |
Inf. Sci. | 5 |
| 2019 | N-FTRN: Neighborhoods based fully convolutional network for Chinese text line recognition
Hongzhu Li, Weiqiang Wang 0001, Ke Lu 0002 |
Multim. Tools Appl. | 3 |
| 2018 | Action Recognition With Coarse-to-Fine Deep Feature Integration and Asynchronous FusionabstractAction recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recognition accuracy by: 1) deriving more precise features for representing actions, and 2) reducing the asynchrony between different information streams. We first introduce a coarse-to-fine network which extracts shared deep features at different action class granularities and progressively integrates them to obtain a more accurate feature representation for input actions. We further introduce an asynchronous fusion network. It fuses information from different streams by asynchronously integrating stream-wise features at different time points, hence better leveraging the complementary information in different streams. Experimental results on action recognition benchmarks demonstrate that our approach achieves the state-of-the-art performance. Weiyao Lin, Ke Lu 0002, Bin Sheng 0001, Jianxin Wu 0001, Bingbing Ni, Hongkai Xiong |
AAAI | 3 |
| 2018 | A Unified CNN-RNN Approach for in-Air Handwritten English Word RecognitionabstractAs a new human-computer interaction application, in-air handwriting allows the user to write in the air in a natural way. In this paper, we propose a unified CNN-RNN approach for in-air handwritten English word recognition (IAHEWR), which integrates the advantages of both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). Specifically, the proposed approach follows an encoder-decoder framework, where the encoder is a deep CNN for efficiently processing the input temporal-sequential features, and the decoder is a RNN for accurately generating the target character sequence. We evaluate the proposed approach on an in-air handwritten English word dataset IAHEW-UCAS2016, and the experimental results demonstrate that the proposed approach achieves the comparable recognition accuracy and much higher computation efficiency when compared with the state-of-the-art approach for IAHEWR. Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
ICME | 3 |
| 2018 | A fast Cascade Shape Regression Method based on CNN-based InitializationabstractCascade shape regression (CSR) methods predict facial landmarks by iteratively updating an initial shape and are state-of-the-art. The initial shape always limits the result and causes local optimum, which is usually obtained from the average face or by randomly picking a face from the training set. In this paper, we propose a CNN-based initial method for CSR. Convolution neural network provides a highly robust initial shape estimation, while the following CSR algorithm fine-tunes the initialization rapidly to achieve higher accuracy. Furthermore, CNN-based initial approach is proposed to get 68-point initial shape, which is calculated from convolutional network 5-point result by the radial basis function interpolation with thin-plate splines (RBF-TPS). Extensive experiments demonstrate that CSR methods are sensitive to the initialization and proposed approach gets favorable results compared to state-of-the-art algorithms and achieves real-time performance. Jian Xue 0002, Ke Lu 0002, Yanfu Yan |
ICPR | 3 |
| 2018 | Group Re-Identification: Leveraging and Integrating Multi-Grain InformationabstractThis paper addresses an important yet less-studied problem: re-identifying groups of people in different camera views. Group re-identification (Re-ID) is very challenging since it is not only interfered by view-point and human pose variations in the traditional single-object Re-ID tasks, but also suffers from group layout and group member variations. To handle these issues, we propose to leverage the information of multi-grain objects: individual person and subgroups of two and three people inside a group image. We compute multi-grain representations to characterize the appearance and spatial features of multi-grain objects and evaluate the importance weight of each object for group Re-ID, so as to handle the interferences from group dynamics. We compute the optimal group-wise matching by using a multi-order matching process based on the multi-grain representation and importance weights. Furthermore, we dynamically update the importance weights according to the current matching results and then compute a new optimal group-wise matching. The two steps are iteratively conducted, yielding the final matching results.Experimental results on various datasets demonstrate the effectiveness of our approach. Weiyao Lin, Bin Sheng 0001, Ke Lu 0002, Junchi Yan, Jingdong Wang 0001, Errui Ding, Hongkai Xiong |
ACM Multimedia | 4 |
| 2018 | Unsupervised color image segmentation with color-alone feature using region growing pulse coupled neural network
Guangzhu Xu, Bang Jun Lei, Ke Lu 0002 |
Neurocomputing | 4 |
| 2018 | Accelerated nonrigid image registration using improved Levenberg-Marquardt method
Jiyang Dong, Ke Lu 0002, Jian Xue 0002, Shuangfeng Dai, Weiguo Pan |
Inf. Sci. | 2 |
| 2018 | Spatiotemporal text localization for videos
Yuanqiang Cai, Weiqiang Wang 0001, Shao Huang, Ke Lu 0002 |
Multim. Tools Appl. | 5 |
| 2018 | In-air handwritten Chinese character recognition with locality-sensitive sparse representation toward optimized prototype classifier
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou |
Pattern Recognit. | 3 |
| 2018 | Data augmentation and directional feature maps extraction for in-air handwritten Chinese character recognition based on convolutional neural network
Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou |
Pattern Recognit. Lett. | 3 |
| 2018 | Single image motion deblurring with reduced ringing effects using variational Bayesian estimation
Shan Cao 0002, Ke Lu 0002, Xiuling Zhou |
Signal Process. | 4 |
| 2018 | Single Image Dehazing Based on the Physical Model and MSRCR AlgorithmabstractTo address the hazy weather image degradation problem, we propose a single image dehazing method based on a physical model and the brightness components of the image by using a multi-scale retinex with color restoration algorithm. The overall dehazing process involves three components, including the atmospheric light value calculation, transmission map estimation, and recovery of the hazy image scene radiance. Our contribution is that we propose a novel algorithm to dehaze a single image by calculating the atmospheric light value and computing the transmission map while considering the dynamic range of the image. Experimental results show that our algorithm can effectively improve the image quality degraded by foggy weather and retain sufficient image details. Jinbao Wang 0001, Ke Lu 0002, Jian Xue 0002, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Scene text detection based on pruning strategy of MSER-trees and Linkage-treesabstractThe extraction and recognition of scene text in images is an important way to understand the semantic information in image. By now, scene text detection is still a challenging problem. In this paper, we present a scene text localization method based on the pruning of Maximally Stable Extremal Region (MSER) tree and Linkage-tree. Concretely, the MSER-tree is first constructed and overlap MSERs are removed by the nonmaximum suppression strategy. Further, the linkage-trees are constructed and pruned based on the features considering the nodes themselves and their siblings. Finally, the non-text lines are eliminated based on to the CNN-based features and intensity contrast. Experimental results on two challenging datasets demonstrate the effectiveness of the proposed method. Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou |
ICME | 3 |
| 2017 | RGB-D object recognition with multimodal deep convolutional neural networksabstractObject recognition from RGB-D images has become a hot topic and gained a significant popularity in recent years due to its numerous applications. In this paper, we propose a novel multimodal deep convolutional neural networks architecture for RGB-D object recognition which composed of three streams with two different types of deep CNNs, where each stream can separately learn from each modality. Finally, we propose a combined architecture of joint network of these three streams to classify the objects. Compared to RGB data, RGB-D images provide additional depth information that can be represented as depth colorization methods or surface normals. Our goal is to exploit both colorization and surface normals information to encode depth images. We show that by utilizing both colorization and surface normals of depth images combined with RGB significantly can improves the classification accuracy. We evaluate our model on one of the most challenging RGB-D object dataset and achieves comparable performance to state-of-the-art methods. Mohammad Muntasir Rahman, Yanhao Tan, Jian Xue 0002, Ke Lu 0002 |
ICME | 4 |
| 2017 | An end-to-end recognizer for in-air handwritten Chinese characters based on a new recurrent neural networksabstractIn-air handwriting is becoming a new human-computer interaction way. It is a challenging task to accurately recognizing in-air handwritten Chinese characters. In this paper, we present an end-to-end recognizer for in-air handwritten Chinese characters by using recurrent neural networks (RNN). Compared with the existing methods, the proposed RNN based methods does not need to explicitly extract features and directly take a sequence of dot locations as input. We have made two aspects of modifications on traditional RNN for improving the recognition accuracy. Concretely, the sum-pooling is performed on the states of each hidden layers, and a faster convergence in training can be obtained. Additionally, an assistant objective function is introduced into the conventional loss function, which brings a slight increase of performance. To evaluate the performance of the proposed method, the experiments are carried out on the IAHCC-UCAS2016 datasets to compare ours with other state-of-art methods. The experimental results show that the proposed RNN model has a fairly high recognition accuracy for in-air handwritten Chinese characters. Haiqing Ren, Weiqiang Wang 0001, Ke Lu 0002, Jianshe Zhou, Qiuchen Yuan |
ICME | 3 |
| 2017 | Convex-relaxed active contour model based on localised kernel mappingabstractIntensity inhomogeneity is one of the major obstacles for intensity‐based segmentation in many applications. The recently proposed kernel mapping (KM) method has exhibited excellent performance on segmenting various types of noisy images while it is not effective to handle intensity inhomogeneity. To overcome this drawback, this study presents a localised KM (LKM) method based on the fact that intensity inhomogeneity can be ignored in a local neighbourhood. The authors’ method first reconstructs the KM formulation of image segmentation in a neighbourhood of each pixel, and then such formulations for all pixels can be integrated together to derive the LKM energy functional. Minimisation of the energy functional is implemented by solving an equivalent convex‐relaxed problem whose optimisation can be quickly achieved via the split Bregman method. Experimental results on two‐phase segmentation and multiphase segmentation demonstrate competitive performance of the LKM method in the presence of intensity inhomogeneity and severe noise. Wenchao Cui, Guoqiang Gong, Ke Lu 0002, Shuifa Sun, Fangmin Dong |
IET Image Process. | 3 |
| 2017 | A 3D polar-radius-moment invariant as a shape circularity measure
Ziping Ma 0001, Jinlin Ma, Bin Xiao 0002, Ke Lu 0002 |
Neurocomputing | 4 |
| 2017 | Performance evaluation of deep feature learning for RGB-D image/video classification
Ling Shao 0001, Ziyun Cai, Li Liu 0004, Ke Lu 0002 |
Inf. Sci. | 4 |
| 2017 | Maximum data delivery probability-oriented routing protocol in opportunistic mobile networks
Huan Zhou 0002, Linping Tong, Tingyao Jiang, Shouzhi Xu, Jialu Fan, Ke Lu 0002 |
Peer-to-Peer Netw. Appl. | 6 |
| 2017 | Learning Correspondence Structures for Person Re-IdentificationabstractThis paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure, which indicates the patchwise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global constraint-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Finally, we also extend our approach by introducing a multi-structure scheme, which learns a set of local correspondence structures to capture the spatial correspondence sub-patterns between a camera pair, so as to handle the spatial misalignments between individual images in a more precise way. Experimental results on various data sets demonstrate the effectiveness of our approach. Weiyao Lin, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Jingdong Wang 0001, Ke Lu 0002 |
IEEE Trans. Image Process. | 7 |
| 2017 | G-IK-SVD: parallel IK-SVD on GPUs for sparse representation of spatial big data
Weijing Song, Ze Deng, Lizhe Wang 0001, Bo Du 0006, Peng Liu 0024, Ke Lu 0002 |
J. Supercomput. | 6 |
| 2016 | MQDF with a novel covariance matrix estimation and discriminant LSRC, which is better for in-air handwritten Chinese character recognitionabstractWith the advance of 3-dimensional sensing devices, the in-air handwriting, as a more natural way for human and computer interaction, is being developed by the UCAS-CVMT Lab. Compared with the conventional handwritten Chinese characters generated by touching, it is more challenging to accurately recognize them due to unconstrained one-stroke writing style. This paper presents two recognizers to address this problem. One is built on the modified quadratic discriminant function (MQDF), where a new kernel method is used to estimate the covariance matrix. Additionally, we present a new classifier built on the locality-sensitive sparse representation based classifier (LSRC). It introduces the discriminant information to reinforce the recognition ability, called discriminant LSRC. We have applied them on the in-air handwritten Chinese character dataset IAHCC-UCAS2015. The experimental results show that the first one has the higher accuracy but huge storage cost, while the DLSRC recognizer has the comparable accuracy and very compact storage space. Weiqiang Wang 0001, Ke Lu 0002 |
ICIP | 3 |
| 2016 | High-order directional features and sparse representation based classification for in-air handwritten Chinese character recognitionabstractThe in-air handwriting is a natural and promising human-computer interaction way. Compared with handwritten Chinese characters on touch screen, the in-air handwritten Chinese characters have their unique characteristics, e.g., each character is always written in a single stroke. In this paper, we propose a high-order directional feature for recognizing in-air handwritten Chinese characters. The proposed highorder features characterize the rate of direction change and the change rate of the rate of direction change, and can be easily be combined with the 8-directional features to get an enhanced version. Additionally, we exploit the locality-sensitive dictionary learning and sparse representation based classifier (LSRC) to recognize in-air handwritten Chinese characters. Since few work has applied the SRC in on-line handwritten Chinese character recognition (OHCCR), we evaluate the proposed system on both an in-air handwritten Chinese character dataset, the IAHCC-UCAS2015 dataset, and a handwritten Chinese character dataset, the SCUT-COUCH2009 database. The experimental results show that the proposed high-order directional features, when combined with the firstorder directional features (8-directional features), can improve the recognition accuracy, and they are more suitable for IAHCCR than traditional OHCCR. Additionally, the results also demonstrate the LSRC is a good choice for IAHCCR. Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Zhangjian Ji |
ICME | 3 |
| 2016 | Egocentric hand detection via region growthabstractWearable cameras used to record daily life are attracting researchers' attention, and a large number of ego-related applications have been developed in recent years. Hand detection is one of the key steps for the tasks like gesture recognition, action recognition and understanding hand-based interaction in egocentric videos, since humans are accustomed to interacting with objects using their hands. In this work a novel region growth approach is proposed for egocentric hand detection. The proposed method first identifies seed regions most likely containing the parts of hand based on the distribution of hand-related matches. Then hand regions are gradually located by extending from the seed regions. Finally a whole hand is obtained according to the scores of adjacent superpixels, which are evaluated based on four egocentric cues: contrast, distance, previous hand position and skin appearance. The experimental results on two publicly available datasets demonstrate this work achieves satisfactory performance. Shao Huang, Weiqiang Wang 0001, Ke Lu 0002 |
ICPR | 3 |
| 2016 | In-air handwritten Chinese character recognition using discriminative projection based on locality-sensitive sparse representationabstractDimensionality reduction methods have been shown to be effective for handwritten Chinese character recognition. In this paper, we propose discriminative projection based on locality-sensitive sparse representation (DPLSR) for in-air handwritten Chinese character recognition. DPLSR based on the locality-sensitive sparse representation based classifier (LSRC), which can provide closed-form solutions and maintain the data locality constraint during the sparse coding stage. In contrast to sparse representation classifier steered discriminative projection (SRC-DP), which did not consider global structure of data and use all training samples as dictionary atoms, DPLSR is able to use fewer atoms and spend less training time to achieve better performance. Experiments are conducted on the IAHCC-UCAS2016 dataset built by us, experimental results demonstrate the effectiveness of proposed method. Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002 |
ICPR | 3 |
| 2016 | Collaborative Q-Learning Based Routing Control in Unstructured P2P Networks
Xiangjun Shen, Jianping Gou, Qirong Mao, Zhengjun Zha, Ke Lu 0002 |
MMM (1) | 6 |
| 2016 | Manifold regularized kernel logistic regression for web image annotation
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002 |
Neurocomputing | 5 |
| 2016 | An overview of multi-modal medical image fusion
Jiao Du, Weisheng Li 0001, Ke Lu 0002, Bin Xiao 0002 |
Neurocomputing | 3 |
| 2016 | Convex optimization based low-rank matrix decomposition for image restoration
Jinbao Wang 0001, Lulu Zhang 0006, Ke Lu 0002 |
Neurocomputing | 4 |
| 2016 | Stereo data sensing, computation and perception
Ke Lu 0002, Yi Zhen |
Neurocomputing | 1 |
| 2016 | A gradient descent boosting spectrum modeling method based on back interval partial least squares
Fangfang Qu, Ke Lu 0002, Xiangyu Wang 0001 |
Neurocomputing | 3 |
| 2016 | GPU-based real-time terrain rendering: Design and implementation
Ke Lu 0002, Weiguo Pan, Shuangfeng Dai |
Neurocomputing | 2 |
| 2016 | Efficient volume rendering methods for out-of-Core datasets by semi-adaptive partitioning
Jian Xue 0002, Ke Lu 0002, Ling Shao 0001, Mohammad Muntasir Rahman |
Inf. Sci. | 3 |
| 2016 | Nonlocal Low-Rank-Based Compressed Sensing for Remote Sensing Image ReconstructionabstractRemote sensing image reconstruction from undersampled data is very much required by the onboard imaging system to cut down data volume and maintain image quality. Nonlocal low-rank regularization deriving from group sparsity, low rank, singular-value thresholding, and nonconvex surrogate functions have recently emerged for image recovery. To use nonlocal low-rank compressed sensing for remote sensing image reconstruction, spectral and temporal redundancy are considered in this letter by utilizing the similarity of correlated bands or historical records. Prior structural knowledge helps to group nonlocal similar blocks more accurately. Oversmoothness of low-rank regularization is improved by injecting referenced structures selectively. The proposed compressed sensing method is tested on satellite images from MODIS, LandSat-7, LandSat-8, IKONOS, and Google Earth to make clear that it outweighs state-of-the-art methods in maintaining fidelity and high visual details. Jingbo Wei, Ke Lu 0002, Lizhe Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Compressed sensing based remote sensing image reconstruction via employing similarities of reference images
Lizhe Wang 0001, Peng Liu 0024, Ke Lu 0002, Dingsheng Liu |
Multim. Tools Appl. | 4 |
| 2016 | Non-local sparse regularization model with application to image denoising
Jinbao Wang 0001, Lulu Zhang 0006, Guang-Mei Xu, Ke Lu 0002 |
Multim. Tools Appl. | 5 |
| 2016 | Large-scale paralleled sparse principal component analysis
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002 |
Multim. Tools Appl. | 5 |
| 2016 | Energy-based automatic recognition of multiple spheres in three-dimensional point cloud
Liang Wang 0021, Chao Shen 0002, Fuqing Duan, Ke Lu 0002 |
Pattern Recognit. Lett. | 4 |
| 2016 | Learning Spatio-Temporal Representations for Action Recognition: A Genetic Programming ApproachabstractExtracting discriminative and robust features from video sequences is the first and most critical step in human action recognition. In this paper, instead of using handcrafted features, we automatically learn spatio-temporal motion features for action recognition. This is achieved via an evolutionary method, i.e., genetic programming (GP), which evolves the motion feature descriptor on a population of primitive 3D operators (e.g., 3D-Gabor and wavelet). In this way, the scale and shift invariant features can be effectively extracted from both color and optical flow sequences. We intend to learn data adaptive descriptors for different datasets with multiple layers, which makes fully use of the knowledge to mimic the physical structure of the human visual cortex for action recognition and simultaneously reduce the GP searching space to effectively accelerate the convergence of optimal solutions. In our evolutionary architecture, the average cross-validation classification error, which is calculated by an support-vector-machine classifier on the training set, is adopted as the evaluation criterion for the GP fitness function. After the entire evolution procedure finishes, the best-so-far solution selected by GP is regarded as the (near-)optimal action descriptor obtained. The GP-evolving feature extraction method is evaluated on four popular action datasets, namely KTH, HMDB51, UCF YouTube, and Hollywood2. Experimental results show that our method significantly outperforms other types of features, either hand-designed or machine-learned. Li Liu 0004, Ling Shao 0001, Xuelong Li 0001, Ke Lu 0002 |
IEEE Trans. Cybern. | 4 |
| 2016 | A Local Structural Descriptor for Image Matching via Normalized Graph Laplacian EmbeddingabstractThis paper investigates graph spectral approaches to the problem of point pattern matching. Specifically, we concentrate on the issue of how to effectively use graph spectral properties to characterize point patterns in the presence of positional jitter and outliers. A novel local spectral descriptor is proposed to represent the attribute domain of feature points. For a point in a given point-set, weight graphs are constructed on its neighboring points and then their normalized Laplacian matrices are computed. According to the known spectral radius of the normalized Laplacian matrix, the distribution of the eigenvalues of these normalized Laplacian matrices is summarized as a histogram to form a descriptor. The proposed spectral descriptor is finally combined with the approximate distance order for recovering correspondences between point-sets. Extensive experiments demonstrate the effectiveness of the proposed approach and its superiority to the existing methods. Jun Tang 0007, Ling Shao 0001, Xuelong Li 0001, Ke Lu 0002 |
IEEE Trans. Cybern. | 4 |
| 2016 | Detecting Salient Objects via Color and Texture Compactness HypothesesabstractIn recent years, the object-level saliency detection has attracted much research attention, due to its usefulness in many high-level tasks. Existing methods are mostly based on the contrast hypothesis, which regards the regions with high contrast in a certain context as salient objects. Although the contrast hypothesis is effective in many scenarios, it cannot handle some difficult cases. As a remedy to address the weakness of contrast hypothesis, we propose a novel compactness hypothesis, which assumes salient regions are more compact than background from the perspectives of both color layout and texture layout. Based on the compactness hypotheses, we implement an effective object-level saliency detection method. In the proposed method, we first construct a weak saliency map based on the compact hypotheses, then collect samples from the weak saliency map to train a dedicated classifier. This classifier is applied on each individual pixel of the input image to produce a confidence score. Finally, the confidence scores are used to form a saliency map. This process is carried out at different scales, and the corresponding results are integrated into the formation of the final saliency map. The proposed approach is evaluated on eight benchmark data sets, where it delivers the competitive performance compared with the state-of-the-art methods. Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002 |
IEEE Trans. Image Process. | 4 |
| 2015 | A novel binarization approach for text in imagesabstractAccurate recognition of scene text and overlaid text is still a challenging issue due to degradation and complex background, and text binarization is crucial for recognition accuracy. This paper presents an effective method to extract characters in images and video frames. Our method assumes that background pixels possess good spatial connectivity and high appearance similarity to boundary pixels in a cropped text string image. It first computes the confidence of pixels as text. Then the confidence map is exploited to partition text regions into characters. Further, each character region is clustered into different layers and background components are removed to generate candidate binarization results. The final result is obtained based on the scores of each layer. Our method is validated by better recognition rates and segmentation accuracy on the ICADR03 dataset and a big dataset of overlaid text. Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002 |
ICIP | 3 |
| 2015 | Robustly tracking objects via multi-task kernel dynamic sparse modelabstractRecently, sparse representation has been successfully applied by some generative tracking methods. However, few methods consider the correlation between the representations of each particle in time domain and space domain. Additionally, most methods use the raw pixels as templates which can not well adapt to the sophisticated object changes. To solve these problems, we consider the sparse representation in kernel space and propose a multi-task kernel dynamic sparse tracking algorithm (MTKDST). As compared to previous methods, our method exploits the dependencies between particles in the space domain and the correlation of particle representation in the time domain to improve the tracking performance. Furthermore, we also adopt the multikernel fusion mechanism to utilize multiple complementary visual features (e.g., spatial color histogram and spatial gradient-orientation histogram) to enhance the robustness of the proposed method. The comprehensive experiments on several challenging image sequences demonstrate that the proposed method outperforms the state-of-the-art approaches in tracking accuracy. Zhangjian Ji, Weiqiang Wang 0001, Ke Lu 0002 |
ICIP | 3 |
| 2015 | In-air handwritten Chinese character recognition using multi-stage classifier based on adaptive discriminative locality alignmentabstractThe in-air handwriting is a natural and useful way for human-computer interaction. Yet, to our knowledge, few work has been done for the in-air handwritten Chinese character recognition (IAHCCR). In this paper, we present a multi-stage recognizer to address the problem of IAHCCR. The proposed methods can also deal with the classical handwritten Chinese character recognition (HCCR). We find that the discriminative locality alignment (DLA) technique in HCCR heavily depends on the choice of parameters in practice. To overcome the disadvantage, we present an adaptive discriminative locality alignment (ADLA), which does not involve the parameter optimization process. At the same time, a new static similar characters collection technique is proposed. We evaluate the proposed methods on the IAHCC-UCAS2014 dataset, an in-air handwritten Chinese character dataset constructed by us, as well as the SCUT-COUCH2009 database, a HCCR dataset. The experimental results demonstrate the effectiveness of the proposed methods on two different kinds of dataset. Xiwen Qu, Weiqiang Wang 0001, Ke Lu 0002, Ning Xu 0008 |
ICIP | 3 |
| 2015 | Large visual words for large scale image classificationabstractRecently, using large visual vocabulary or codebooks to quantize and partition the set of local feature descriptors into large set of disjoint subsets termed visual words (or large visual words) has become an important research topic in solving many computer vision problems including near duplicate image retrieval, object retrieval, etc. Generally, large visual words means a heavy burden on the cost of time and memory space for both the construction of large vocabulary and the searching process, especially for large scale applications. In this paper, we present an efficient generation approach of large visual words with a very compact vocabulary, namely two dictionaries learned with sparse non-negative matrix factorization (NMF). After piecewise sparse decomposition of features with two learned dictionaries, we map a pair of indices of the dictionary's bases corresponding to the maximum elements of the two sparse codes to a large set of visual words upon the assumption that data with similar properties will share the same base with the largest sparse coefficient. With the help of an inverted file structure built through the large visual words, K-nearest neighbors (KNN) can be efficiently retrieved. Therefore, we can classify images very efficiently with the incorporation of our fast KNN search based on large visual words into SVM-KNN method. Experiments on the public Oxford dataset, and ACM Multimedia 2013 Yahoo! image classification challenge dataset show that our approach is both effective and efficient. Sheng Tang, Ke Lu 0002, Yongdong Zhang 0001 |
ICIP | 3 |
| 2015 | Detecting Salient Objects via Spatial and Appearance Compactness HypothesesabstractObject-level saliency detection has been attracting a lot of attention, due to its potential enhancement in many high-level vision tasks. Many previous methods are based on the contrast hypothesis which regards the regions with high contrast in a certain context as salient. Although the contrast hypothesis is valid in many cases, it cannot handle some difficult cases. To make up for the weakness of contrast hypothesis, we propose a novel compactness hypothesis which assumes salient regions are more compact than background spatially and in appearance. Based on compactness hypotheses, we implement an effective object-level saliency detection method, which is demonstrated to be effective even in difficult cases. In addition, we present an adaptive multiple saliency maps fusion framework which can automatically select saliency maps of high quality according to three quality assessment rules. We evaluate the proposed method on four benchmark datasets and the comparable performance as the state-of-the-art methods has been achieved. Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002 |
ACM Multimedia | 3 |
| 2015 | A novel approach of lung segmentation on chest CT images using graph cuts
Shuangfeng Dai, Ke Lu 0002, Jiyang Dong |
Neurocomputing | 2 |
| 2015 | 3D face reconstruction from skull by regression modeling in shape parameter spaces
Fuqing Duan, Donghua Huang, Yun Tian 0002, Ke Lu 0002, Zhongke Wu |
Neurocomputing | 4 |
| 2015 | An adaptively weighted algorithm for camera calibration with 1D objects
Liang Wang 0021, Fuqing Duan, Ke Lu 0002 |
Neurocomputing | 3 |
| 2015 | Single image dehazing with a physical model and dark channel prior
Jinbao Wang 0001, Lulu Zhang 0006, Ke Lu 0002 |
Neurocomputing | 4 |
| 2015 | 3D orbit selection for regional observation GEO SAR
Yangte Gao, Jiaoshan Li, Ke Lu 0002 |
Neurocomputing | 5 |
| 2015 | Weakly-supervised scene parsing with multiple contextual cues
Teng Li 0001, Xinyu Wu 0001, Bingbing Ni, Ke Lu 0002, Shuicheng Yan |
Inf. Sci. | 4 |
| 2015 | Hybrid architecture for 3D visualization of ultrasonic data
Weiguo Pan, Jian Xue 0002, Ke Lu 0002, Shuangfeng Dai |
Inf. Sci. | 3 |
| 2015 | Particle Swarm Optimization based dictionary learning for remote sensing big data
Lizhe Wang 0001, Hao Geng, Peng Liu 0024, Ke Lu 0002, Joanna Kolodziej, Rajiv Ranjan 0001, Albert Y. Zomaya |
Knowl. Based Syst. | 4 |
| 2015 | Compressed Sensing of a Remote Sensing Image Based on the Priors of the Reference ImageabstractBasic compressed-sensing algorithms for image reconstructions mainly deal with the computation of sparse regularization. Remote sensing applications often have multisource or multitemporal images whose different components are acquired separately. Therefore, this letter considers the reconstruction of a remote sensing image using an auxiliary image from another sensor or another time as the reference. For this application, a new compressed-sensing object function is developed that uses a reference image as a prior. In the new model, the sparsity constraints in the transform domain come from the target image, and the gradient priors in the spatial domain come from the auxiliary reference image. The hybrid regularization is optimized by basing the algorithm on the Bregman split method. The proposed method shows better performances when compared with other three popular compressed-sensing algorithms. Lizhe Wang 0001, Ke Lu 0002, Peng Liu 0024 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Nonlocal variational image segmentation models on graphs using the Split Bregman
Ke Lu 0002, Daru Pan, Weiguo Pan |
Multim. Syst. | 1 |
| 2015 | Transmission of multimedia contents in opportunistic networks with social selfish nodes
Daru Pan, WeiJing Chen, Ke Lu 0002 |
Multim. Syst. | 4 |
| 2015 | Guest editorial: selected papers from ICIMCS 2013
Meng Wang 0001, Ke Lu 0002, Gang Hua 0001, Cees Snoek |
Multim. Syst. | 2 |
| 2015 | An efficient multi-threshold AdaBoost approach to detecting faces in images
Weiqiang Wang 0001, Ke Lu 0002 |
Multim. Tools Appl. | 4 |
| 2015 | Multiview Hessian regularized logistic regression for action recognition
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002 |
Signal Process. | 5 |
| 2015 | An improved fractional-order differentiation model for image denoising
Jinbao Wang 0001, Lulu Zhang 0006, Ke Lu 0002 |
Signal Process. | 4 |
| 2015 | Toward Naturalistic 2D-to-3D ConversionabstractNatural scene statistics (NSSs) models have been developed that make it possible to impose useful perceptually relevant priors on the luminance, colors, and depth maps of natural scenes. We show that these models can be used to develop 3D content creation algorithms that can convert monocular 2D videos into statistically natural 3D-viewable videos. First, accurate depth information on key frames is obtained via human annotation. Then, both forward and backward motion vectors are estimated and compared to decide the initial depth values, and a compensation process is applied to further improve the depth initialization. Then, the luminance/chrominance and initial depth map are decomposed by a Gabor filter bank. Each subband of depth is modeled to produce a NSS prior term. The statistical color-depth priors are combined with the spatial smoothness constraint in the depth propagation target function as a prior regularizing term. The final depth map associated with each frame of the input 2D video is optimized by minimizing the target function over all subbands. In the end, stereoscopic frames are rendered from the color frames and their associated depth maps. We evaluated the quality of the generated 3D videos using both subjective and objective quality assessment methods. The experimental results obtained on various sequences show that the presented method outperforms several state-of-the-art 2D-to-3D conversion methods. Xun Cao, Ke Lu 0002, Qionghai Dai, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2015 | Learning View-Model Joint Relevance for 3D Object Retrievalabstract3D object retrieval has attracted extensive research efforts and become an important task in recent years. It is noted that how to measure the relevance between 3D objects is still a difficult issue. Most of the existing methods employ just the model-based or view-based approaches, which may lead to incomplete information for 3D object representation. In this paper, we propose to jointly learn the view-model relevance among 3D objects for retrieval, in which the 3D objects are formulated in different graph structures. With the view information, the multiple views of 3D objects are employed to formulate the 3D object relationship in an object hypergraph structure. With the model data, the model-based features are extracted to construct an object graph to describe the relationship among the 3D objects. The learning on the two graphs is conducted to estimate the relevance among the 3D objects, in which the view/model graph weights can be also optimized in the learning process. This is the first work to jointly explore the view-based and model-based relevance among the 3D objects in a graph-based framework. The proposed method has been evaluated in three data sets. The experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness on retrieval accuracy of the proposed 3D object retrieval method. Ke Lu 0002, Jian Xue 0002, Jiyang Dong, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | A Real-Time Hand Posture Recognition System Using Deep Neural NetworksabstractHand posture recognition (HPR) is quite a challenging task, due to both the difficulty in detecting and tracking hands with normal cameras and the limitations of traditional manually selected features. In this article, we propose a two-stage HPR system for Sign Language Recognition using a Kinect sensor. In the first stage, we propose an effective algorithm to implement hand detection and tracking. The algorithm incorporates both color and depth information, without specific requirements on uniform-colored or stable background. It can handle the situations in which hands are very close to other parts of the body or hands are not the nearest objects to the camera and allows for occlusion of hands caused by faces or other hands. In the second stage, we apply deep neural networks (DNNs) to automatically learn features from hand posture images that are insensitive to movement, scaling, and rotation. Experiments verify that the proposed system works quickly and accurately and achieves a recognition accuracy as high as 98.12%. Ao Tang, Ke Lu 0002, Jie Huang 0011, Houqiang Li |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2015 | Fashion Parsing With Video ContextabstractIn this paper, we propose a novel semi- supervised learning strategy to address human parsing. Existing human parsing datasets are relatively small due to the required tedious human labeling. We present a general, affordable and scalable solution, which harnesses the rich contexts in those easily available web videos to boost any existing human parser. First, we crawl a large number of unlabeled videos from the web. Then for each video, the cross-frame contexts are utilized for human pose co- estimation , and then video co-parsing to obtain satisfactory human parsing results for all frames. More specifically, SIFT flow and super-pixel matching are used to build correspondences across different frames, and these correspondences then contextualize the pose estimation and human parsing in individual frames. Finally these parsed video frames are used as the reference corpus for the non-parametric human parsing component of the whole solution. To further improve the accuracy of video co-parsing, we propose an active learning method to incorporate human guidance, where the labelers are required to assess the accuracies of the pose estimation results of certain selected video frames. Then we take reliable frames as the seed frames to guide the video pose co-estimation. Our human parsing framework can then easily incorporate the human feedback to train a better fashion parser. Extensive experiments on two benchmark fashion datasets as well as a newly collected challenging Fashion Icon dataset well demonstrate the encouraging performance gain from our general pipeline for human parsing. Si Liu 0001, Xiaodan Liang, Luoqi Liu, Ke Lu 0002, Liang Lin 0004, Xiaochun Cao, Shuicheng Yan |
IEEE Trans. Multim. | 4 |
| 2014 | Hamming embedding with fragile bits for image searchabstractRecently, several binary descriptors are proposed, which represent interest points in image using binary codes. In these binary feature schemes, two descriptors are considered as a match, if the Hamming distance between them is below a threshold. Applying Hamming distance to measure the similarity between binary descriptors can extremely promote the computational efficiency. However, our experimental results presents that there exists a large number of bits in the binary feature vector cannot maintain the robustness while image conditions change. Rather than ignore the impacts of those unstable bits, we take into account the difference of robustness among the feature bits and propose a novel similarity measurement, which called the Fragile Bit Ratio (FBR). FBR is used in binary feature matching to measure how two features differ. High FBRs are associated with genuine matches between two binary features and low FBRs are associated with impostor ones. Based on this metric, we propose a new binary feature matching scheme to fuse the Hamming distance and Fragile Bit Ratio. In our approach, we match the descriptors using the Hamming distance threshold roughly, and then filtered by the Fragile Bits Ratio to refine the candidate set. In experiments, using Fragile Bits Radio can effectively remove the false matches and highly improve the accuracy of image search. Furthermore, our method can easily be integrated into the other well-established binary features schemes. Dongye Zhuang, Dongming Zhang 0004, Jintao Li 0001, Ke Lu 0002, Qi Tian 0001 |
ICIP | 4 |
| 2014 | Video Text Extraction Using the Fusion of Color Gradient and Log-Gabor FilterabstractVideo text which contains rich semantic information can be utilized for video indexing and summarization. However, compared with scanned documents, text recogniton for video text is still a challenging problem due to complex background. Segmenting text line into single characters before text extraction can achieve higher recognition accuracy, since background of single character is less complex compared with whole text line. Therefore, we first perform character segmentation, which can accurately locate the character gap in the text line. More specifically, we get a fusion map which fuses the results of color gradient and log-gabor filter. Then, candidate segmentation points are obtained by vertical projection analysis of the fusion map. We get segmentation points by finding minimum projection value of candidate points in a limited range. Finally, we get the binary image of the single character image by applying K-means clustering and combine their results to form binary image of the whole text line. The binary image is further refined by inward filling and the fusion map. The experimental results on a large amount of data show that the proposed method can contribute to better binarization result which leads to a higher character recognition rate of OCR engine. Zhike Zhang, Weiqiang Wang 0001, Ke Lu 0002 |
ICPR | 3 |
| 2014 | Fashion Parsing with Video ContextabstractIn this paper, we explore how to utilize the video context to facilitate fashion parsing. Instead of annotating a large amount of fashion images, we present a general, affordable and scalable solution, which harnesses the rich contexts in easily available fashion videos to boost any existing fashion parser. First, we crawl a large unlabelled fashion video corpus with fashion frames. Then for each fashion video, the cross-frame contexts are utilized for human pose co-estimation, and then video co-parsing to obtain satisfactory fashion parsing results for all frames. More specifically, Sift Flow and super-pixel matching are used to build correspondences across frames, and these correspondences then con- textualize the pose estimations and fashion parsing in individual frames. Finally, these parsed video frames are used as the reference corpus for the non-parametric fashion parsing component of the whole solution. Extensive experiments on two benchmark fashion datasets as well as a newly collected challenging Fashion Icon (FI) dataset demonstrate the encouraging performance gain from our general pipeline for fashion parsing. Si Liu 0001, Xiaodan Liang, Luoqi Liu, Ke Lu 0002, Liang Lin 0004, Shuicheng Yan |
ACM Multimedia | 4 |
| 2014 | A graph distance based metric for data oriented workflow retrieval with variable time constraints
Yinglong Ma 0001, Ke Lu 0002 |
Expert Syst. Appl. | 3 |
| 2014 | Robust object removal with an exemplar-based image inpainting approach
Jing Wang 0093, Ke Lu 0002, Daru Pan, Bing-Kun Bao |
Neurocomputing | 2 |
| 2014 | Single-image motion deblurring using an adaptive image prior
Ke Lu 0002, Bing-Kun Bao, Lulu Zhang 0006, Jinbao Wang 0001 |
Inf. Sci. | 2 |
| 2014 | 3D model retrieval and classification by semi-supervised learning with content-based similarity
Ke Lu 0002, Jian Xue 0002, Weiguo Pan |
Inf. Sci. | 1 |
| 2014 | Skull Identification via Correlation Measure Between Skull and Face ShapeabstractSkull identification is an important subject for research in forensic medicine. Current research can be divided into two categories: 1) craniofacial superimposition and 2) craniofacial reconstruction. Both categories rely essentially on the accurate extraction and representation of the intrinsic relationship between the skull and face in terms of the morphology, which still remain unsolved. They have high uncertainty and a low identification capability. This paper proposes a novel skull identification method that matches an unknown skull with enrolled 3D faces, in which the mapping between the skull and face is obtained using canonical correlation analysis. Unlike existing techniques, this method needs no accurate relationship between the skull and face, and measures only the correlation between them. In order to measure the correlation more reliably and improve the identification capability of the correlation analysis model, a region fusion strategy is adopted. Experimental results validate the proposed method, and show that the region-based method can significantly boost the matching accuracy. The correct identification rate reaches 94% when using a CT data set. This paper can provide a theory support for research on craniofacial superimposition and craniofacial reconstruction. Fuqing Duan, Yan Li 0121, Yun Tian 0002, Ke Lu 0002, Zhongke Wu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2014 | Learning-Based Bipartite Graph Matching for View-Based 3D Model RetrievalabstractDistance measure between two sets of views is one central task in view-based 3D model retrieval. In this paper, we introduce a distance metric learning method for bipartite graph matching-based 3D object retrieval framework. In this method, the relationship among 3D models is formulated by a graph structure with semisupervised learning to estimate the model relevance. More specially, we model two sets of views by using a bipartite graph, on which their optimal matching is estimated. Then, we learn a refined distance metric by using the user’s relevance feedback. The proposed method has been evaluated on four data sets and the experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed method. Ke Lu 0002, Rongrong Ji, Jinhui Tang 0001, Yue Gao 0002 |
IEEE Trans. Image Process. | 1 |
| 2014 | A Graph Derivation Based Approach for Measuring and Comparing Structural Semantics of OntologiesabstractOntology reuse offers great benefits by measuring and comparing ontologies. However, the state of art approaches for measuring ontologies neglects the problems of both the polymorphism of ontology representation and the addition of implicit semantic knowledge. One way to tackle these problems is to devise a mechanism for ontology measurement that is stable, the basic criteria for automatic measurement. In this paper, we present a graph derivation representation based approach (GDR) for stable semantic measurement, which captures structural semantics of ontologies and addresses those problems that cause unstable measurement of ontologies. This paper makes three original contributions. First, we introduce and define the concept of semantic measurement and the concept of stable measurement. We present the GDR based approach, a three-phase process to transform an ontology to its GDR. Second, we formally analyze important properties of GDRs based on which stable semantic measurement and comparison can be achieved successfully. Third but not the least, we compare our GDR based approach with existing graph based methods using a dozen real world exemplar ontologies. Our experimental comparison is conducted based on nine ontology measurement entities and distance metric, which stably compares the similarity of two ontologies in terms of their GDRs. Yinglong Ma 0001, Ling Liu 0001, Ke Lu 0002, Beihong Jin, Xiangjie Liu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Representative Discovery of Structure Cues for Weakly-Supervised Image SegmentationabstractWeakly-supervised image segmentation is a challenging problem with multidisciplinary applications in multimedia content analysis and beyond. It aims to segment an image by leveraging its image-level semantics (i.e., tags). This paper presents a weakly-supervised image segmentation algorithm that learns the distribution of spatially structural superpixel sets from image-level labels. More specifically, we first extract graphlets from a given image, which are small-sized graphs consisting of superpixels and encapsulating their spatial structure. Then, an efficient manifold embedding algorithm is proposed to transfer labels from training images into graphlets. It is further observed that there are numerous redundant graphlets that are not discriminative to semantic categories, which are abandoned by a graphlet selection scheme as they make no contribution to the subsequent segmentation. Thereafter, we use a Gaussian mixture model (GMM) to learn the distribution of the selected post-embedding graphlets (i.e., vectors output from the graphlet embedding). Finally, we propose an image segmentation algorithm, termed representative graphlet cut, which leverages the learned GMM prior to measure the structure homogeneity of a test image. Experimental results show that the proposed approach outperforms state-of-the-art weakly-supervised image segmentation methods, on five popular segmentation data sets. Besides, our approach performs competitively to the fully-supervised segmentation models. Yue Gao 0002, Yingjie Xia, Ke Lu 0002, Jialie Shen 0001, Rongrong Ji |
IEEE Trans. Multim. | 4 |
| 2013 | Binary Code Ranking with Weighted Hamming DistanceabstractBinary hashing has been widely used for efficient similarity search due to its query and storage efficiency. In most existing binary hashing methods, the high-dimensional data are embedded into Hamming space and the distance or similarity of two points are approximated by the Hamming distance between their binary codes. The Hamming distance calculation is efficient, however, in practice, there are often lots of results sharing the same Hamming distance to a query, which makes this distance measure ambiguous and poses a critical issue for similarity search where ranking is important. In this paper, we propose a weighted Hamming distance ranking algorithm (WhRank) to rank the binary codes of hashing methods. By assigning different bit-level weights to different hash bits, the returned binary codes are ranked at a finer-grained binary code level. We give an algorithm to learn the data-adaptive and query-sensitive weight for each hash bit. Evaluations on two large-scale image data sets demonstrate the efficacy of our weighted Hamming distance for binary code ranking. Lei Zhang 0119, Yongdong Zhang 0001, Jinhui Tang 0001, Ke Lu 0002, Qi Tian 0001 |
CVPR | 4 |
| 2013 | Extract foreground objects based on sparse model of spatiotemporal spectrumabstractIn this paper, we present a novel foreground object detection method based on the sparse model of the spectrum of spatiotemporal DCT domain, which is robust for high dynamic scenes. First, we adopt the three-dimensional Discrete Cosine Transform (DCT) to calculate the spatiotemporal spectrum representation of the current frame. Then, identification of foreground pixels is formulated as the analysis of the sparse solution of an optimization problem, where foreground pixels correspond to an outlier of the sparse model. Finally, the background updating method is presented to adaptively update the dictionary of sparse model corresponding to background representation. The experimental results on four challenging video sequences show that the proposed method is more robust to high dynamic changes of scenes compared with four representative methods. Zhangjian Ji, Weiqiang Wang 0001, Ke Lu 0002 |
ICIP | 3 |
| 2013 | Social event detection with robust high-order co-clusteringabstractThis paper is devoted to detecting social, real-world events from the sharing images/videos on social media sites like Flickr and YouTube. The fast growing contents make the social media sites become gold mines for social event detection, but we still need to overcome the challenge of processing the associated heterogeneous metadata, such as time-stamp, location, visual content and textual content. Different from the traditional early or late fusion with different types of metadata, we represent them into a star-structured $K$-partite graph, that is, social media itself is regarded as the central vertices set and different types of metadata are treated as the auxiliary vertices sets which are pairwise independent with each other but correlated with the central one. Based on this graph, Social Event Detection with Robust High-Order Co-Clustering (SED-RHOCC) algorithm is proposed and it includes two steps: 1) coarse event detection, 2) clusters and samples refinement. In the first step, by revealing the inter-relationship on the constructed star-structured $K$-partite graph and the intra-relationship within some metadata sets such as time-stamp, we co-cluster social media and the associated metadata separately and iteratively to avoid information loss in early/late fusion. After that, a post process is utilized to refine the clusters and social media samples in the second step. MediaEval Social Event Detection Dataset [1] and its subset are selected to demonstrate the effectiveness of our proposed approach in handling the datasets with and without non-event samples. Bing-Kun Bao, Weiqing Min, Ke Lu 0002, Changsheng Xu |
ICMR | 3 |
| 2013 | Paint the City Colorfully: Location Visualization from Multiple Themes
Quan Fang, Changsheng Xu, Ke Lu 0002 |
MMM (1) | 4 |
| 2013 | Foreground Detection Utilizing Structured Sparse Model via l1, 2 Mixed NormsabstractForeground object detection is a crucial technique of intelligent surveillance systems, and it is still a challenging problem in complex scenes with illumination variations and dynamic backgrounds. Intuitively, the foreground object pixels are often not sparsely distributed but tend to be clustered. Motivated by this hypothesis, we present a new structured sparse model to extract foreground objects, which introduces the spatial neighborhood information into a unified optimization framework by l1,2 mixed norms. Simultaneously, we also give the solving method of the proposed model in details. Moreover, we apply the model to the sparse signal recovery and background subtraction in videos. In the experiments, better performance is obtained over previous methods. The experimental results validate the hypothesis and the effectiveness of the proposed method. Zhangjian Ji, Weiqiang Wang 0001, Ke Lu 0002 |
SMC | 3 |
| 2013 | A flexible 3D cerebrovascular extraction from TOF-MRA images
Yun Tian 0002, Fuqing Duan, Ke Lu 0002, Zhongke Wu, Qingjun Wang, Lin Sun 0002, Lizhi Xie |
Neurocomputing | 3 |
| 2013 | Measuring ontology information by rules based transformation
Yinglong Ma 0001, Ke Lu 0002, Beihong Jin |
Knowl. Based Syst. | 2 |
| 2013 | View-Based Discriminative Probabilistic Modeling for 3D Object Retrieval and RecognitionabstractIn view-based 3D object retrieval and recognition, each object is described by multiple views. A central problem is how to estimate the distance between two objects. Most conventional methods integrate the distances of view pairs across two objects as an estimation of their distance. In this paper, we propose a discriminative probabilistic object modeling approach. It builds probabilistic models for each object based on the distribution of its views, and the distance between two objects is defined as the upper bound of the Kullback-Leibler divergence of the corresponding probabilistic models. 3D object retrieval and recognition is accomplished based on the distance measures. We first learn models for each object by the adaptation from a set of global models with a maximum likelihood principle. A further adaption step is then performed to enhance the discriminative ability of the models. We conduct experiments on the ETH 3D object dataset, the National Taiwan University 3D model dataset, and the Princeton Shape Benchmark. We compare our approach with different methods, and experimental results demonstrate the superiority of our approach. Meng Wang 0001, Yue Gao 0002, Ke Lu 0002, Yong Rui |
IEEE Trans. Image Process. | 3 |
| 2013 | Edge-SIFT: Discriminative Binary Descriptor for Scalable Partial-Duplicate Mobile SearchabstractAs the basis of large-scale partial duplicate visual search on mobile devices, image local descriptor is expected to be discriminative, efficient, and compact. Our study shows that the popularly used histogram-based descriptors, such as scale invariant feature transform (SIFT) are not optimal for this task. This is mainly because histogram representation is relatively expensive to compute on mobile platforms and loses significant spatial clues, which are important for improving discriminative power and matching near-duplicate image patches. To address these issues, we propose to extract a novel binary local descriptor named Edge-SIFT from the binary edge maps of scale- and orientation-normalized image patches. By preserving both locations and orientations of edges and compressing the sparse binary edge maps with a boosting strategy, the final Edge-SIFT shows strong discriminative power with compact representation. Furthermore, we propose a fast similarity measurement and an indexing framework with flexible online verification. Hence, the Edge-SIFT allows an accurate and efficient image search and is ideal for computation sensitive scenarios such as a mobile image search. Experiments on a large-scale dataset manifest that the Edge-SIFT shows superior retrieval accuracy to Oriented BRIEF (ORB) and is superior to SIFT in the aspects of retrieval precision, efficiency, compactness, and transmission cost. Shiliang Zhang, Qi Tian 0001, Ke Lu 0002, Qingming Huang, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Effectively localize text in natural scene images
Ke Lu 0002, Weiqiang Wang 0001 |
ICPR | 2 |
| 2012 | Multimodal Graph-Based Reranking for Web Image SearchabstractThis paper introduces a web image search reranking approach that explores multiple modalities in a graph-based learning scheme. Different from the conventional methods that usually adopt a single modality or integrate multiple modalities into a long feature vector, our approach can effectively integrate the learning of relevance scores, weights of modalities, and the distance metric and its scaling for each modality into a unified scheme. In this way, the effects of different modalities can be adaptively modulated and better reranking performance can be achieved. We conduct experiments on a large dataset that contains more than 1000 queries and 1 million images to evaluate our approach. Experimental results demonstrate that the proposed reranking approach is more robust than using each individual modality, and it also performs better than many existing methods. Meng Wang 0001, Hao Li 0030, Dacheng Tao, Ke Lu 0002, Xindong Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2008 | A Geometric Active Contours Model for Multiple Objects Segmentation
Peng Zhang 0037, Ke Lu 0002 |
ICIC (1) | 3 |