VLDB 2026 Research / reviewers in the wild / expert
Dehui Kong
dblp:14/38
· DBLP profile ↗
80ranked-venue papers
4as first author
51since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 23 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoEmpaTeam: Enhancing Cognitive Empathy using LLM-based Avatars and Dynamic Role Play in Virtual RealityabstractCognitive empathy, the ability to understand others‘ perspectives, is essential for effective communication, reducing biases, and constructive negotiation. However, this skill is declining in a performance-driven society, which prioritizes efficiency over perspective-taking. Here, the training of cognitive empathy is challenging because it is a subtle, hard-to-perceive soft skill. To address this, we developed CoEmpaTeam, a VR-based system that enables users to train their cognitive empathy by using LLM-driven avatars with different personalities. Through dynamic role play, users actively engage in perspective-taking, experiencing situations through another person’s eyes. CoEmpaTeam deploys three avatars who significantly differ in their personality, validated by a technical evaluation and an online experiment (n=90). Next, we evaluated the system through a lab experiment with 32 participants who performed three sessions across two weeks, followed by a one-week diary study. Our results showed a significant increase in cognitive empathy, which, according to participants, transferred into their real lives. Dehui Kong, Martin Feick, Shi Liu 0002, Alexander Maedche |
CHI | 1 |
| 2026 | SubGAva: 3D Gaussian Primitive Subdivision for Photo-realistic and Animatable Human Avatars from Monocular VideoabstractReconstructing photo-realistic and animatable human avatars from monocular RGB videos is a long-standing and challenging problem. Existing methods based on Neural Radiance Field (NeRF) or 3D Gaussian Splatting (3DGS) can model animatable avatars, but often face difficulties in capturing high-frequency dynamic appearance details under monocular settings. To address this issue, we propose SubGAva, an avatar modeling approach based on Gaussian primitive subdivision. Our method employs an efficient geometry–appearance fusion strategy to characterize the coupled variations of geometry and appearance induced by pose changes. In addition, we introduce a controllable Gaussian subdivision mechanism together with a high-frequency guided loss, which improves the reconstruction of fine-grained surface details. Experimental results on People-Snapshot and DynVideo datasets show that our method yields superior rendering quality and more realistic dynamic appearance compared to existing approaches. Dehui Kong |
ICMR | 2 |
| 2026 | Progressive knowledge evolution under multi-granularity knowledge distillation for infrared action recognition
Dehui Kong |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Deep Residual Discriminative Dictionary Learning for image classification
Lichun Wang 0002, Jianjia Xin, Kai Xu 0012, Huiyong Zhang, Shaofan Wang 0001, Dehui Kong |
Knowl. Based Syst. | 7 |
| 2026 | EKDSC: Long-tailed recognition based on expert knowledge distillation for specific categories
Yaping Bai, Dehui Kong, Suqiao Yang |
Neural Networks | 3 |
| 2026 | LADC-Net: Lesion-aware dual-calibrated network for long-tailed medical image classification
Yaping Bai, Dehui Kong, Suqiao Yang |
Pattern Recognit. | 3 |
| 2026 | HAhb-KG: Hierarchical Augmented Knowledge Graph for Human Behavior Assisting Cross-Modal Learning Action DetectionabstractAction detection in untrimmed, densely annotated video datasets is a challenging task due to the presence of composite actions and co-occurring actions in videos. To facilitate action detection in such intricate scenarios, leveraging ample prior information from the data and comprehending the context of actions in the video are the most important two clues. Specifically, the co-occurrence probability of actions can effectively capture the temporal relationships and associations among actions, aiding the model in recognizing multiple actions occurring simultaneously. Additionally, aggregating action information from different levels of the data into a comprehensive graph and describing human actions from various semantic layers can significantly reduce ambiguities in action detection. Based on this, a novel knowledge graph, Hierarchical Augmented Knowledge Graph for human behaviour (HAhb-KG), is proposed, which brings together action-related prior knowledge on different levels into a unified hierarchical graph. The graph describes human behaviour from various semantic aspects by defining diversified graph nodes, and augments the nodes and relationships with corresponding images and probability of co-occurrence respectively, to introduce textual modality information and weigh the associations between actions. In order to mine the knowledge related to the input video in the knowledge graph, HAhb-KG oriented knowledge understanding framework is proposed to embed multi-modal knowledge as a valuable supplement to visual information. Incorporated with the framework, a cross-modal learning action detection model is designed to achieve high accuracy in action detection tasks, which validates the effectiveness of HAhb-KG. Our method achieves gains of 1.45(mAP) and 2.28(mAP) in action detection experiments on the Charades and TSU datasets, respectively, which show that the proposed method outperforms existing knowledge-based action detection methods. Dehui Kong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | ASK-HOI: Affordance-Scene Knowledge Prompting for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection task aims to learn how humans interact with surrounding objects by inferring fine-grained triples of$\left\langle \rm {\emph {human, action, object}} \right\rangle$, which plays a vital role in computer vision tasks such as human-centered scene understanding and visual question answering. However, HOI detection suffers from class long-tailed distributions and zero-shot problems. Current methods typically identify HOI only from input images or label spaces in a data-driven manner, lacking sufficient knowledge prompts, and consequently limits their potential for real-world scenes. Hence, to fill this gap, this paper introduces affordance and scene knowledge as prompts on different granularities to the HOI detector to improve its recognition ability. Concretely, we first construct a large-scale affordance-scene knowledge graph, named ASKG, whose knowledge can be divided into two categories according to the fields of image information, i.e., the knowledge related to affordances of object instances and the knowledge associated with the scene. Subsequently, the knowledge of affordance and scene specific to the input image is extracted by an ASKG-based prior knowledge embedding module. Since this knowledge corresponds to the image at different granularities, we then propose an instance field adaptive fusion module and a scene field adaptive fusion module to enable visual features fully absorb the knowledge prompts. These two encoded features of different fields and knowledge embeddings are finally fed into a proposed HOI recognition module to predict more accurate HOI results. Extensive experiments on both HICO-DET and V-COCO benchmarks demonstrate that the proposed method leads to competitive results compared with the state-of-the-art methods. Dongpan Chen, Dehui Kong, Junna Gao, Qianxing Li |
IEEE Trans. Multim. | 2 |
| 2026 | Text-KeyPoint Human Representation Based Multi-Level Semantic Guided Model for Human Mesh RecoveryabstractHuman mesh recovery (HMR) from the monocular RGB images has attracted a lot of attention from computer vision. Affected by the occlusion between joints and the fact that a 2D pose may match with multiple 3D poses, HMR remains a difficult task. Some mainstream neural network methods realize information exchange in the process of recovering key points extracted from images into meshes, enhance the perception of the vertices and improve the structural rationality of the estimation results. However, the information exchange of vertices without reasonable guidance of global semantic information lacks attention to the global rationality of the estimation results. In order to solve this problem, we propose Text-KeyPoint human representation, which augments key points with semi-structured textual information generated from manual rules to describe the human body at different levels, thus introducing global semantic information. Based on Text-KeyPoint human representation, we propose a novel semantic guided model, which fuses multi-level text features extracted from semi-structured text and multi-level geometric features pooled from key points to provide multi-level semantic guidance in the process of human recovery, thus improving the global rationality of results. Comprehensive experimental results on widely used HMR benchmarks H3.6M, 3DPW and SURREAL show that the proposed method achieves competitive performance compared with the state-of-the-art methods. Qianxing Li, Dehui Kong, Wenshan Shen |
IEEE Trans. Multim. | 2 |
| 2026 | Semantic Prototype Guided Sparse Temporal Interaction for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to detect action segments in untrimmed videos only using video-level labels. Existing methods typically follow the multi-instance learning (MIL) paradigm with a top-k strategy, often resulting in incomplete action localization. Moreover, the local and discontinuous nature of actions causes action segments to be isolated and lack sufficient temporal interaction. To address these issues, this article introduces semantic prototypes to enrich video representations, enabling the model to aggregate category-level action cues across videos and recover semantically relevant but weakly activated segments, thereby improving action completeness. A prototype contrastive loss is further employed to improve feature discriminability. Moreover, a sparse temporal interaction unit is designed to jointly model short-term context and long-range dependencies. The boundary-guided loss utilizes the temporal interaction outputs to explicitly constrain semantic responses around action boundaries, promoting sharp and temporally consistent transitions. Based on these, this article proposes a semantic prototype guided sparse temporal interaction network (S2Net), achieving a unified video modeling from full semantic understanding to fine-grained boundary perception. Extensive experiments on THUMOS14 and ActivityNet1.3 demonstrate that S2Net achieves more accurate and complete action localization. Dehui Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | Multi-Granular Action Detection Based on Highlighted Categories Hierarchical Prior KnowledgeabstractIn untrimmed video data of indoor scenes, actions exhibit complex temporal relationships, such as co-occurrence and compositional dependencies, and multi-granularity semantic hierarchies. Detecting actions under these intricate relationships is a challenging task. Many existing methods primarily rely on complex temporal annotations to address the problem of multi-granularity action detection. However, this approach often lacks explicit modeling of hierarchical dependencies between actions. This oversight results in limited capability to handle multi-granularity action detection across complex temporal scales. To address the limitations, this article proposes a novel multi-granular action detection method based on highlighted a set of categories hierarchical prior knowledge graph, which is first introduced in this article. For representing the hierarchical prior knowledge graph, a Quantified Action Hierarchical Prior Construction (QAHC) approach is proposed, through which the semantics of action categories are integrated with the actions’ temporal patterns to express the hierarchical dependencies and logical relationships among actions both qualitatively and quantitatively in the form of a directed acyclic graph. Based on the hierarchical prior knowledge graph, we design a Prior Highlight Guided Action Detection (PHAD) network. It employs a prior highlighted mechanism to select sufficient and necessary prior knowledge in real-time to guide the efficient detection of multi-granularity actions. Experimental results show that our method achieves excellent accuracy on the Charades (+1.49) and STS (+1.12) datasets. Dehui Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask PromptsabstractAffordance refers to the interactable functional properties of an object, and affordance segmentation aims to pixel-level segment the object functional parts in a given image, which is crucial for various interactive vision tasks. Existing methods address the affordance segmentation problem by utilizing only image features, they can hardly solve the problems of interference between adjacent object pixels in complex scenes, and inability to generalize to the open-world. To tackle these problems, we propose a novel open-vocabulary affordance segmentation task and a benchmark dataset, and propose an approach with object shape mask prompts. The mask is used as prior for different granularity visual feature enhancement and fine-grained text prompt embedding. Specifically, we first propose a mask prompt generation module, which generates refined object shape masks, as well as text prompts for mask-focused regions. Based on the masks, we propose a mask prompt feature enhancement module. It uses masks to encode instance features, and then aggregates them with global features to enhance the visual feature representation. The enhanced visual features are combined with text prompts of different granularity to generate class-agnostic affordance mask proposals. We finally classify these proposals in a proposed affordance prediction module. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior performance on a proposed benchmark dataset. Our approach is also competitive on other open-vocabulary part segmentation datasets. Dongpan Chen, Dehui Kong |
AAAI | 2 |
| 2025 | ELSPU: Expanding Labeled Samples via Semantic Features for Positive Unlabeled LearningabstractPositive Unlabeled (PU) learning, as a form of weakly supervised learning, is characterized by only having a fraction of positive samples labeled, while all the others remain unlabeled. The objective of PU learning is to train a binary classifier that can effectively distinguish positive and negative samples with the incomplete labeling. To achieve a more realistic distribution of samples, we propose a novel approach leveraging the semantic information of the latent classes to expand positive samples via CLIP. Especially, we design and propose a kind of adapter training method, which enables the proposed framework to adapt the medical images well. We validate the effectiveness of the proposed ELSPU method. Extensive experiments demonstrate that ELSPU significantly outperforms baseline methods, with an average accuracy improvement of 1.96% on common benchmarks, highlighting its competitive performance. Qingren Zhang, Dehui Kong |
SMC | 3 |
| 2025 | MMF-Net: A novel multi-feature and multi-level fusion network for 3D human pose estimationabstractAbstract Human pose estimation based on monocular video has always been the focus of research in the human computer interaction community, which suffers mainly from depth ambiguity and self‐occlusion challenges. While the recently proposed learning‐based approaches have demonstrated promising performance, they do not fully explore the complementarity of features. In this paper, the authors propose a novel multi‐feature and multi‐level fusion network (MMF‐Net), which extracts and combines joint features, bone features and trajectory features at multiple levels to estimate 3D human pose. In MMF‐Net, firstly, the bone length estimation module and the trajectory multi‐level fusion module are used to extract the geometric size information of the human body and multi‐level trajectory information of human motion, respectively. Then, the fusion attention‐based combination (FABC) module is used to extract multi‐level topological structure information of the human body, and effectively fuse topological structure information, geometric size information and trajectory information. Extensive experiments show that MMF‐Net achieves competitive results on Human3.6M, HumanEva‐I and MPI‐INF‐3DHP datasets. Qianxing Li, Dehui Kong |
IET Comput. Vis. | 2 |
| 2025 | 3d human pose estimation based on conditional dual-branch diffusion
Zhuowei Bai, Dehui Kong, Dongpan Chen, Qianxing Li |
Multim. Syst. | 3 |
| 2025 | GCLNet: Generalized Contrastive Learning for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to precisely locate action instances in given videos by video-level classification supervision, which is partly related to action classification. Most existing localization works directly utilize feature encoders pre-trained for video classification tasks to extract video features, resulting in non-targeted features that lead to incomplete or over-complete action localization. Therefore, we propose Generalized Contrast Learning Network (GCLNet), in which two novel strategies are proposed to improve the pre-trained features. First, to address the issue of over-completeness, GCLNet introduces text information with good context independence and category separability to enrich the expression of video features, as well as proposes a novel generalized contrastive learning approach for similarity metrics, which facilitates pulling closer the features belonging to the same category while pushing farther apart those from different categories. Consequently, it enables more compact intra-class feature learning and ensures accurate action localization. Second, to tackle the problem of incomplete, we exploit the respective advantages of RGB and Flow features in scene appearance and temporal motion expression, designing a hybrid attention strategy in GCLNet to enhance each channel features mutually. This process greatly improves the features through establishing cross-channel consensus. Finally, we conduct extensive experiments on THUMOS14 and ActivityNet1.2, respectively, and the results show that our proposed GCLNet can produce more representative action localization features. Dehui Kong |
IEEE Trans. Big Data | 2 |
| 2025 | Dual-Branch Hypergraph Convolutional Network Learning Spectral-Spatial-Semantic Features for Hyperspectral Image ClassificationabstractHyperspectral Image (HSI) classification, which aims to assign pixel-level categories for given HSI data, has achieved remarkable success with deep learning architectures. However, such approaches typically require large amounts of annotated data, which are often scarce in practical applications. To address the challenge of limited annotated samples, this paper proposes a novel framework that leverages both semantic and structural prior knowledge to extract more expressive HSI features. This paper proposes a Dual-Branch Hypergraph Convolutional Network (DB-HGCN) for comprehensive spectral-spatial-semantic feature extraction, which employs multi-hop constrained superpixel-level hypergraphs to effectively model the complex high-order correlations inherent in HSI data. The proposed architecture consists of two complementary branches: (1) a Semantic-enhanced Hypergraph Convolutional Network (SSeHGCN) that incorporates category-specific semantic knowledge through a vision-language model to enhance spatial representations, and (2) a Spatial-Spectral enhanced Hypergraph Convolutional Network (SSpHGCN) that captures both intra-superpixel visual features and their interrelationships via a novel Spatial-Spectral Feature and Relation Fusion Module (SSRFM). Extensive experiments on four benchmark datasets demonstrate the superiority of the proposed approach, with DB-HGCN achieving state-of-the-art overall classification accuracy (OA) using only five labeled samples per class. This significant performance gain highlights the effectiveness of our method in addressing the data scarcity challenge in HSI classification. Shuran Jing, Yijie Ding, Dehui Kong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multi-Anchor Offset Representation Based Coarse-to-Fine Diffusion Model for Human Pose Estimationabstract3D human pose estimation (3DHPE) in images aims at estimating 3D joint positions from images. The existing 3DHPE methods usually define the loss function as the error measured by Euclidean distance between the locations of the predicted joints and the ground truth of joints, which confuses two different kinds of errors with obviously different characteristics and should not be processed equally: the error caused by different pose structures and the others. However, The existing human pose representations are not suitable to distinguish these two kinds of errors. In order to tackle this problem, we propose a novel Multi-Anchor Offset Representation (MAOR) for human pose, which locates the position of each joint using its offsets from a group of selected high-precision joints named Multi-Anchor. Making use of MAOR, the pose error related to the distortion of spatial structure can be measured independently from other errors, which is helpful in promoting the accuracy of pose estimation. We then propose a novel MAOR-based coarse-to-fine diffusion model (MAOR-DiffPose) for pose estimation, which optimizes different types of errors of poses step by step. Firstly, a MAOR-based Denoising Process (MDP) is devised to explicitly optimize spatial structures of 3D poses by using MAOR to describe poses and improves the inductive learning ability of MAOR-DiffPose by extracting view-independent features. Secondly, a Joint Coordinate Denoising Process assisted by MAOR (JCDPaM) is devised to expand the input features meaningfully by combining MAOR with the pose representation based on joint coordinates and optimize the joint coordinates of 3D poses with the assistance of MAOR. MAOR-DiffPose realizes accurate 3DHPE by iterating MDP and JCDPaM modules. Comprehensive experimental results on widely used 3DHPE benchmarks Human3.6M and MPI-INF-3DHP show that the proposed method achieves competitive performance compared with the state-of-the-art methods. Qianxing Li, Dehui Kong, Dongpan Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | FusionArch: A Fusion-Based Accelerator for Point-Based Point Cloud Neural NetworksabstractPoint-based Point Cloud Neural Networks (PCNNs) have attracted much attention for their higher accuracy than voxel-based and multi-view-based PCNNs. Nevertheless, the increasing scale of point cloud data poses a challenge for real-time processing. Numerous previous works focus on accelerating PCNN inference but only optimize specific stages, limiting their generality to different networks with diverse performance bottlenecks. In this paper, we take nearly all stages of PCNNs into account, and propose 3 orthogonal algorithms, including Fusion-FPS, Fusion-Computation, and Fusion-Aggregation. We introduce Fusion-FPS to alter the sequential execution flow by reducing the Farthest Point Sampling (FPS) across layers to once and organize all neighbor search stages in parallel. To exclude redundant feature computations of “Filling Points”, we propose Fusion-Computation, identifying the presence and locations of “Filling Points” and directly borrowing the nearest neighbor features for them. To eliminate redundant memory accesses caused by shared neighbors in aggregation, we present Fusion-Aggregation, which clusters nearby centroids and coalesces their replicated accesses. In support of our algorithms, we co-design FusionArch, an architecture that implements our strategies and further optimizes memory access via a Local Fusion-Aggregation Table (LFT). We evaluate FusionArch on both server-level and edge-level platforms on 5 PCNNs across 4 applications and show remarkable accuracy and performance gains. On average, FusionArch achieves$2.6\times,5.6\times, 13.0\times$speedup and$17\times, 22\times, 62.4\times$energy savings over PointAcc.Server, NVIDIA AIOO GPU and Intel Xeon CPU, respectively. Moreover, it outperforms PRADA, PointAcc.Edge, Mesorasi and GPU with speedups of$2.4\times, 2.9\times, 5.3\times, 5.5\times$, and energy savings of$4.4\times, 7.2\times, 12.4\times, 11.5\times$, respectively. Xueyuan Liu 0001, Zhuoran Song, Guohao Dai 0001, Gang Li 0015, Can Xiao, Dehui Kong, Xiaoyao Liang |
DATE | 7 |
| 2024 | Dual-Branch Network with Online Knowledge Distillation for 3D Hand Pose Estimation
Yingqi He, Dehui Kong |
ICANN (3) | 3 |
| 2024 | Low-Light Raw Image Enhancement on a Dataset Suffering Light EffectsabstractDeep learning-based methods have achieved remarkable success in low-light image enhancement (LLIE). But most existing works are based on sRGB data and do not focus on the light effects in bright regions when enhancing low-light regions. This inevitably leads to excessive enhancement and saturation of bright regions, resulting in reduced contrast and inaccurate color. To address this problem, a low-light raw dataset covering diverse lighting conditions is proposed to overcome the limitations of the existing datasets and to supervise the training of our model. Then, we design a new enhancement network that incorporates global information to learn mapping curves from low-light images to Ground Truth (GT). A novel loss function is also proposed to help achieve high-quality enhancement for a low-light raw image suffering light effects. In terms of qualitative evaluations, our approach performs best in suppressing light effects and boosting the intensity of dark regions compared with other state-of-the-art low-light algorithms. In quantitative tests, it is also shown that the proposed method has the highest peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM), suggesting a superior enhancement performance. Guipeng Zhang, Dehui Kong |
ICASSP | 4 |
| 2024 | Dual-Branch Knowledge Enhancement Network with Vision-Language Model for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, pre-trained Vision-Language Models (VLM) have shown their great recognition ability in HOI detection task. However, these VLM based methods are struggle to transfer knowledge to achieve desired performance. To this end, we propose a Dual-Branch Knowledge Enhancement Network with VLM (DBKEN-VLM) within the two-stage paradigm to enhance the effectiveness of VLM. Specifically, we propose a semantic mining decoder to supplement contextual and action-related semantic information into our model. It forms a dual-branch knowledge enhancement network with spatial guided decoder. Furthermore, we propose a two-level fusion strategy for the dualbranch network to facilitate better knowledge transfer of VLM. One is feature-level fusion, producing more instructive interaction features; another is decision-level fusion, further enhancing the capability of VLM for HOI detection. The proposed method achieves competitive performance compared to recent methods on two benchmark datasets, HICO-DET and V-COCO. Guangpu Zhou, Dehui Kong, Dongpan Chen, Zhuowei Bai |
IJCNN | 2 |
| 2024 | A Dual-Branch Structural Network for Human Pose Estimation Based on Millimeter Wave RadarabstractRadar-based Human Pose Estimation (R-HPE) aims to locate the body joints of each individual in a given radar image. This is relevant for various applications such as action recognition, person re-identification, and human-object interaction. Unlike traditional RGB-based human pose estimation, radar-based human pose estimation can effectively preserve human privacy and remain stable under low-light conditions and darkness. However, research on radar-based human pose estimation is limited, and existing methods fail to adequately model radar features, resulting in lower accuracy of corresponding pose estimation algorithms. Therefore, this paper proposes a dual-branch structured network, which can extract dimension-independent features and dimension-dependent features separately and then combine them for more precise decision-making. This allows the network to learn richer and more diverse feature representations, thereby improving the quality of feature extraction. Meanwhile, a Multi-Dimensional Feature Fusion Network extracts more detailed feature representations. Furthermore, it is combined with a Transformer module to further enhance the model's ability to extract local features and global modeling capabilities, thereby improving the accuracy of human pose estimation. Extensive experiments conducted on the HuPR[10] dataset demonstrate that our model outperforms existing state-of-the-art models in terms of human pose estimation performance. Hanxin Chen, Dehui Kong |
SMC | 3 |
| 2024 | An entity alignment approach coupling NGBoost and SHAP for constructing spatio-temporal evolution knowledge graph from historical atlasesabstractThe historical atlases provide a wealth of information about the evolution of geography over time and space. The alignment of geographical entities across varying time periods is a crucial aspect of extracting meaningful insights into the spatio-temporal dynamics of geography. This paper proposes a geographic entity alignment approach coupling Natural Gradient Boosting (NGBoost) with SHapley Additive exPlanation (SHAP). Taking the historical atlas of China as a case study, a geographic entity alignment model based on NGBoost is constructed considering the different kinds of similarity features of geographic entities, including semantic, distance, shape, size and topology. The contribution of similarity features in the NGBoost model is analyzed using the SHAP framework so as to improve the explanatory capacity of the model. The spatio-temporal evolution relationships of geographic entities are generated by association rules depending on alignment types and represented as quadruples, for constructing geographic knowledge graphs. The proposed NGBoost method was found a superior accuracy by comparing with BP neural networks, random forests, and other alternative methods for aligning geographic entities. The constructed geographic spatio-temporal evolution knowledge graphs offer valuable support for the queries of evolutionary knowledge. Yongquan Yang, Min Cao 0006, Dehui Kong, Min Chen 0008 |
Int. J. Geogr. Inf. Sci. | 3 |
| 2024 | Intentional Understanding and Human-Computer Collaboration: A Smart Pen for Solid Geometry TeachingabstractThe current teaching mode in geometry mainly focuses on two-dimensional levels, and the teaching tools utilized are static and not interactive. This paper proposes the use of a smart pen for three-dimensional geometry experimental teaching and a multimodal intention understanding and human-computer collaboration algorithm for the smart pen. The primary innovations of this paper lie in the development of a smart pen and a virtual platform for geometry education tailored for geometry experimental instruction. This system can promptly perceive and comprehend user behavior in real-time. Furthermore, a standardized topological equivalence model is proposed as the basis for a point selection strategy. By establishing correspondence between the modeled point selection model and the actual operation scene, the behavioral intent imposed on the model is applied to the operation object of the actual scene. Additionally, CNN-based and information entropy-based multimodal fusion intention understanding models are proposed for different input modalities to capture the operational intention of users by fusing their multimodal input data. The algorithm further improves the accuracy rate through an error correction mechanism based on implicit interaction to achieve better human-computer collaboration. The algorithm proposed in this paper has resulted in a 0.47-second improvement in point selection average time and achieved an intent understanding accuracy of 98.21%. This improvement leads to better fault tolerance and fluency during human-computer interaction, reduces the cognitive load on the user, and improves the overall user experience. Dehui Kong, Zhiquan Feng, Zishuo Xia, Weina Li |
Int. J. Hum. Comput. Interact. | 1 |
| 2024 | Joint multi-scale transformers and pose equivalence constraints for 3D human pose estimation
Yongpeng Wu 0002, Dehui Kong, Junna Gao |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | ADOSMNet: a novel visual affordance detection network with object shape mask guided feature encoders
Dongpan Chen, Dehui Kong, Shaofan Wang 0001 |
Multim. Tools Appl. | 2 |
| 2024 | OASNet: Object Affordance State Recognition Network With Joint Visual Features and Relational Semantic EmbeddingsabstractTraditional affordance learning tasks aim to understand object’s interactive functions in an image, such as affordance recognition and affordance detection. However, these tasks cannot determine whether the object is currently interacting, which is crucial for many follow-up tasks, including robotic manipulation and planning task. To fill this gap, this paper proposes a novel object affrodance state (OAS) recognition task, i.e., simultaneously recognizing an object’s affordances and the partner objects that are interacting with it. Accordingly, to facilitate the application of deep learning technology, an OAS recognition task related dataset OAS10k is constructed by collecting and labeling over 10k images. In the dataset, a sample is defined as a set of an image and its OAS labels, each label is represented as$\left \langle{ \rm {\textit {subject, subject's affrodance, interacted object}} }\right \rangle $. These triplet labels have rich relational semantic information, which can improve OAS recognition performance. We hence construct a directed OAS knowledge graph of affordance states, and extract an OAS matrix from it for modelling the semantic relationships of the triplets. Based on the matrix, we propose an OAS recognition network (OASNet), which utilizes GCN to capture the relational semantic embeddings, and uses a transformer to fuse them with the visual features from an image to recognize the affordance states of objects in the image. Experimental results on OAS10k dataset and other triplet label recognition datasets demonstrate that the proposed OASNet achieves the best performance compared to the state-of-the-art methods. The dataset and codes will be released onhttps://github.com/mxmdpc/OAS. Dongpan Chen, Dehui Kong, Lichun Wang 0002, Junna Gao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | SYENet: A Simple Yet Effective Network for Multiple Low-Level Vision Tasks with Real-time Performance on Mobile DeviceabstractWith the rapid development of AI hardware accelerators, applying deep learning-based algorithms to solve various low-level vision tasks on mobile devices has gradually become possible. However, two main problems still need to be solved: task-specific algorithms make it difficult to integrate them into a single neural network architecture, and large amounts of parameters make it difficult to achieve real-time inference. To tackle these problems, we propose a novel network, SYENet, with only 6K parameters, to handle multiple low-level vision tasks on mobile devices in a real-time manner. The SYENet consists of two asymmetrical branches with simple building blocks. To effectively connect the results by asymmetrical branches, a Quadratic Connection Unit(QCU) is proposed. Furthermore, to improve performance, a new Outlier-Aware Loss is proposed to process the image. The proposed method proves its superior performance with the best PSNR as compared with other networks in real-time applications such as Image Signal Processing(ISP), Low-Light Enhancement(LLE), and Super-Resolution(SR) with 2K60FPS throughput on Qualcomm 8 Gen 1 mobile SoC(System-on-Chip). Particularly, for ISP task, SYENet got the highest score in MAI 2022 Learned Smartphone ISP challenge. Weiran Gou, Ziyao Yi, Shaoqing Li, Zibin Liu, Dehui Kong |
ICCV | 6 |
| 2023 | Deep Moore-Penrose Inverse Network with Refinement Strategy for One-class ClassificationabstractMultilayer least-square (LS)-based one-class classification networks (MLS-OCNs) have gained great attention for the purpose of identifying anomalies and outliers. However, many MLS-OCNs encounter the issue of loosely connected feature coding because they use two separate mechanisms for feature encoding and final pattern recognition. This paper proposes a solution to this problem by introducing a multilayer algorithm called deep Moore-Penrose inverse network with refinement (DMPINR). In particular, DMPINR employs an end-to-end learning process based on the Moore-Penrose inverse (MPI) to identify optimal latent space and classify objects simultaneously. To enhance the robustness of representations, the DMPINR technique pulls back the residual error from the output layer to the hidden layers sequentially, recalculating the parameters of these hidden layers using MPI. The experimental results on ten popular OCC datasets demonstrate that the proposed approach outperforms many existing MLS-OCNs in G-Mean and F1scores. Junna Gao, Dehui Kong, Weisi Lin, Wandong Zhang |
SMC | 2 |
| 2023 | Learning visual-and-semantic knowledge embedding for zero-shot image classification
Dehui Kong, Xiliang Li, Shaofan Wang 0001 |
Appl. Intell. | 1 |
| 2023 | MAG: a smart gloves system based on multimodal fusion perception
Hong Cui, Zhiquan Feng, Jinglan Tian, Dehui Kong, Zishuo Xia, Weina Li |
CCF Trans. Pervasive Comput. Interact. | 4 |
| 2023 | HyperGraph based human mesh hierarchical representation and reconstruction from a single imageabstractReconstructing 3D human mesh from monocular images has been extensively studied. However, the existing non-parametric reconstruction methods are inefficient when modeling vertex relationship concerning human information due to they generally adopt an uniform template mesh. To this end, this paper proposes a novel hypergraph-based human mesh hierarchical representation that enables the expression of vertices at body, parts, and vertices perspectives, corresponding to global, local and individual, respectively. And then a novel HyperGraph Attention-based human mesh reconstruction network ( HGaMRNet ) is put forward in turn, which mainly consists of two modules and can efficiently capture human information from different granularities of the mesh. Specifically, the first module, Body2Parts, decouples a body into local parts, and leverages Mix-Attention (MAT) based feature encoder to learn visual cues and semantic information of the parts for capturing complex human kinematic relationships from monocular images. The second module, Part2Vertices, transfers part features to the corresponding vertices through an adaptive incidence matrix , and utilizes a HyperGraph Attention network to update the vertex features. This is conductive to learning the fine-grained morphological information of a human mesh. All in one, supported by the hypergraph-based hierarchical representation of the human mesh, HGaMRNet balances the effects of neighbor vertex from different levels properly and eventually promotes the reconstruction accuracy of human mesh. Experiments conducted on both Human3.6M and 3DPW datasets show that HGaMRNet outperforms most of the existing image-based human mesh reconstruction methods. Chenhui Hao, Dehui Kong |
Comput. Graph. | 2 |
| 2023 | CIGNet: Category-and-Intrinsic-Geometry Guided Network for 3D coarse-to-fine reconstruction
Junna Gao, Dehui Kong, Shaofan Wang 0001 |
Neurocomputing | 2 |
| 2023 | Multi-scale latent feature-aware network for logical partition based 3D voxel reconstruction
Dehui Kong, Shaofan Wang 0001, Qianxing Li |
Neurocomputing | 2 |
| 2023 | A Survey of Visual Affordance Recognition Based on Deep LearningabstractVisual affordance recognition is an important research topic in robotics, human-computer interaction, and other computer vision tasks. In recent years, deep learning-based affordance recognition methods have achieved remarkable performance. However, there is no unified and intensive survey of these methods up to now. Therefore, this article reviews and investigates existing deep learning-based affordance recognition methods from a comprehensive perspective, hoping to pursue greater acceleration in this research domain. Specifically, this article first classifies affordance recognition into five tasks, delves into the methodologies of each task, and explores their rationales and essential relations. Second, several representative affordance recognition datasets are investigated carefully. Third, based on these datasets, this article provides a comprehensive performance comparison and analysis of the current affordance recognition methods, reporting the results of different methods on the same datasets and the results of each method on different datasets. Finally, this article summarizes the progress of affordance recognition, outlines the existing difficulties and provides corresponding solutions, and discusses its future application trends. Dongpan Chen, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Big Data | 2 |
| 2023 | Hierarchical Coupled Discriminative Dictionary Learning for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize images of novel classes, but does not use any images belonging to the novel classes during model training, which is realized by exploiting the auxiliary semantic information. Recently, most ZSL methods focus on learning visual-semantic embeddings to transfer knowledge from the seen classes to the novel classes. Visual-semantic embedding is usually established based on the visual features of images and the semantic information of classes, i.e., class attributes. However, image features are extracted at the individual level, while class attributes are obtained at the group level, so the granularity of these features is different, which makes it difficult to match the two kinds of features. To tackle such problem, we propose hierarchical coupled discriminative dictionary learning (HCDDL) method to hierarchically establish visual-semantic embedding at class-level and image-level with a coarse-to-fine way. Firstly, a class-level coupled dictionary is trained to build basic and coarse-grained connection between visual space and semantic space. Using the class-level coupled dictionary, image attributes are generated. Based on the fine-grained image attributes and images features, an image-level coupled dictionary is learned. In addition, during the learning of hierarchical coupled dictionaries, the discriminative losses are adopted to ensure dictionaries learn more accurate representation, which is beneficial to the recognition task. Recognition of unseen images is performed through searching the class nearest to the unseen image in multiple spaces. Experiments on four widely used benchmark datasets show the effectiveness of the proposed method, and sufficient ablation experiments demonstrate that the coarse-to-fine way leads to good performances. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | DASI: Learning Domain Adaptive Shape Impression for 3D Object ReconstructionabstractPrevious 3D object reconstruction methods from 2D images involve two issues: the lack of in-depth exploration of the prior knowledge of 3D shapes, and the difficulty of dealing with the serious occluded parts. Inspired by human’s perception on real-world objects which is composed of an overall impression (known asshape impression) and an enhanced cognition, we propose a deep network (denoted by DASI) to learn the Domain Adaptive Shape Impression for 3D reconstruction from arbitrary view images. DASI consists of two modules: shape reconstruction module and shape refinement module. The former module reconstructs a coarse volume by learning a domain adaptive shape impression as embedding in image-based reconstruction. We first leverage 3D objects to learn a shape impression being associated with prior knowledge of 3D objects. To attain consensus on shape impression from 2D images, we regard the 3D shape and the 2D image as two different domains. By adapting the two domains, the shape impression learned from 3D objects is transferred to 2D images and guides the images-based reconstruction. The latter module refines the objects by modeling the whole 3D volume to local 3D patches and exploring their intrinsic geometry relationships. Quantitative and qualitative experimental results on two benchmark datasets demonstrate that DASI outperforms several state-of-the-arts for 3D reconstruction from single and multi-view 2D images. Junna Gao, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | HPGCN: Hierarchical poselet-guided graph convolutional network for 3D pose estimation
Yongpeng Wu 0002, Dehui Kong, Shaofan Wang 0001 |
Neurocomputing | 2 |
| 2022 | Grassmannian graph-attentional landmark selection for domain adaptation
Shaofan Wang 0001, Dehui Kong |
Multim. Tools Appl. | 3 |
| 2022 | GAN for vision, KG for relation: A two-stage network for zero-shot action recognition
Dehui Kong, Shaofan Wang 0001 |
Pattern Recognit. | 2 |
| 2022 | Real-Time Human Action Recognition Using Locally Aggregated Kinematic-Guided Skeletonlet and Supervised Hashing-by-Analysis Modelabstract3-D action recognition is referred to as the classification of action sequences which consist of 3-D skeleton joints. While many research works are devoted to 3-D action recognition, it mainly suffers from three problems: 1) highly complicated articulation; 2) a great amount of noise; and 3) low implementation efficiency. To tackle all these problems, we propose a real-time 3-D action-recognition framework by integrating the locally aggregated kinematic-guided skeletonlet (LAKS) with a supervised hashing-by-analysis (SHA) model. We first define the skeletonlet as a few combinations of joint offsets grouped in terms of the kinematic principle and then represent an action sequence using LAKS, which consists of a denoising phase and a locally aggregating phase. The denoising phase detects the noisy action data and adjusts it by replacing all the features within it with the features of the corresponding previous frame, while the locally aggregating phase sums the difference between an offset feature of the skeletonlet and its cluster center together over all the offset features of the sequence. Finally, the SHA model combines sparse representation with a hashing model, aiming at promoting the recognition accuracy while maintaining high efficiency. Experimental results on MSRAction3D, UTKinectAction3D, and Florence3DAction datasets demonstrate that the proposed method outperforms state-of-the-art methods in both recognition accuracy and implementation efficiency. Shaofan Wang 0001, Dehui Kong, Lichun Wang 0002 |
IEEE Trans. Cybern. | 3 |
| 2022 | A Spatial Relationship Preserving Adversarial Network for 3D Reconstruction from a Single Depth ViewabstractRecovering the geometry of an object from a single depth image is an interesting yet challenging problem. While previous learning based approaches have demonstrated promising performance, they don’t fully explore spatial relationships of objects, which leads to unfaithful and incomplete 3D reconstruction. To address these issues, we propose a Spatial Relationship Preserving Adversarial Network (SRPAN) consisting of 3D Capsule Attention Generative Adversarial Network (3DCAGAN) and 2D Generative Adversarial Network (2DGAN) for coarse-to-fine 3D reconstruction from a single depth view of an object. Firstly, 3DCAGAN predicts the coarse geometry using an encoder-decoder based generator and a discriminator. The generator encodes the input as latent capsules represented as stacked activity vectors with local-to-global relationships (i.e., the contribution of components to the whole shape), and then decodes the capsules by modeling local-to-local relationships (i.e., the relationships among components) in an attention mechanism. Afterwards, 2DGAN refines the local geometry slice-by-slice, by using a generator learning a global structure prior as guidance, and stacked discriminators enforcing local geometric constraints. Experimental results show that SRPAN not only outperforms several state-of-the-art methods by a large margin on both synthetic datasets and real-world datasets, but also reconstructs unseen object categories with a higher accuracy. Dehui Kong, Shaofan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Latent Feature-Aware and Local Structure-Preserving Network for 3D Completion from a Single Depth View
Dehui Kong, Shaofan Wang 0001 |
ICANN (2) | 2 |
| 2021 | Zero-shot Recognition with Image Attributes Generation using Hierarchical Coupled Dictionary LearningabstractZero-shot learning (ZSL) aims to recognize images from unseen (novel) classes with the training images from seen classes. The attributes of each class is exploited as auxiliary semantic information. Recently most ZSL approaches focus on learning visual-semantic embeddings to transfer knowledge from the seen classes to the unseen classes. However, few works study whether the auxiliary semantic information in the class-level is extensive enough or not for the ZSL task. To tackle such problem, we propose a hierarchical coupled dictionary learning (HCDL) approach to hierarchically align the visual-semantic structures in both the class-level and the image-level. Firstly, the class-level coupled dictionary is trained to establish a basic connection between visual space and semantic space. Then, the image attributes are generated based on the basic connection. Finally, the fine-grained information can be embedded by training the image-level coupled dictionary. Zero-shot recognition is performed in multiple spaces by searching the nearest neighbor class of the unseen image. Experiments on two widely used benchmark datasets show the effectiveness of the proposed approach. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
MMAsia | 4 |
| 2021 | A Local-Global Commutative Preserving Functional Map for Shape CorrespondenceabstractExisting non-rigid shape matching methods mainly involve two disadvantages. (a) Local details and global features of shapes can not be carefully explored. (b) A satisfactory trade-off between the matching accuracy and computational efficiency can be hardly achieved. To address these issues, we propose a local-global commutative preserving functional map (LGCP) for shape correspondence. The core of LGCP involves an intra-segment geometric submodel and a local-global commutative preserving submodel, which accomplishes the segment-to-segment matching and the point-to-point matching tasks, respectively. The first submodel consists of an ICP similarity term and two geometric similarity terms which guarantee the correct correspondence of segments of two shapes, while the second submodel guarantees the bijectivity of the correspondence on both the shape level and the segment level. Experimental results on both segment-to-segment matching and point-to-point matching show that, LGCP not only generate quite accurate matching results, but also exhibit a satisfactory portability and a high efficiency. Qianxing Li, Shaofan Wang 0001, Dehui Kong |
MMAsia | 3 |
| 2021 | Deep3D reconstruction: methods, data, and challengesabstractThree-dimensional (3D) reconstruction of shapes is an important research topic in the fields of computer vision, computer graphics, pattern recognition, and virtual reality. Existing 3D reconstruction methods usually suffer from two bottlenecks: (1) they involve multiple manually designed states which can lead to cumulative errors, but can hardly learn semantic features of 3D shapes automatically; (2) they depend heavily on the content and quality of images, as well as precisely calibrated cameras. As a result, it is difficult to improve the reconstruction accuracy of those methods. 3D reconstruction methods based on deep learning overcome both of these bottlenecks by automatically learning semantic features of 3D shapes from low-quality images using deep networks. However, while these methods have various architectures, in-depth analysis and comparisons of them are unavailable so far. We present a comprehensive survey of 3D reconstruction methods based on deep learning. First, based on different deep learning model architectures, we divide 3D reconstruction methods based on deep learning into four types, recurrent neural network, deep autoencoder, generative adversarial network, and convolutional neural network based methods, and analyze the corresponding methodologies carefully. Second, we investigate four representative databases that are commonly used by the above methods in detail. Third, we give a comprehensive comparison of 3D reconstruction methods based on deep learning, which consists of the results of different methods with respect to the same database, the results of each method with respect to different databases, and the robustness of each method with respect to the number of views. Finally, we discuss future development of 3D reconstruction methods based on deep learning. Dehui Kong, Shaofan Wang 0001, Zhiyong Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | Joint Transferable Dictionary Learning and View Adaptation for Multi-view Human Action RecognitionabstractMulti-view human action recognition remains a challenging problem due to large view changes. In this article, we propose a transfer learning-based framework called transferable dictionary learning and view adaptation (TDVA) model for multi-view human action recognition. In the transferable dictionary learning phase, TDVA learns a set of view-specific transferable dictionaries enabling the same actions from different views to share the same sparse representations, which can transfer features of actions from different views to an intermediate domain. In the view adaptation phase, TDVA comprehensively analyzes global, local, and individual characteristics of samples, and jointly learns balanced distribution adaptation, locality preservation, and discrimination preservation, aiming at transferring sparse features of actions of different views from the intermediate domain to a common domain. In other words, TDVA progressively bridges the distribution gap among actions from various views by these two phases. Experimental results on IXMAS, ACT4 2 , and NUCLA action datasets demonstrate that TDVA outperforms state-of-the-art methods. Dehui Kong, Shaofan Wang 0001, Lichun Wang 0002 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2021 | DLGAN: Depth-Preserving Latent Generative Adversarial Network for 3D ReconstructionabstractAlthough deep networks based methods outperform traditional 3D reconstruction methods which require multiocular images or class labels to recover the full 3D geometry, they may produce incomplete recovery and unfaithful reconstruction when facing occluded parts of 3D objects. To address these issues, we propose Depth-preserving Latent Generative Adversarial Network (DLGAN) which consists of 3D Encoder-Decoder based GAN (EDGAN, serving as a generator and a discriminator) and Extreme Learning Machine (ELM, serving as a classifier) for 3D reconstruction from a monocular depth image of an object. Firstly, EDGAN decodes a latent vector from the 2.5D voxel grid representation of an input image, and generates the initial 3D occupancy grid under common GAN losses, a latent vector loss and a depth loss. For the latent vector loss, we design 3D deep AutoEncoder (AE) to learn a target latent vector from ground truth 3D voxel grid and utilize the vector to penalize the latent vector encoded from the input 2.5D data. For the depth loss, we utilize the input 2.5D data to penalize the initial 3D voxel grid from 2.5D views. Afterwards, ELM transforms float values of the initial 3D voxel grid to binary values under a binary reconstruction loss. Experimental results show that DLGAN not only outperforms several state-of-the-art methods by a large margin on both a synthetic dataset and a real-world dataset, but also predicts more occluded parts of 3D objects accurately without class labels. Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Hardness-Aware Dictionary Learning: Boosting Dictionary for RecognitionabstractSparse representation is a powerful tool in many visual applications since images can be represented effectively and efficiently with a dictionary. Conventional dictionary learning methods usually treat each training sample equally, which would lead to the degradation of recognition performance when the samples from same category distribute dispersedly. This is because the dictionary focuses more on easy samples (known as highly clustered samples), and those hard samples (known as widely distributed samples) are easily ignored. As a result, the test samples which exhibit high dissimilarities to most of intra-category samples tend to be misclassified. To circumvent this issue, this paper proposes a simple and effective hardness-aware dictionary learning (HADL) method, which considers training samples discriminatively based on the AdaBoost mechanism. Different from learning one optimal dictionary, HADL learns a set of dictionaries and corresponding sub-classifiers jointly in an iterative fashion. In each iteration, HADL learns a dictionary and a sub-classifier, and updates the weights based on the classification errors given by current sub-classifier. Those correctly classified samples are assigned with small weights while those incorrectly classified samples are assigned with large weights. Through the iterated learning procedure, the hard samples are associated with different dictionaries. Finally, HADL combines the learned sub-classifiers linearly to form a strong classifier, which improves the overall recognition accuracy effectively. Experiments on well-known benchmarks show that HADL achieves promising classification results. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
IEEE Trans. Multim. | 4 |
| 2021 | Discriminative matrix-variate restricted Boltzmann machine classification model
Pengyu Tian, Dehui Kong, Lichun Wang 0002, Shaofan Wang 0001 |
Wirel. Networks | 3 |
| 2020 | An Efficient Multi-scale Method for Single Image Super ResolutionabstractConvolutional neural network has shown its superior performance in single image super resolution(SR) as in other computer vision applications. One of its disadvantage is that the computation needs will be significantly higher than that of other applications for the resolution of input is often high definition and it is growing rapidly. In this paper, the low-level vision attributed of super resolution is emphasized by using a multi-scale neural network. The proposed method do not take a very deep framework which is not suitable in ultra-high resolution or mobile terminals. To obtain the geometry structure, a deformable convolution layer is applied in shallow feature extraction which reduces the complexity of the network structure. The deformable convolution scheme constraints the explosion of parameter space and provides shallow features at different scales. The backbone structure is pyramid which offers more scale information for reconstruction. These ideas ensure the efficiency of the model which is also demonstrated on extensive experiments. Dehui Kong, Ke Xu 0014, Tongtong Zhu, Hengqi Liu, Jianjun Song |
IWCMC | 1 |
| 2020 | Matrix-variate variational auto-encoder with applications to image process
Huixia Yan, Junbin Gao, Dehui Kong, Lichun Wang 0002, Shaofan Wang 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | An Unsupervised Real-Time Framework of Human Pose Tracking From Range Image SequencesabstractPose tracking from range image sequences remains a difficult task due to strong noise and serious self-occlusion of human body. Existing work either rely on extremely large and precisely annotated datasets, or rely on accurate human mesh model and GPU acceleration. In this paper, we propose an unsupervised real-time framework of pose tracking from range image sequences. Our framework consists of a visible hybrid model (VHM), a componentwise correspondence optimization (CCO) and a dynamic database lookup (DDL). VHM consists of component sphere sets and component visible spherical point sets which exhibits both simplicity and high accuracy. CCO converts the matching between VHM and input point cloud into several subproblems regarding local rotations of components and a global translation of body abdominal joint, each of which has an efficient closed form solution. DDL is designed to recover correct pose when tracking fails, which effectively mitigates accumulative error during tracking. Experiments on SMMC, PDT, EVAL datasets indicate that our framework not only achieves better or competitive precision compared with state-of-the-art methods, but also produces real-time efficiency in personal computers without GPU acceleration. Yongpeng Wu 0002, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | 3D human pose estimation from range images with depth difference and geodesic distance
Dehui Kong, Shaofan Wang 0001, Zhiyong Wang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Effective human action recognition using global and local offsets of skeleton joints
Dehui Kong, Shaofan Wang 0001, Lichun Wang 0002 |
Multim. Tools Appl. | 2 |
| 2019 | Unsupervised Learning of Human Pose Distance Metric via Sparsity Locality Preserving ProjectionsabstractHuman poses admit complicated articulations and multigranular similarity. Previous works on learning human pose metric utilize sparse models, which concentrate large weights on highly similar poses and fail to depict an overall structure of poses with multigranular similarity. Moreover, previous works require a large number of similar/dissimilar annotated pairwise poses, which is an tedious task and remains inaccurate due to different subjective judgments of experts. Motivated by graph-based neighbor assignment techniques, we propose an unsupervised model called sparsity locality preserving projection with adaptive neighbors (SLPPAN), for learning human pose distance metric. By using a property of the graph Laplacian, SLPPAN introduces a fixed-rank constraint to enforce an adaptive graph structure of poses and learns the neighbor assignment, the similarity measurement, and pose metric simultaneously. Experiments on pose retrieval of the CMU Mocap database demonstrate that SLPPAN outperforms traditional pose metric learning methods by capturing viewpoint variations of human poses. Experiments on keyframe extraction of the MSRAction3D database demonstrate that SLPPAN outperforms current methods by precisely detecting important frames of action sequences. Shaofan Wang 0001, Yongjia Xin, Dehui Kong |
IEEE Trans. Multim. | 3 |
| 2018 | A-optimal convolutional neural network
Zihong Yin, Dehui Kong, Guoxia Shao, Xinran Ning, Huidong Jin 0001 |
Neural Comput. Appl. | 2 |
| 2017 | Infrared dim target detection based on total variation regularization and principal component pursuit
Xiaoyang Wang 0005, Zhenming Peng, Dehui Kong, Ping Zhang 0023, Yanmin He |
Image Vis. Comput. | 3 |
| 2017 | Infrared Dim and Small Target Detection Based on Stable Multisubspace Learning in Heterogeneous SceneabstractInfrared (IR) dim and small target detection in a highly complex background play an important role in many applications, and remain a challenging problem. In this paper, a novel method named stable multisubspace learning is presented to deal with this problem. The new method takes into account the inner structure of actual images so that it overcomes the shortage of the traditional method. First, by analyzing the multisubspace structure of heterogeneous background data, a corresponding image model is proposed using subspace learning strategy. This model is also stable to noise interference. Second, an efficient optimization algorithm is designed to solve the proposed IR image model. By adding the proper postprocessing procedure, we can get the detection result. Experiments on simulation scenes and real scenes show that the proposed method has superior detection ability under heterogeneous background. Xiaoyang Wang 0005, Zhenming Peng, Dehui Kong, Yanmin He |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Extracting hand articulations from monocular depth images using curvature scale space descriptorsabstractWe propose a framework of hand articulation detection from a monocular depth image using curvature scale space (CSS) descriptors. We extract the hand contour from an input depth image, and obtain the fingertips and finger-valleys of the contour using the local extrema of a modified CSS map of the contour. Then we recover the undetected fingertips according to the local change of depths of points in the interior of the contour. Compared with traditional appearance-based approaches using either angle detectors or convex hull detectors, the modified CSS descriptor extracts the fingertips and finger-valleys more precisely since it is more robust to noisy or corrupted data; moreover, the local extrema of depths recover the fingertips of bending fingers well while traditional appearance-based approaches hardly work without matching models of hands. Experimental results show that our method captures the hand articulations more precisely compared with three state-of-the-art appearance-based approaches. Shaofan Wang 0001, Dehui Kong |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | Sparse Pose Regression via Componentwise Clustering Feature Point RepresentationabstractWe propose two-dimensional pose estimation from a single range image of the human body, using sparse regression with a componentwise clustering feature point representation (CCFPR) model. CCFPR includes primary feature points and secondary feature points. The primary feature points consist of the torso center and five extremal points of human body, and further serve to classify all body pixels as the points of six body components. The secondary feature points are given by the cluster centers of each of the five components other than the torso, using K-means cluster. The human pose is obtained by learning a sparse projection matrix, which maps CCFPR to the skeleton points of human body, based on the assumption that each skeleton point be represented by a combination of a few feature points of associated body components. Experimental results on both virtual data and real data show that, under the sparse regression model with a suitably selected cluster number, CCFPR outperforms the random decision forest approach and prediction results of Kinect sensor v2 . Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | Synthesis of sign language co-articulation based on key frames
Lichun Wang 0002, Dehui Kong |
Multim. Tools Appl. | 3 |
| 2015 | High-Resolution Light Field Capture With Coded ApertureabstractAcquiring light field with larger angular resolution and higher spatial resolution in low cost is the goal of light field capture. Combining or modifying traditional optical cameras is a usual method for designing light field capture equipment, among which most models should deliberate trade-off between angular and spatial resolution, but augmenting coded aperture avoids this consideration by multiplexing information from different views. On the basis of coded aperture, this paper suggests an improved light field camera model that has double measurements and one mask. The two compressive measurements are respectively realized by a coded aperture and a random convolution CMOS imager, the latter is used as imaging sensor of the camera. The single mask design permits high light efficiency, which enables the sampling images to have high clarity. The double measurement design keeps more correlation information, which is conductive to enhancing the reconstructed light field. The higher clarity and more correlation of samplings mean higher quality of rebuilt light field, which also means higher resolution under condition of a lower PSNR requirement for rebuilt light field. Experimental results have verified advantage of the proposed design: compared with the representative mask-based light field camera models, the proposed model has the highest reconstruction quality and a higher light efficiency. Lichun Wang 0002, Dehui Kong |
IEEE Trans. Image Process. | 3 |
| 2015 | Connectivity-preserving geometry images
Shaofan Wang 0001, Dehui Kong, Juan Xue, Weijia Zhu, Hubert Roth |
Vis. Comput. | 2 |
| 2014 | Image super-resolution based on multi-space sparse representation
Guodong Jing, Yunhui Shi, Dehui Kong, Wenpeng Ding |
Multim. Tools Appl. | 3 |
| 2014 | Chinese Sign Language animation generation considering context
Lichun Wang 0002, Dehui Kong |
Multim. Tools Appl. | 4 |
| 2014 | Adaptive particle shape setting and normal calculation methods in fluid rendering
Dehui Kong, Yong Zhang 0029 |
Multim. Tools Appl. | 2 |
| 2014 | Similarity Assessment Model for Chinese Sign Language VideosabstractThis paper proposes a model for measuring similarity between videos which content is Chinese Sign Language (CSL), vision and sign language semantic are considered for the model. Vision component of the model is distance based on Volume Local Binary Patterns (VLBP), which is robust for motion and illumination. Semantic component of the model computes semantic distance based on definition of sign language semantic, which is defined as hand shape, location, orientation and movements. While quantizing the sign language semantic, contour is used to measure shape and orientation; trajectory is used for measuring location and movement. Experiment results show that proposed assessment model is effective and assessing result given by the model is close to subjective scoring. Lichun Wang 0002, Dehui Kong |
IEEE Trans. Multim. | 3 |
| 2013 | A Real-Time Fluid Rendering Method with Adaptive Surface Smoothing and Realistic Splash
Yong Zhang 0029, Dehui Kong |
MMM (2) | 3 |
| 2012 | Making smooth transitions based on a multi-dimensional transition database for joining Chinese sign-language videos
Lichun Wang 0002, Dehui Kong |
Multim. Tools Appl. | 3 |
| 2011 | Fast mode dependent directional transform via butterfly-style transform and integer lifting steps
Wenpeng Ding, Ruiqin Xiong, Yunhui Shi, Dehui Kong |
J. Vis. Commun. Image Represent. | 4 |
| 2011 | Estimate of The Bézout Number For Linear Piecewise Algebraic Curves Over Arbitrary TriangulationsabstractA piecewise algebraic curve is a curve determined by the zero set of a bivariate spline function. This paper gives an upper bound of the Bézout number, the maximum number of intersections between two linear piecewise algebraic curves whose intersections are finite, over arbitrary triangulations. Shaofan Wang 0001, Renhong Wang 0001, Dehui Kong |
SIAM J. Discret. Math. | 3 |
| 2009 | An improved description method of the bumpy texture
Dehui Kong |
Sci. China Ser. F Inf. Sci. | 2 |
| 2007 | Adaptive Out-of-Core Simplification of Large Point CloudsabstractWith the increasing of data complexity, the needs for out-of-core simplification become evident. However, most of the existing point-based simplification algorithms adopt in-core scheme. We present an adaptive out-of-core algorithm for simplifying point-sampled models. Our approach uses quadric matrix to analyze the detailed regions of the initial simplified model that generated by an out-of-core uniform clustering. And then we use point-pair contraction to further simplify the flat regions and point-split to refine the detailed regions. Since the algorithm is input insensitive, it obtains high quality with low memory requirement. Dehui Kong |
ICME | 3 |
| 2004 | Audio to visual signal mappings with HMMabstractThere has been a large amount of research on speech driven face animation. Particularly, recently research efforts have been demonstrated that the hidden Markov model techniques could achieve a high level of success in the field of audio/visual mapping without language information. In this paper, firstly a linear model based facial representation method was applied, which extracts face features as a global feature. Secondly, a HMM based method was presented, which includes a two-level frame to promote the audio-visual mapping result. Xibin Jia, Dehui Kong |
ICASSP (5) | 4 |
| 2004 | Spatial prediction based intra-codingabstractAccording to the H.264 video coding standard, spatial prediction is used for intra block coding. The luma prediction may be based on a 4/spl times/4 block, for which there are nine prediction modes, or a 16/spl times/16 macroblock., for which there are four prediction modes. For chroma prediction, there are also four prediction modes. In this paper, a new method is proposed for improving the intra prediction algorithm, the first step is to apply direction to the results of the DC prediction and the second step is to use simplified modes to reduce computational complexity. Nan Zhang 0006, Dehui Kong, Wenying Yue |
ICME | 3 |
| 2003 | A New Facial Feature Extraction Method Based on Linear Combination Model abstractA new facial feature extraction method is proposed. Based on linear combination model, the method locates feature points in facial images precisely. The model uses the knowledge of prototypic faces to interpret novel faces. To get the knowledge, the prototypes are labeled manually on the feature points. Generally, the construction of the linear combination model depends on pixel-wise alignments of prototypes, and the alignments are computed by an optical flow algorithm or bootstrapping algorithm which is a full-scale optimization and not includes local information such as facial feature points. To combine local facial feature with the linear combination model, a restrained optical flow algorithm is proposed to compute the pixel-wise alignments. With the information of labeled feature points, the model matches the input facial images and extracts the feature points automatically. Implementing the feature extraction method on the MPI face database, the experimental results show that the method has good performance. Yongli Hu, Dehui Kong |
Web Intelligence | 3 |
| 2002 | An Improved Algorithm for Hairstyle DynamicsabstractThis paper introduces an efficient and flexible hair modeling method to develop intricate hairstyle dynamics. A prominent contribution of the present work is that it proposes an evaluation approach for the spring coefficient, i.e., spring coefficient can be obtained through a combination of the large deflection deformation model and spring hinge model. This is based on the fact that there is a directly proportional relationship between the spring coefficient and stiffness coefficient, a variable determined by hair shape. What is more, the damping coefficient is no longer regarded as a constant, but a function of hair density, and this treatment has turned out to be successful in solving the problem of hair-hair collision. As a result, a dynamic model, which fits a great variety of hairstyles, is proposed. Wenjun Lao, Dehui Kong |
ICMI | 2 |
| 2001 | Animating Human Face under Arbitrary IlluminationabstractThe authors present a method to animate a human face under arbitrary lighting conditions. We first acquire the reflectance models of a human face in various expressions by employing a robot to precisely control the light position. By developing a relighting tool, named LUXMASTER, we achieve great freedom to change the illumination condition and get a variety of photorealistic novel renderings of the human face. Then, we introduce morphing techniques to add the time dimension and produce a clip of facial animation with varying illumination and expression. In order to make the morphing work easily, we design an expressive, intuitive and efficient tool, called FLUIDMAN, which at the same time enables an enormous possibility of visual effects. Tong-Bo Chen, Wan-Jun Huang, Dehui Kong |
PG | 4 |