Maodi Hu

dblp:01/7951 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
8since 2021 · last 2025
0009-0003-2207-5134ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Domain-level relation extraction for informative taxonomy learning
Maodi Hu, Donghuan Song, Zhixiong Zhang 0002
Data Min. Knowl. Discov.1
2024 KDPG-Enhanced MRC Framework for Scientific Entity Recognition in Survey Papers
abstract
Scientific survey papers play a pivotal role in advancing knowledge and scientific progress by providing concise summaries and analyses of research trends and findings. To facilitate better knowledge organization and analysis, we have undertaken the challenge of defining the scientific entity recognition task for survey papers and carefully curated a dataset that closely emulates real-world scenarios. The scientific entity recognition task presents unique challenges, including multi-label, low-resource, and nested scenarios. To address these challenges, we propose a unified framework based on the machine reading comprehension (MRC) paradigm. This framework not only supports nested and multi-label settings but also enables the effective transfer of information from high-resource categories to low-resource ones, ensuring adaptability and robustness. To further enhance performance, we introduce the Knowledge-Driven Prototype Guidance (KDPG) module, seamlessly integrated into a two-phase learning strategy. The KDPG module leverages prior knowledge and acts as an initial prototype-based manifold constraint, effectively harnessing the power of few-shot learning capabilities. Through this integration, our approach complements the classification learning tasks for entity recognition, resulting in improved accuracy and efficiency. Our experimental results validate the effectiveness of the proposed KDPG-enhanced MRC framework, showcasing its leading performance on publicly available datasets and our collected scientific survey paper dataset.
Maodi Hu, Zhijun Chang, Zhixiong Zhang 0002
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Confusion Region Mining for Crowd Counting
abstract
Existing works mainly focus on crowd and ignore the confusion regions which contain extremely similar appearance to crowd in the background, while crowd counting needs to face these two sides at the same time. To address this issue, we propose a novel end-to-end trainable confusion region discriminating and erasing network called CDENet. Specifically, CDENet is composed of two modules of confusion region mining module (CRM) and guided erasing module (GEM). CRM consists of basic density estimation (BDE) network, confusion region aware bridge and confusion region discriminating network. The BDE network first generates a primary density map, and then the confusion region aware bridge excavates the confusion regions by comparing the primary prediction result with the ground-truth density map. Finally, the confusion region discriminating network learns the difference of feature representations in confusion regions and crowds. Furthermore, GEM gives the refined density map by erasing the confusion regions. We evaluate the proposed method on four crowd counting benchmarks, including ShanghaiTech Part_A, ShanghaiTech Part_B, UCF_CC_50, and UCF-QNRF, and our CDENet achieves superior performance compared with the state-of-the-arts.
Jiawen Zhu 0003, Wenda Zhao 0003, Libo Yao, You He 0002, Maodi Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.5
2022 Video Object Segmentation via Structural Feature Reconfiguration
Zhenyu Chen 0001, Ping Hu 0001, Lu Zhang 0053, Huchuan Lu, You He 0002, Maodi Hu
ACCV (7)8
2022 Gated Hypergraph Neural Network for Scene-Aware Recommendation
Tianchi Yang, Luhao Zhang, Chuan Shi 0001, Cheng Yang 0002, Siyong Xu, Ruiyu Fang, Maodi Hu, Huaijun Liu, Dong Wang 0022
DASFAA (2)7
2022 A Joint Framework for Explainable Recommendation with Knowledge Reasoning and Graph Representation
Luhao Zhang, Ruiyu Fang, Tianchi Yang, Maodi Hu, Chuan Shi 0001, Dong Wang 0022
DASFAA (3)4
2022 Co-clustering Interactions via Attentive Hypergraph Neural Network
abstract
With the rapid growth of interaction data, many clustering methods have been proposed to discover interaction patterns as prior knowledge beneficial to downstream tasks. Considering that an interaction can be seen as an action occurring among multiple objects, most existing methods model the objects and their pair-wise relations as nodes and links in graphs. However, they only model and leverage part of the information in real entire interactions, i.e., either decompose the entire interaction into several pair-wise sub-interactions for simplification, or only focus on clustering some specific types of objects, which limits the performance and explainability of clustering. To tackle this issue, we propose to Co-cluster the Interactions via Attentive Hypergraph neural network (CIAH). Particularly, with more comprehensive modeling of interactions by hypergraph, we propose an attentive hypergraph neural network to encode the entire interactions, where an attention mechanism is utilized to select important attributes for explanations. Then, we introduce a salient method to guide the attention to be more consistent with real importance of attributes, namely saliency-based consistency. Moreover, we propose a novel co-clustering method to perform a joint clustering for the representations of interactions and the corresponding distributions of attribute selection, namely cluster-based consistency. Extensive experiments demonstrate that our CIAH significantly outperforms state-of-the-art clustering methods on both public datasets and real industrial datasets.
Tianchi Yang, Cheng Yang 0002, Luhao Zhang, Chuan Shi 0001, Maodi Hu, Huaijun Liu, Dong Wang 0022
SIGIR5
2021 Topic-aware Heterogeneous Graph Neural Network for Link Prediction
abstract
Heterogeneous graphs (HGs), consisting of multiple types of nodes and links, can characterize a variety of real-world complex systems. Recently, heterogeneous graph neural networks (HGNNs), as a powerful graph embedding method to aggregate heterogeneous structure and attribute information, has earned a lot of attention. Despite the ability of HGNNs in capturing rich semantics which reveal different aspects of nodes, they still stay at a coarse-grained level which simply exploits structural characteristics. In fact, rich unstructured text content of nodes also carries latent but more fine-grained semantics arising from multi-facet topic-aware factors, which fundamentally manifest why nodes of different types would connect and form a specific heterogeneous structure. However, little effort has been devoted to factorizing them.
Siyong Xu, Cheng Yang 0002, Chuan Shi 0001, Yuan Fang 0001, Tianchi Yang, Luhao Zhang, Maodi Hu
CIKM8
2013 Cross-View Gait Recognition with Short Probe Sequences: from View Transformation Model to View-Independent stance-Independent Identity Vector
abstract
Considering it is difficult to guarantee that at least one continuous complete gait cycle is captured in real applications, we address the multi-view gait recognition problem with short probe sequences. With unified multi-view population hidden markov models (umvpHMMs), the gait pattern is represented as fixed-length multi-view stances. By incorporating the multi-stance dynamics, the well-known view transformation model (VTM) is extended into a multi-linear projection model in a four-order tensor space, so that a view-independent stance-independent identity vector (VSIV) can be extracted. The main advantage is that the proposed VSIV is stable for each subject regardless of the camera location or the sequence length. Experiments show that our algorithm achieves encouraging performance for cross-view gait recognition even with short probe sequences.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001
Int. J. Pattern Recognit. Artif. Intell.1
2013 Estimation of view angles for gait using a robust regression method
Yunhong Wang 0001, Zhaoxiang Zhang 0001, Maodi Hu
Multim. Tools Appl.4
2013 Incremental Learning for Video-Based Gait Recognition With LBP Flow
abstract
Gait analysis provides a feasible approach for identification in intelligent video surveillance. However, the effectiveness of the dominant silhouette-based approaches is overly dependent upon background subtraction. In this paper, we propose a novel incremental framework based on optical flow, including dynamics learning, pattern retrieval, and recognition. It can greatly improve the usability of gait traits in video surveillance applications. Local binary pattern (LBP) is employed to describe the texture information of optical flow. This representation is called LBP flow, which performs well as a static representation of gait movement. Dynamics within and among gait stances becomes the key consideration for multiframe detection and tracking, which is quite different from existing approaches. To simulate the natural way of knowledge acquisition, an individual hidden Markov model (HMM) representing the gait dynamics of a single subject incrementally evolves from a population model that reflects the average motion process of human gait. It is beneficial for both tracking and recognition and makes the training process of the HMM more robust to noise. Extensive experiments on widely adopted databases have been carried out to show that our proposed approach achieves excellent performance.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little
IEEE Trans. Cybern.1
2013 View-Invariant Discriminative Projection for Multi-View Gait-Based Human Identification
abstract
Existing methods for multi-view gait-based identification mainly focus on transforming the features of one view to the features of another view, which is technically sound but has limited practical utility. In this paper, we propose a view-invariant discriminative projection (ViDP) method, to improve the discriminative ability of multi-view gait features by a unitary linear projection. It is implemented by iteratively learning the low dimensional geometry and finding the optimal projection according to the geometry. By virtue of ViDP, the multi-view gait features can be directly matched without knowing or estimating the viewing angles. The ViDP feature projected from gait energy image achieves promising performance in the experiments of multi-view gait-based identification. We suggest that it is possible to construct a gait-based identification system for arbitrary probe views, by incorporating the information of gallery data with sufficient viewing angles. In addition, ViDP performs even better than the state-of-the-art view transformation methods, which are trained for the combination of gallery and probe viewing angles in every evaluation.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little, Di Huang 0001
IEEE Trans. Inf. Forensics Secur.1
2012 Combinational Subsequence Matching for Human Identification from General Actions
Maodi Hu, Yunhong Wang 0001, James J. Little
ACCV (3)1
2011 Multi-view multi-stance gait identification
abstract
View transformation in gait analysis has attracted more and more attentions recently. However, most of the existing methods are based on the entire gait dynamics, such as Gait Energy Image (GEI). And the distinctive characteristics of different walking phases are neglected. This paper proposes a multi-view multi-stance gait identification method using unified multi-view population Hidden Markov Models (pHMM-s), in which all the models share the same transition probabilities. Hence, the gait dynamics in each view can be normalized into fixed-length stances by Viterbi decoding. To optimize the view-independent and stance-independent identity vector, a multi-linear projection model is learned from tensor decomposition. The advantage of using tensor is that different types of information are integrated in the final optimal solution. Extensive experiments show that our algorithm achieves promising performances of multi-view gait identification even with incomplete gait cycles.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001
ICIP1
2011 Gait-Based Gender Classification Using Mixed Conditional Random Field
abstract
This paper proposes a supervised modeling approach for gait-based gender classification. Different from traditional temporal modeling methods, male and female gait traits are competitively learned by the addition of gender labels. Shape appearance and temporal dynamics of both genders are integrated into a sequential model called mixed conditional random field (CRF) (MCRF), which provides an open framework applicable to various spatiotemporal features. In this paper, for the spatial part, pyramids of fitting coefficients are used to generate the gait shape descriptors; for the temporal part, neighborhood-preserving embeddings are clustered to allocate the stance indexes over gait cycles. During these processes, we employ evaluation functions like the partition index and Xie and Beni's index to improve the feature sparseness. By fusion of shape descriptors and stance indexes, the MCRF is constructed in coordination with intra- and intergender temporary Markov properties. Analogous to the maximum likelihood decision used in hidden Markov models (HMMs), several classification strategies on the MCRF are discussed. We use CASIA (Data set B) and IRIP Gait Databases for the experiments. The results show the superior performance of the MCRF over HMMs and separately trained CRFs.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001
IEEE Trans. Syst. Man Cybern. Part B1
2010 Combining Spatial and Temporal Information for Gait Based Gender Classification
abstract
In this paper, we address the problem of gait based gender classification. The Gabor feature which is a new attempt for gait analysis, not only improves the robustness to the segmental noise, but also provides a feasible way to purge the additional influence factors like clothing and carrying condition changes before supervised learning. Furthermore, through the agency of Maximization of Mutual Information (MMI), the low dimensional discriminative representation is obtained as the Gabor-MMI feature. After that, gender related Gaussian Mixture Model-Hidden Markov Models (GMM-HMMs) are constructed for classification work. In this case, supervised learning reduces the dimension of parameter space, and significantly increases the gap between likelihoods of the gender models. In order to assess the performance of our proposed approach, we compare it with other methods on the standard CASIA Gait Databases (Dataset B). Experimental results demonstrate that our approach achieves better Correct Classification Rate (CCR) than the state of the art methods.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001
ICPR1
2009 A New Approach for Gender Classification Based on Gait Analysis
abstract
In this paper, we propose a novel pattern to represent spatio-temporal information of gait appearance which is called Gait Principal Component Image (GPCI). GPCI is a grey-level image which compresses the spatiotemporal information by amplifying the dynamic variation of different body part. The detection of gait period is based on LLE coefficients and it is also a new attempt. KNN classifier is employed for gender classification. The framework can be applied in real-time setting because of its rapidity and robustness. The experimental results on IRIP Gait Database (32 males, 28 females) show that the proposed approach achieves a high accuracy in automatic gender classification.
Maodi Hu, Yunhong Wang 0001
ICIG1