Zhongheng Li

dblp:190/4345 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-7091-9600ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 X-STA: Cross-Modal Spatial-Temporal Alignment Network for Unified Audio-Visual Segmentation
abstract
Audio-Visual Segmentation (AVS) aims to segment sound sources from video frames using synchronized audio cues. This task requires not only localizing the sound sources within frames but also accurately delineating their shapes. Existing AVS methods often rely on assumptions of spatial-temporal consistency between audio-visual content and are typically designed for specific learning paradigms. However, this specialization limits their ability to handle multi-granularity supervision signals and and adapt to diverse task requirements. For this purpose, we propose a Cross-modal Spatial-Temporal Alignment (X-STA) network to alleviate spatial-temporal inconsistency and overcome paradigm-specific constraint. Our X-STA introduces three key components: a novel multi-stage Cross-modal Adapter (xAdapter) that transfers knowledge from a pre-trained SAM through multi-grained representation adaptation, an innovative Cross-modal Prompter (xPrompter) that provides geometry-aware constraints for AVS through dynamic prompting strategies, and a Cross-modal Self-supervised (xSelf) mechanism that refines temporal alignment and enables self-supervised AVS. These components collectively facilitate explicit reasoning about location and geometric shape of the sound source by refining the alignment of cross-modal spatial-temporal cues. Our method achieves competitive performance across several baselines on widely-used AVS datasets, demonstrating its effectiveness in addressing the complexities of AVS.
Hanyu Xuan, Tongxing Liu, Wenxiang Dong, Zhongheng Li, Shuo Chen 0003
IEEE Signal Process. Lett.4
2024 Coordinate Descent Optimized Trace Difference Model for Joint Clustering and Feature Extraction
Fei Wang 0008, Zhongheng Li, Zheng Wang 0037, Feiping Nie 0001
Pattern Recognit.3
2023 Fine-grained Activities of People Worldwide
abstract
Every day, humans perform many closely related activities that involve subtle discriminative motions, such as putting on a shirt vs. putting on a jacket, or shaking hands vs. giving a high five. Activity recognition by ethical visual AI could provide insights into our patterns of daily life, however existing activity recognition datasets do not capture the massive diversity of these human activities around the world. To address this limitation, we introduce Collector, a free mobile app to record video while simultaneously annotating objects and activities of consented subjects. This new data collection platform was used to curate the Consented Activities of People (CAP) dataset, the first large-scale, fine-grained activity dataset of people worldwide. The CAP dataset contains 1.45M video clips of 512 fine grained activity labels of daily life, collected by 780 subjects in 33 countries. We provide activity classification and activity detection benchmarks for this dataset, and analyze baseline results to gain insight into how people around with world perform common activities. The dataset, benchmarks, evaluation tools, public leaderboards and mobile apps are available for use at https://visym.github.io/cap.
Jeffrey Byrne, Greg Castañón, Zhongheng Li, Gil J. Ettinger
WACV3
2023 Efficient random subspace decision forests with a simple probability dimensionality setting scheme
Fei Wang 0008, Zhongheng Li, Peilin Jiang, Fuji Ren, Feiping Nie 0001
Inf. Sci.3
2023 An Effective Clustering Optimization Method for Unsupervised Linear Discriminant Analysis
abstract
The recent work Unsupervised Linear Discriminant Analysis (Un-LDA) completes its clustering process during the alternating optimization by converting equivalently the objective and finally using the K-means algorithm. However, the K-means algorithm has its inherent drawbacks. It is hard for the K-means algorithm to deal well with some complex clustering cases where there are too many real clusters or non-convex clusters. In this paper, a novel clustering optimization method is presented to accomplish the clustering process in Un-LDA and the resulting method can be named Un-LDA(CD). Specifically, instead of the K-means algorithm, an elaborately designed coordinate descent algorithm is adopted to obtain the clusters after the objective function goes through a series of simple but deft equivalent conversions. Extensive experiments have demonstrated that the coordinate descent clustering solution for Un-LDA can outperform the original K-means based solution on the tested data sets especially those complex data sets with a pretty large number of real clusters.
Fei Wang 0008, Fuji Ren, Zhongheng Li, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.4
2023 Efficient Multi-View K-Means Clustering With Multiple Anchor Graphs
abstract
Multi-view clustering has attracted a lot of attention due to its ability to integrate information from distinct views, but how to improve efficiency is still a hot research topic. Anchor graph-based methods and k-means-based methods are two current popular efficient methods, however, both have limitations. Clustering on the derived anchor graph takes a while for anchor graph-based methods, and the efficiency of k-means-based methods drops significantly when the data dimension is large. To emphasize these issues, we developed an efficient multi-view k-means clustering method with multiple anchor graphs (EMKMC). It first constructs anchor graphs for each view and then integrates these anchor graphs using an improved k-means strategy to obtain sample categories without any extra post-processing. Since EMKMC combines the high-efficiency portions of anchor graph-based methods and k-means-based methods, its efficiency is substantially higher than current fast methods, especially when dealing with large-scale high-dimensional multi-view data. Extensive experiments demonstrate that, compared to other state-of-the-art methods, EMKMC can boost clustering efficiency by several to thousands of times while maintaining comparable or even exceeding clustering effectiveness.
Ben Yang, Xuetao Zhang 0001, Zhongheng Li, Feiping Nie 0001, Fei Wang 0008
IEEE Trans. Knowl. Data Eng.3
2021 Flexible multi-view semi-supervised learning with unified graph
Zhongheng Li, Qianyao Qiang, Bin Zhang 0022, Fei Wang 0008, Feiping Nie 0001
Neural Networks1
2021 Unsupervised Linear Discriminant Analysis for Jointly Clustering and Subspace Learning
abstract
Linear discriminant analysis (LDA) is one of commonly used supervised subspace learning methods. However, LDA will be powerless faced with the no-label situation. In this paper, the unsupervised LDA (Un-LDA) is proposed and first formulated as a seamlessly unified objective optimization which guarantees convergence during the iteratively alternative solving process. The objective optimization is in both the ratio trace and the trace ratio forms, forming a complete framework of a new approach to jointly clustering and unsupervised subspace learning. The extension of LDA into Un-LDA enables to not only complete unsupervised subspace learning via the explicitly presented subspace projection matrix but also simultaneously finish clustering and even clustering out-of-sample data via the explicitly presented transformation matrix. To overcome the difficulty in solving the non-convex objective optimization, we mathematically prove that the Un-LDA optimization in both forms can be transformed into the simple K-means clustering optimization when the subspace is determined. The Un-LDA optimization is eventually completed by alternatively optimizing the clusters using K-means and the subspace using the supervised LDA methods and iterating this whole process until convergence or stopping criterion. The experiments demonstrate that our proposed Un-LDA algorithms are comparable or even much superior to the counterparts.
Fei Wang 0008, Feiping Nie 0001, Zhongheng Li, Weizhong Yu, Rong Wang 0001
IEEE Trans. Knowl. Data Eng.4
2020 Robust and Scalable Entity Alignment in Big Data
abstract
Entity alignment has always had significant uses within a multitude of diverse scientific fields. In particular, the concept of matching entities across networks has grown in significance in the world of social science as communicative networks such as social media have expanded in scale and popularity. With the advent of big data, there is a growing need to provide analysis on graphs of massive scale. However, with millions of nodes and billions of edges, the idea of alignment between a myriad of graphs of similar scale using features extracted from potentially sparse or incomplete datasets becomes daunting. In this paper we will propose a solution to the issue of large-scale alignments in the form of a multi-step pipeline. Within this pipeline we introduce scalable feature extraction for robust temporal attributes, accompanied by novel and efficient clustering algorithms in order to find groupings of similar nodes across graphs. The features and their clusters are fed into a versatile alignment stage that accurately identifies partner nodes among millions of possible matches. Our results show that the pipeline can process large data sets, achieving efficient runtimes within the memory constraints.
James Flamino, Christopher Abriola, Benjamin Zimmerman, Zhongheng Li, Joel Douglas
IEEE BigData4
2020 A forest of trees with principal direction specified oblique split on random subspace
Fei Wang 0008, Feiping Nie 0001, Weizhong Yu, Rong Wang 0001, Zhongheng Li
Neurocomputing6
2020 A linear multivariate binary decision tree classifier based on K-means splitting
Fei Wang 0008, Feiping Nie 0001, Zhongheng Li, Weizhong Yu, Fuji Ren
Pattern Recognit.4