VLDB 2026 Research / reviewers in the wild / expert
Yameng Liu
dblp:222/3214
· DBLP profile ↗
19ranked-venue papers
1as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFDiff: Diffusion probabilistic model for medical image segmentation with multi-scale features and frequency-aware attention
Xing-Li Zhang 0001, Yameng Liu, Zhihui Wang 0003 |
Comput. Vis. Image Underst. | 2 |
| 2026 | Evolution Rather Than Degradation: Structure-Guided Elastic Consensus Learning for Multimodal Knowledge Graph Completion
Yameng Liu, Shuai Zheng 0005, Zhenfeng Zhu, Yunhui Xu, Yao Zhao 0001, Kunlun He |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | A Novel Environment Object Modeling Method for Vehicular ISAC ScenariosabstractIntegrated Sensing and Communication (ISAC), as a fundamental technology of 6G, empowers Vehicle-to-Everything (V2X) systems with enhanced sensing capabilities. One of its promising applications is the reliance on constructed maps for vehicle positioning. Traditional positioning methods primarily rely on Line-of-Sight (LOS), but in urban vehicular scenarios, obstructions often result in predominantly Non-Line-of-Sight (NLOS) conditions. Existing researched indicate that NLOS paths, characterized by one-bounce reflection on building wall with determined delay and angle, can support sensing and positioning. However, experimental validation remains insufficient. To address this gap, channel measurements are conducted in an urban street to explore the existence of strong reflected paths in the presence of a vehicle target. The results show significant power contribution from NLOS paths, with large Environmental Objects (EOs) playing a key role in shaping NLOS propagation. Then, a novel model for EO reflection is proposed to extend the Geometry-Based Stochastic Model (GBSM) for ISAC channel standardization. Simulation results validate the model's ability to capture EO's power and position characteristics, showing that higher EO-reflected power and closer distance to Rx reduce Delay Spread (DS), which is more favorable for positioning. This model provides theoretical guidance and empirical support for ISAC positioning algorithms and system design in vehicular scenarios. Hanyuan Jiang, Yuxiang Zhang 0002, Yameng Liu, Jianhua Zhang 0001, Lei Tian 0004, Tao Jiang 0025 |
WCNC | 3 |
| 2024 | Contextual Correspondence Matters: Bidirectional Graph Matching for Video Summarization
Yunzuo Zhang, Yameng Liu |
ECCV (87) | 2 |
| 2024 | M2SUM: Multi-Granularity Scale-Adaptive Video Summarizer towards Informative Context Representation LearningabstractVideo summarization intends to automatically select meaningful segments from untrimmed videos. Although previous efforts have achieved remarkable progress, they still struggle to robustly aggregate and effectively process multi-granularity contextual information within videos, which hinders understanding towards video content. To address these issues, we propose M2SUM, which is composed of three dominant components including the embedding learning attention (ELA) module, multi-granularity aggregator (MGA), and semantic scale-adaption (SSA) module. ELA dynamically enhances pre-trained visual features by considering the similarity relationship across frame-level and video-level embeddings. MGA incorporates self-attention and temporal convolution into a unified learnable module, robustly learning long-range and short-range multi-granularity temporal dependencies. SSA is exploited to adaptively perform representation fusion after deep interaction across multi-granularity temporal dependencies. According to the fused representations, M2SUM predicts importance scores and generates video summaries. Extensive experiments on standard datasets have proved the effectiveness and superiority of our method in F-score and rank-based evaluations. Yunzuo Zhang, Yameng Liu, Weili Kang |
ICASSP | 2 |
| 2024 | ECPNet: An Enhanced Curve Perception Network for Lane DetectionabstractLane detection methods based on anchors have received increasing attention, but fixed-shape anchors make it difficult to model complex lane line shapes. To solve this problem, we propose an Enhanced Curve Perception Network (ECPNet). Specifically, we propose a Layer-by-layer Context Fusion (LCF) module to fully utilize both high-level and low-level features in lane detection by establishing short hop connections across feature layers of diverse scales. Then, we propose a novel Structural Correction Prediction (SCP) module, which enhances the detection ability of the model on the curve structure lane by dynamically guiding the selection of anchor classification patterns. In addition, ECPNet adaptive calibration pays attention to channel features through the Cross-Channel Attention (CCA) mechanism. Experiments on the two most representative datasets demonstrate that the proposed method achieves state-of-the-art performance in a variety of environments, especially in curvy lanes. Yunzuo Zhang, Cunyu Wu, Yameng Liu |
ICASSP | 5 |
| 2024 | Multi-Frequency Channel Measurement in Smart Factories: Comparative Analysis with 3GPP InF ScenariosabstractSmart factory scenarios in the Industrial Internet of Things (IIoT) introduce novel challenges to wireless communication, characterized by large-scale device interaction, strong reflection, and dense scattering, diverging from typical factories. To ascertain the channel disparities between smart factories and typical factories, we analyze channel characteristics across sub-6 GHz (e.g., 6 GHz) and mmWave frequencies (e.g., 28 and 39 GHz) in smart factory scenarios, establishing a multi-frequency path loss model. Furthermore, the comparison between the large-scale fading characteristics, including path loss and root mean square delay spread (RMS DS), observed in smart factory scenarios and those in the typical indoor factory (InF) scenarios defined by the 3rd-Generation Partnership Project (3GPP). Additionally, channel characteristics of the metal and non-metal scatterer areas in the smart factories are extracted to analyze the impact of metal scatterers on angle spread (AS) and RMS DS. Qingmei Mo, Yuxiang Zhang 0002, Jialin Wang 0001, Yameng Liu, Zeyong Chai, Huiwen Gong, Jianhua Zhang 0001 |
VTC Fall | 4 |
| 2024 | Latest progress for 3GPP ISAC channel modeling standardization
Yuanpeng Pei, Yameng Liu |
Sci. China Inf. Sci. | 4 |
| 2024 | Attention-guided multi-granularity fusion model for video summarization
Yunzuo Zhang, Yameng Liu, Cunyu Wu |
Expert Syst. Appl. | 2 |
| 2024 | VSS-Net: Visual Semantic Self-Mining Network for Video SummarizationabstractVideo summarization, with the target to detect valuable segments given untrimmed videos, is a meaningful yet understudied topic. Previous methods primarily consider inter-frame and inter-shot temporal dependencies, which might be insufficient to pinpoint important content due to limited valuable information that can be learned. To address this limitation, we elaborate on a Visual Semantic Self-mining Network (VSS-Net), a novel summarization framework motivated by the widespread success of cross-modality learning tasks. VSS-Net initially adopts a two-stream structure consisting of a Context Representation Graph (CRG) and a Video Semantics Encoder (VSE). They are jointly exploited to establish the groundwork for further boosting the capability of content awareness. Specifically, CRG is constructed using an edge-set strategy tailored to the hierarchical structure of videos, enriching visual features with local and non-local temporal cues from temporal order and visual relationship perspectives. Meanwhile, by learning visual similarity across features, VSE adaptively acquires an instructive video-level semantic representation of the input video from coarse to fine. Subsequently, the two streams converge in a Context-Semantics Interaction Layer (CSIL) to achieve sophisticated information exchange across frame-level temporal cues and video-level semantic representation, guaranteeing informative representations and boosting the sensitivity to important segments. Eventually, importance scores are predicted utilizing a prediction head, followed by key shot selection. We evaluate the proposed framework and demonstrate its effectiveness and superiority against state-of-the-art methods on the widely used benchmarks. Yunzuo Zhang, Yameng Liu, Weili Kang, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Joint Multi-Level Feature Network for Lightweight Person Re-IdentificationabstractLearning fine-grained features is crucial to the performance improvement of person re-identification (Re-ID). Although existing methods have made significant progress, utilizing multi-level information to obtain fine-grained features has not been explored in this field. To alleviate this issue, we propose a lightweight person Re-ID method named Joint Multi-Level Feature Network (JMLFNet) to obtain robust feature representation for the Re-ID task. Specifically, we design a Multi-Attention Block (MAB) and embed it into the lightweight backbone network to improve performance, which can make the network focus on the key parts of pedestrian images. Meanwhile, we propose a Multi-Level Feature Extraction (MLFE) method to extract multi-granularity features of high-level semantic information and low-level detail information, which can effectively capture the feature diversity of pedestrian images. Furthermore, we design a Feature Fusion Block (FFB), which is fused the fine-grained features of high-level and low-level information to better obtain the discriminative feature representation of pedestrian images. Extensive experiments conducted on popular datasets Market1501 and DukeMTMC-reID demonstrate that the proposed JMLFNet has competitive performance compared with the state-of-the-art methods. Yunzuo Zhang, Weili Kang, Yameng Liu, Pengfei Zhu 0005 |
ICASSP | 3 |
| 2023 | Multi-Scale Semantic and Detail Extraction Network for Lightweight Person Re-Identification
Yunzuo Zhang, Weili Kang, Yameng Liu, Pengfei Zhu 0005 |
Comput. Vis. Image Underst. | 3 |
| 2023 | Key frame extraction based on quaternion Fourier transform with multiple features fusion
Yunzuo Zhang, Ruixue Liu, Pengfei Zhu 0005, Yameng Liu |
Expert Syst. Appl. | 5 |
| 2023 | Self-Attention Guidance and Multiscale Feature Fusion-Based UAV Image Object DetectionabstractObject detection on UAV images is a recent research hotspot. Existing object detection methods have achieved good results on general scenes, but there are inherent challenges with UAV images. The detection accuracy of UAV images is limited by complex backgrounds, significant scale differences, and densely arranged small objects. To solve these problems, we propose a UAV image object detection network based on Self-attention Guidance and Multi-scale Feature fusion (SGMFNet). Firstly, we design a Global-Local Feature Guidance module (GLFG). This module can effectively combine local information and global information, which makes the model focus on the object area and reduces the impact of complex background. Secondly, an improved Parallel Sampling Feature Fusion module (PSFF) is designed to efficiently fuse multi-scale features. Thirdly, we design an Inverse-residual Feature Enhancement module (IFE), which is embedded in the front of the newly added detection head to enhance feature extraction on small objects. Finally, we conduct a large number of experiments on the VisDrone2019 dataset. The results show that the proposed SGMFNet outperforms other popular methods, and has achieved good results in many scenarios. Yunzuo Zhang, Cunyu Wu, Tian Zhang 0008, Yameng Liu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | MAR-Net: Motion-Assisted Reconstruction Network for Unsupervised Video SummarizationabstractVideo summarization targets to extract the most important segments from a video by spatiotemporal analysis. Previous methods primarily learn content within videos based on appearance information, with a rare discussion on the effective utilization of motion information, which is equally essential to video understanding. In this letter, we expound upon a Motion-Assisted Reconstruction Network (MAR-Net), which synergistically models appearance and motion information within videos for unsupervised video summarization without any manual annotations. MAR-Net notably comprises a Bidirectional Modality Encoder (BiME) and a Video Context Navigator (VCN). By integrating uni-modal and cross-modal feature aggregation into a unified module, BiME allows for exploring sophisticated dependency relationships among features through a bidirectional attention mechanism. VCN can promote the semantic consistency between the cross-modal contexts and the input video by a consistency loss term, alleviating the noisy impact within the motion stream. Empirical results conducted on benchmark datasets demonstrate that MAR-Net outperforms other state-of-the-art methods. Yunzuo Zhang, Yameng Liu, Weili Kang |
IEEE Signal Process. Lett. | 2 |
| 2022 | Joint Reinforcement and Contrastive Learning for Unsupervised Video SummarizationabstractThis letter presents a joint Reinforcement and Contrastive Learning framework termed RCL for unsupervised video summarization, aiming at addressing the existing two shortcomings: (i) poor feature representation, and (ii) inefficient context modeling capability. Concretely, the proposed framework consists of an Optimized Coding Module (OCM) and a Dissimilarity-Guided Attention Graph (DGAG). The OCM is grounded on Gate Recurrent Unit (GRU), which encodes the content within shots into concise representations. Different from the existing approaches, contrastive learning is introduced to promote discriminative and informative feature learning. Afterward, the DGAG adaptively performs feature aggregation by evaluating the semantic dissimilarity across shots to eliminate chaotic message passing for accurate context modeling. Finally, the procedure of importance score prediction is formulated as a node classification task, and these scores are utilized for a summary generation. Extensive experiments on the benchmark datasets demonstrate the superior performance of the proposed method. Yunzuo Zhang, Yameng Liu, Pengfei Zhu 0005, Weili Kang |
IEEE Signal Process. Lett. | 2 |
| 2021 | FLAME: A Fast Large-scale Almost Matching Exactly Approach to Causal InferenceabstractA classical problem in causal inference is that of matching, where treatment units need to be matched to control units based on covariate information. In this work, we propose a method that computes high quality almost-exact matches for high-dimensional categorical datasets. This method, called FLAME (Fast Large-scale Almost Matching Exactly), learns a distance metric for matching using a hold-out training data set. In order to perform matching efficiently for large datasets, FLAME leverages techniques that are natural for query processing in the area of database management, and two implementations of FLAME are provided: the first uses SQL queries and the second uses bit-vector techniques. The algorithm starts by constructing matches of the highest quality (exact matches on all covariates), and successively eliminates variables in order to match exactly on as many variables as possible, while still maintaining interpretable high-quality matches and balance between treatment and control groups. We leverage these high quality matches to estimate conditional average treatment effects (CATEs). Our experiments show that FLAME scales to huge datasets with millions of observations where existing state-of-the-art methods fail, and that it achieves significantly better performance than other matching methods. Tianyu Wang 0008, Marco Morucci, M. Usaid Awan, Yameng Liu, Sudeepa Roy 0001, Cynthia Rudin, Alexander Volfovsky |
J. Mach. Learn. Res. | 4 |
| 2019 | Interpretable Almost-Exact Matching for Causal InferenceabstractMatching methods are heavily used in the social and health sciences due to their interpretability. We aim to create the highest possible quality of treatment-control matches for categorical data in the potential outcomes framework. The method proposed in this work aims to match units on a weighted Hamming distance, taking into account the relative importance of the covariates; the algorithm aims to match units on as many relevant variables as possible. To do this, the algorithm creates a hierarchy of covariate combinations on which to match (similar to downward closure), in the process solving an optimization problem for each unit in order to construct the optimal matches. The algorithm uses a single dynamic program to solve all of the units’ optimization problems simultaneously. Notable advantages of our method over existing matching procedures are its high-quality interpretable matches, versatility in handling different data distributions that may have irrelevant variables, and ability to handle missing data by matching on as many available covariates as possible. Awa Dieng, Yameng Liu, Sudeepa Roy 0001, Cynthia Rudin, Alexander Volfovsky |
AISTATS | 2 |
| 2019 | Interpretable Almost Matching Exactly With Instrumental Variables
M. Usaid Awan, Yameng Liu, Marco Morucci, Sudeepa Roy 0001, Cynthia Rudin, Alexander Volfovsky |
UAI | 2 |