Mengyu Yang

dblp:284/0734 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoPHo: Classifier-guided Conditional Topology Generation with Persistent Homology
abstract
The structure of topology underpins much of the research on performance and robustness, yet available topology data are typically scarce, necessitating the generation of synthetic graphs with desired properties for testing or release. Prior diffusion-based approaches either embed conditions into the diffusion model, requiring retraining for each attribute and hindering real-time applicability, or use classifier-based guidance post-training, which does not account for topology scale and practical constraints. In this paper, we show from a discrete perspective that gradients from a pre-trained graph?level classifier can be incorporated into the discrete reverse diffusion posterior to steer generation toward specified structural properties. Based on this insight, we propose Classifier-guided Conditional Topology Generation with Persistent Homology (CoPHo), which builds a persistent homology filtration over intermediate graphs and interprets features as guidance signals that steer generation toward the desired properties at each denoising step. Experiments on four generic/network datasets demonstrate that CoPHo outperforms existing methods at matching target metrics, and we further validate its transferability on the QM9 molecular dataset. The code is available at https://github.com/Lrbomchz/CoPHo.
Gongli Xi, Ye Tian 0008, Mengyu Yang, Yuchao Zhang 0004, Xiangyang Gong, Xirong Que, Wendong Wang 0003
KDD (1)3
2026 SEC: Enabling MLLMs for Low-Latency IoT Video Analysis via Semantic-Aware Edge-Cloud Collaboration
abstract
The rapid proliferation of IoT-enabled cameras has driven increasing demand for low-latency, intelligent video understanding in real-world applications such as smart cities and industrial automation. While Multimodal Large Language Models (MLLMs) offer unprecedented capabilities in semantic reasoning and natural language-based video comprehension, their deployment in latency-sensitive IoT environments remains challenging due to high computational costs and sequential decoding bottlenecks. Moreover, conventional edge-cloud video analysis frameworks often rely on semantic-agnostic frame sampling, leading to information loss or redundant data transmission. In this paper, we proposeSEC, a semantic-aware edge-cloud collaborative framework for efficient and accurate video analysis.SECintroduces a task-aware key frame selection mechanism at the edge to maximize semantic relevance while minimizing bandwidth usage, and a novel adaptive speculative decoding framework with tree-based parallel generation on the cloud to accelerate MLLM inference. Extensive experiments under realistic edge-cloud deployment settings on four video understanding benchmarks demonstrate that the proposedSECframework achieves superior performance, significantly reducing the end-to-end inference latency while improving the accuracy, a rare win-win in latency-critical IoT systems.
Mengyu Yang, Ye Tian 0008, Peizhuang Cong, Lanshan Zhang, Gongli Xi, Song Wang 0006, Wendong Wang 0003
IEEE Internet Things J.1
2025 Celestial Equilibrium Theory-Based Optimal Deployment of SRv6 for Traffic Engineering
abstract
Segment Routing over IPv6 (SRv6) is a promising source routing solution with wide-ranging applications in the field of Traffic Engineering. By adding SR tags into IPv6 packets, traffic can be directed to various SR segments, effectively distributing traffic. However, upgrading all network nodes to SRv6 nodes simultaneously is impractical. This paper addresses the incremental deployment of SRv6 from traffic engineering, aiming to minimize MLU within the network. We introduce the theory of celestial equilibrium and model the problem as a celestial equilibrium-like model with a global distribution of SR nodes and their corresponding areas of influence. To address this problem model, we propose a novel algorithm based on the EM algorithm, EM-SRTE. In our proposed framework, step E leverages a reinforcement learning algorithm combined with a self-attention module for graph learning to optimize the selection of SR nodes. Meanwhile, step M utilizes a similar algorithm to optimize the region range of SR nodes. Experimental results using publicly available datasets demonstrate that our model outperforms state-of-the-art baselines.
Ye Tian 0008, Yuan Yang 0001, Mengyu Yang, Wendong Wang 0003, Xiangyang Gong
ICC4
2025 Clink! Chop! Thud! - Learning Object Sounds From Real-World Interactions
abstract
Can a model distinguish between the sound of a spoon hitting a hardwood floor versus a carpeted one? Everyday object interactions produce sounds unique to the objects involved. We introduce the sounding object detection task to evaluate a model's ability to link these sounds to the objects directly involved. Inspired by human perception, our multimodal object-aware framework learns from in-the-wild egocentric videos. To encourage an object-centric approach, we first develop an automatic pipeline to compute segmentation masks of the objects involved to guide the model's focus during training towards the most informative regions of the interaction. A slot attention visual encoder is used to further enforce an object prior. We demonstrate state of the art performance on our new task along with existing multimodal action understanding tasks.
Mengyu Yang, Haozheng Pei, Siddhant Agarwal, Arun Balajee Vasudevan, James Hays
ICCV1
2025 SDP: Spiking Diffusion Policy for Robotic Manipulation with Learnable Channel-Wise Membrane Thresholds
Zhixing Hou, Maoxu Gao, Mengyu Yang, Chio-In Ieong
PRCV (11)4
2025 REM: Enabling Real-Time Neural-Enhanced Video Streaming on Mobile Devices Using Macroblock-Aware Lookup Table
abstract
The demand for mobile video streaming has seen a substantial surge in recent years. However, current platforms heavily depend on network capacity to ensure the delivery of high-quality video streams. The emergence of neural-enhanced video streaming presents a promising solution to address this challenge by leveraging client-side computation, thereby reducing bandwidth consumption. Nonetheless, deploying advanced super-resolution (SR) models on mobile devices is hindered by the computational demands of existing SR models. In this paper, we propose REM, a novel neural-enhanced mobile video streaming framework. REM utilizes a customized lookup table to facilitate real-time neural-enhanced video streaming on mobile devices. Initially, we conduct a series of measurements to identify abundant macroblock redundancies across frames in a video stream. Subsequently, we introduce a dynamic macroblock selection algorithm that prioritizes important macroblocks for neural enhancement. The SR-enhanced results are stored in the lookup table and efficiently reused to meet real-time requirements and minimize resource overhead. By considering macroblock-level characteristics of the video frames, the lookup table enables efficient and fast processing. Additionally, we design a lightweight macroblock-aware SR module to expedite inference. Finally, we perform extensive experiments on various mobile devices. The results demonstrate that REM enhances overall processing throughput by up to 10.2 times and reduces power consumption by up to 58.6% compared to state-of-the-art methods. Consequently, this leads to a 38.06% improvement in the quality of experience for mobile users.
Baili Chai, Di Wu 0001, Mengyu Yang, Miao Hu 0001
IEEE Trans. Mob. Comput.4
2024 Near-Lossless Gradient Compression for Data-Parallel Distributed DNN Training
abstract
Data parallelism has become a cornerstone in scaling up the training of deep neural networks (DNNs). However, the communication overhead associated with synchronizing gradients across multiple nodes has emerged as a significant bottleneck, adversely affecting training efficiency and leading to a surge in large-scale distributed model training costs. By leveraging insights into the statistical characteristics of gradients, we present GComp, a near-lossless gradient compression scheme designed to reduce the communication burden during data-parallel training significantly. GComp develops an optimized Huffman encoding/decoding strategy to compress gradient exponents effectively. Additionally, it introduces an innovative multi-level quantization method for mantissa, complemented by a pruning strategy that eliminates zero-valued gradients. These integrated approaches significantly reduce the volume of data for synchronization, while virtually not affecting the DNN model's training accuracy. We conduct comprehensive evaluations of GComp, demonstrating that our method can decrease the communication volume by as much as 67.1%, and enhance training speed by up to 1.9×.
Xue Li 0024, Cheng Guo 0007, Kun Qian 0021, Menghao Zhang 0001, Mengyu Yang, Mingwei Xu 0001
SoCC5
2024 AdaViPro: Region-Based Adaptive Visual Prompt For Large-Scale Models Adapting
abstract
Recently, prompt-based methods have emerged as a new alternative ‘parameter-efficient fine-tuning’ paradigm, which only fine-tunes a small number of additional parameters while keeping the original model frozen. However, despite achieving notable results, existing prompt methods mainly focus on ‘what to add’, while overlooking the equally important aspect of ‘where to add’, typically relying on the manually crafted placement. To this end, we propose a region-based Adaptive Visual Prompt, named AdaViPro, which integrates the ‘where to add’ optimization of the prompt into the learning process. Specifically, we reconceptualize the ‘where to add’ optimization as a problem of regional decision-making. During inference, AdaViPro generates a regionalized mask map for the whole image, which is composed of 0 and 1, to designate whether to apply or discard the prompt in each specific area. Therefore, we employ Gumbel-Softmax sampling to enable AdaViPro’s end-to-end learning through standard back-propagation. Extensive experiments demonstrate that our AdaViPro yields new efficiency and accuracy trade-offs for adapting pre-trained models.
Mengyu Yang, Ye Tian 0008, Lanshan Zhang, Xuming Ran, Wendong Wang 0003
ICIP1
2024 The Un-Kidnappable Robot: Acoustic Localization of Sneaking People
abstract
How easy is it to sneak up on a robot? We examine whether we can detect people using only the incidental sounds they produce as they move, even when they try to be quiet. To do so, we first collect a robotic dataset of high-quality 4-channel audio paired with 360° RGB data of people moving in different indoor settings. Using this dataset, we train models to predict if there is a moving person nearby and then their location using only audio. We implement our method on a robot, allowing it to track a single person moving quietly using only passive audio sensing. For demonstration videos, see our project page.
Mengyu Yang, Patrick Grady, Samarth Brahmbhatt, Arun Balajee Vasudevan, Charles C. Kemp, James Hays
ICRA1
2024 DTA: Deformable Temporal Attention for Video Recognition
abstract
Recently, transformer models have demonstrated superior performance in video tasks. However, a prevalent limitation in most current video Transformers lies in their tendency to overlook inherent temporal regions of interest, such as motion trajectories, leading to susceptibility to redundant information during temporal modeling. Existing methods that pay attention to motion trajectories have high computational demands, lacking in lightweight efficiency. To strike a balance between effective modeling of temporal regions of interest and computational efficiency, we propose a video transformer backbone with deformable temporal attention (DTA). Inspired by the work on deformable receptive fields, DTA employs a lightweight decision network to enhance the flexibility of temporal attention. The decision network computes the offsets of tokens in the input feature map, enabling them to move to temporally relevant regions of interest and efficiently model temporal information. We conducted extensive experiments on three popular datasets and surpassed the baseline. Additionally, we performed ablation experiments specifically targeting the model structure and parameters. These results confirm the effectiveness of the proposed deformable temporal attention mechanism.
Xiaohan Lei, Mengyu Yang, Gongli Xi, Yang Liu 0325, Jiulin Li, Lanshan Zhang, Ye Tian 0008
IJCNN2
2024 Semantic Fusion Based Graph Network for Video Scene Detection
abstract
Video scene detection, an initial step of video analysis, temporally divides heterogeneous video into semantic segments, which is widely used in video summarization, search, browsing and retrieval. Video scene detection always cuts video into shots first and then groups these shots into segments. In this process, how to solve complex dependency relationship among shots is a barrier. The existing methods consider using Recurrent Neural Networks and Hidden Markov Model to simulate the dependency relationship between shots. However, linear approaches work a little on hierarchical video structure. In this paper, a GNN-based network is proposed to model complex structures of videos instead. Besides, the semantic gap between low-level features and high-level semantics is also a big obstacle of video scene detection. Here, three visual semantic elements in shot, i.e., environment, object, action and audio feature are extracted as shot representation. Later, we utilize a multi-modal fusion strategy, which combines early fusion and late fusion, to bridge the semantic gap between low-level features and high-level semantics. The proposed method was evaluated on BBC Planet Earth dataset and Open Video Scene Detection (OSVD) dataset, the experimental results demonstrate that the proposed method outperforms the state-of-the-art in video scene detection task.
Ye Tian 0008, Yang Liu 0325, Mengyu Yang, Lanshan Zhang
IJCNN3
2024 Leveraging Coarse-to-Fine Grained Representations in Contrastive Learning for Differential Medical Visual Question Answering
Di Wang 0011, Zhicheng Jiao, Haodi Zhong, Mengyu Yang, Quan Wang 0006
MICCAI (5)6
2024 WaveDN: A Wavelet-based Training-free Zero-shot Enhancement for Vision-Language Models
abstract
Vision-Language Models (VLMs) built on contrastive learning, such as CLIP, demonstrate great transferability and excel in downstream tasks like zero-shot classification and retrieval. To further enhance the performance of VLMs, existing methods have introduced additional parameter modules or fine-tuned VLMs on downstream datasets. However, these methods often fall short in scenarios where labeled data for downstream tasks is either unavailable or insufficient for fine-tuning, and the training of additional parameter modules may considerably impair the existing transferability of VLMs on open-set tasks. To alleviate this issue, we introduce WaveDN, a wavelet-based distribution normalization method that can boost the VLMs' performance on downstream tasks without parametric modules or labeled data. Initially, wavelet distributions are extracted from the embeddings of the sampled, unlabeled test samples. Subsequently, WaveDN conducts a hierarchical normalization across the wavelet coefficients of all embeddings, thereby incorporating the distributional characteristics of the test data. Finally, the normalized embeddings are reconstructed via inverse wavelet transformation, facilitating the computation of similarity metrics between the samples. Through extensive experiments on two downstream tasks, using a total of 14 datasets covering text-image and text-audio modal data, WaveDN has demonstrated superiority compared to state-of-the-art methods.
Jiulin Li, Mengyu Yang, Ye Tian 0008, Lanshan Zhang, Yongchun Lu, Jice Liu, Wendong Wang 0003
ACM Multimedia2
2024 Global Patch-wise Attention is Masterful Facilitator for Masked Image Modeling
abstract
Masked image modeling (MIM), as a self-supervised learning paradigm in computer vision, has gained widespread attention among researchers. MIM operates by training the model to predict masked patches of the image. Given the sparse nature of image semantics, it is imperative to devise a masking strategy that steers the model towards reconstructing high-semantic regions. However, conventional mask strategies often miss these high-semantic regions or lack alignment with the masks and semantics. To solve this, we propose the Global Patch-wise Attention (GPA) framework, a transferable and efficient framework for MIM pre-training. We observe that the attention between patches can be the metric of identifying high-semantic regions, which can guide the model to learn more effective representations. Therefore, we firstly define the global patch-wise attention via vision transformer blocks. Then we design the soft-to-hard mask generation to guide the model gradually focusing on high semantic regions identified by GPA (GPA as a teacher). Finally, we design an extra task to predict GPA (GPA as a feature). Experiments conducted under various settings demonstrate that our proposed GPA framework enables MIM to learn better representations, which benefit the model across a wide range of downstream tasks. Furthermore, our GPA framework can be easily and effectively transferred to various MIM architectures.
Gongli Xi, Ye Tian 0008, Mengyu Yang, Lanshan Zhang, Xirong Que, Wendong Wang 0003
ACM Multimedia3
2023 Object-Based Multipath Transmission Scheduling Algorithm in Multi-Modal Scenarios
abstract
At present, with the rapid advancement of mul-timedia technology, more and more multimodal applications have emerged. The data communication of multimodal applications involves multiple modalities, and the transmitted data has deadline requirements and block transmission characteristics. However, existing transport layer protocols cannot meet the transmission requirements of applications by perceiving the data attributes and it is difficult to avoid high-priority modalities from excessively seizing transmission resources, resulting in transmission starvation in other modalities. Therefore, this paper considers the data blocks and their transmission requirements of the upper layer application as independent data objects and proposes an object multipath transmission scheduling al-gorithm (OMTS) that can consider the fairness of multimodal transmission. OMTS uses reinforcement learning algorithms to comprehensively consider the transmission requirements of data objects and the quality of multipath network transmission, to determine the transmission order and path allocation strategy of objects. In addition, we also design a model structure that separates scheduling and learning, allowing the algorithm to learn more valuable scheduling strategies through continuous interaction with the environment. The comparative experiment of data transmission through the network simulation environment shows that OMTS is superior to existing scheduling algorithms.
Ye Tian 0008, Mengyu Yang, Xiangyang Gong
GLOBECOM3
2023 Understanding and Improving Perceptual Quality of Volumetric Video Streaming
abstract
Volumetric video is fully three-dimensional and provides users with highly immersive and interactive experience. However, it is difficult to stream volumetric video over the Internet due to sheer video size and limited network bandwidth. Existing solutions suffered from poor perceptual quality and low coding efficiency. In this paper, we first conduct a comprehensive user study to understand the effectiveness of popular perceptual quality metrics for volumetric video. It is observed that those metrics cannot well capture the impact of user viewing behaviors. Considering the findings that users are more sensitive to the distortion of 2D image rendered from 3D point cloud, a new metric called Volu-FMAF is proposed to better represent perceptual quality of volumetric video. Next, we propose a novel neural-based volumetric video streaming framework RenderVolu and design a distortion-aware rendered image super-resolution network, called RenDA-Net, to further improve user perceptual quality. Last, we conduct extensive experiments with real datasets to validate our proposed method, and the results show that our method can boost the perceptual quality of volumetric video by 171% to 190%, and achieves a speedup of 108x in terms of decoding efficiency compared to the state-of-the-art approaches.
Mengyu Yang, Di Wu 0001, Miao Hu 0001, Yipeng Zhou
ICME1
2023 Cost-effective Modality Selection for Video Popularity Prediction
abstract
Video popularity prediction, also known as predicting the future popularity of a video, is a task that generally uses various modalities of the video, such as visual, audio, text, and metadata, to make predictions. However, the traditional approach puts all the modalities directly into a model for processing, which introduces noise and redundancy into the model. To address this issue, we propose Cost-effective Modality Selection(CeMS) for video popularity prediction, which can adaptively select cost-effective modalities as input based on the features of the metadata. Specifically, a policy network first obtains the prior information from metadata to analyze the differences in representation ability and computation cost between various modalities. Then, with the policy estimation, an optimal set of cost-effective modalities is determined, and the backbone network corresponding to the selected modalities is activated for feature extraction. Finally, the semantic information of all selected modalities is fed into the decoder to yield the predicted popularity. Experiments demonstrate that our proposed approach yields 63%-69% reduction in computation when compared to the traditional baseline that simply uses all the modalities. Additionally, we achieve consistent improvements in accuracy over the state-of-the-art methods.
Yang Liu 0325, Mengyu Yang, Ye Tian 0008, Lanshan Zhang, Xirong Que, Wendong Wang 0003
IJCNN2
2023 View while Moving: Efficient Video Recognition in Long-untrimmed Videos
abstract
Recent adaptive methods for efficient video recognition mostly follow the two-stage paradigm of "preview-then-recognition" and have achieved great success on multiple video benchmarks. However, this two-stage paradigm involves two visits of raw frames from coarse-grained to fine-grained during inference (cannot be parallelized), and the captured spatiotemporal features cannot be reused in the second stage (due to varying granularity), being not friendly to efficiency and computation optimization.To this end, inspired by human cognition, we propose a novel recognition paradigm of "View while Moving" for efficient long-untrimmed video recognition.In contrast to the two-stage paradigm, our paradigm only needs to access the raw frame once.The two phases of coarse-grained sampling and fine-grained recognition are combined into unified spatiotemporal modeling, showing great performance.Moreover, we investigate the properties of semantic units in video and propose a hierarchical mechanism to efficiently capture and reason about the unit-level and video-level temporal semantics in long-untrimmed videos respectively.Extensive experiments on both long-untrimmed and short-trimmed videos demonstrate that our approach outperforms state-of-the-art methods in terms of accuracy as well as efficiency, yielding new efficiency and accuracy trade-offs for video spatiotemporal modeling.
Ye Tian 0008, Mengyu Yang, Lanshan Zhang, Zhizhen Zhang, Yang Liu 0325, Xiaohui Xie, Xirong Que, Wendong Wang 0003
ACM Multimedia2
2023 A Fuzzy Error Based Fine-Tune Method for Spatio-Temporal Recognition Model
Jiulin Li, Mengyu Yang, Yang Liu 0325, Gongli Xi, Lanshan Zhang, Ye Tian 0008
PRCV (1)2
2023 An accurate shared bicycle detection network based on faster R-CNN
abstract
Abstract Detecting shared bicycles is an essential and challenging task. Deep learning has been widely used in object detection tasks in urban scenes, such as vehicle detection. However, deep learning algorithms still face many difficulties and challenges in shared bicycle detection. For example, the problem of large deformation of shared bicycles and the problem of small targets because the camera is far away from the shared bicycles. In order to solve these problems, this study introduces the feature fusion module and deformable convolution into the object detection network, which improves the efficiency of shared bicycle detection. This study proposes an enhanced faster R‐CNN network (A classic two‐stage object detection network) for shared bicycle detection and a shared bicycle dataset (SBD) is constructed for model training and testing. Compared with the original faster R‐CNN, the mean average precision (mAP) of the enhanced method on SBD is improved by 13%, which indicates that the method provided in this study is more suitable for detecting shared bicycles. This study also conducts experiments on the Microsoft Common Objects (COCO) dataset, where this method achieves 40.2% of the mAP, which is 5.8% higher than faster R‐CNN before improvement.
Lingqiao Li, Mengyu Yang
IET Image Process.3
2023 FedCL: Federated contrastive learning for multi-center medical image classification
Zhenbing Liu, Fengfeng Wu, Mengyu Yang, Xipeng Pan
Pattern Recognit.4
2021 Soloist: Generating Mixed-Initiative Tutorials from Existing Guitar Instructional Videos Through Audio Processing
abstract
Learning musical instruments using online instructional videos has become increasingly prevalent. However, pre-recorded videos lack the instantaneous feedback and personal tailoring that human tutors provide. In addition, existing video navigations are not optimized for instrument learning, making the learning experience encumbered. Guided by our formative interviews with guitar players and prior literature, we designed Soloist, a mixed-initiative learning framework that automatically generates customizable curriculums from off-the-shelf guitar video lessons. Soloist takes raw videos as input and leverages deep-learning based audio processing to extract musical information. This back-end processing is used to provide an interactive visualization to support effective video navigation and real-time feedback on the user's performance, creating a guided learning experience. We demonstrate the capabilities and specific use-cases of Soloist within the domain of learning electric guitar solos using instructional YouTube videos. A remote user study, conducted to gather feedback from guitar players, shows encouraging results as the users unanimously preferred learning with Soloist over unconverted instructional videos.
Bryan Wang, Mengyu Yang, Tovi Grossman
CHI2
2021 TriBERT: Human-centric Audio-visual Representation Learning
abstract
The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-modal variants (e.g., ViLBERT) have limited themselves to visual-linguistic data. Relatively few have explored its use in audio-visual modalities, and none, to our knowledge, illustrate them in the context of granular audio-visual detection or segmentation tasks such as sound source separation and localization. In this work, we introduce TriBERT -- a transformer-based architecture, inspired by ViLBERT, which enables contextual feature learning across three modalities: vision, pose, and audio, with the use of flexible co-attention. The use of pose keypoints is inspired by recent works that illustrate that such representations can significantly boost performance in many audio-visual scenarios where often one or more persons are responsible for the sound explicitly (e.g., talking) or implicitly (e.g., sound produced as a function of human manipulating an object). From a technical perspective, as part of the TriBERT architecture, we introduce a learned visual tokenization scheme based on spatial attention and leverage weak-supervision to allow granular cross-modal interactions for visual and pose modalities. Further, we supplement learning with sound-source separation loss formulated across all three streams. We pre-train our model on the large MUSIC21 dataset and demonstrate improved performance in audio-visual sound source separation on that dataset as well as other datasets through fine-tuning. In addition, we show that the learned TriBERT representations are generic and significantly improve performance on other audio-visual tasks such as cross-modal audio-visual-pose retrieval by as much as 66.7% in top-1 accuracy.
Tanzila Rahman, Mengyu Yang, Leonid Sigal
NeurIPS2