Mengyang Liu

dblp:143/0651 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CODERL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
abstract
Xue Jiang, Yihong Dong, Mengyang Liu, Deng Hongyi, Tian Wang, Yongding Tao, Zhi Jin, Wenpin Jiao, Ge Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yihong Dong, Mengyang Liu, Hongyi Deng, Yongding Tao, Zhi Jin 0001, Wenpin Jiao, Ge Li 0001
ACL (1)3
2026 PipeNN: A Predictor-Free Pipeline for Energy-Latency Co-Optimization of Heterogeneous Mobile DAG-DNN Inference
Yukun Tian, Tianwei Jiang, Ruiting Zhou, Fang Dong 0001, Mengyang Liu
ICDCS8
2026 Lightweight and Efficient Restoration of Beiting Murals: A Degression Receptive Field Network with Multi-scale Reparameterization and Wavelet-Fourier Frequency Domain Fusion
Mengyang Liu, Shumei Bao, Yuhao Ke, Youcai Cheng
ICIC (10)1
2025 Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling
abstract
Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. To address these challenges, we propose a novel multi-condition guided framework for character image animation, employing several well-designed input modules to enhance the implicit decoupling capability of the model. First, the optical flow guider calculates the background optical flow map as guidance information, which enables the model to implicitly learn to decouple the background motion into background constants and background momentum during training, and generate a stable background by setting zero background momentum during inference. Second, the depth order guider calculates the order map of the characters, which transforms the depth information into the positional information of multiple characters. This facilitates the implicit learning of decoupling different characters, especially in accurately separating the occluded body parts of multiple characters. Third, the reference pose map is input to enhance the ability to decouple character texture and pose information in the reference image. Furthermore, to fill the gap of fair evaluation of multi-character image animation, we propose a new benchmark comprising about 4,000 frames. Extensive qualitative and quantitative evaluations demonstrate that our method excels in generating high-quality character animations, especially in scenarios of complex backgrounds and multiple characters.
Jingyun Xue, Hongfa Wang, Qi Tian 0003, Yue Ma 0016, Andong Wang, Zhiyuan Zhao 0002, Shaobo Min, Kaihao Zhang, Harry Shum, Wei Liu 0005, Mengyang Liu, Wenhan Luo
ICLR12
2025 InfiniCL: Elastic Continual Learning for Resource-Constrained Edge Devices
abstract
On-device continual learning (CL) enables lifelong and privacy-preserving learning for various edge intelligent applications. Increasing the number of model parameters as new learning tasks emerge is effective in ensuring learning quality but inefficient in memory cost, especially for resource-constrained devices. In this paper, we introduce InfiniCL, the first ondevice CL system that dynamically balances memory cost and learning quality. A key idea behind InfiniCL is elastic continual learning: selectively freezing layers in the expanding model and periodically distilling the model, preventing unbounded memory growth while preserving learning quality for new tasks. This novel CL paradigm opens a new challenging problem: how to decide the memory allocation of the model and data to achieve better learning Quality of Service (QoS) under the limited memory budget? To alleviate this challenge, we further propose a Bayesian Optimization-driven algorithm to jointly optimize layer freezing selection and data-model memory allocation. Evaluations show that InfiniCL outperforms state-of-the-art methods on diverse memory constraints, achieving 5.34-7.15% and 2.72-9.36% higher accuracy on CIFAR-100 and ImageNet-100, respectively.
Chenyu Lu, Mengyang Liu, Fang Dong 0001, Borui Li 0001, Ruiting Zhou, Shiyao Ji
IWQoS2
2025 ADPTD: Adaptive Data Partition With Unbiased Task Dispatching for Video Analytics at the Edge
abstract
Recently, edge-assisted methods have been proposed as a promising technique to deliver fast and accurate on-device video analytics by partitioning frame data and dispatching them to edge servers for parallel execution. However, the data partition (DP) reduces the detection latency but decreases accuracy since objects may cross the boundaries of adjacent blocks. The effect of DP on the accuracy and latency depends on multiple vital parameters (e.g., target size, density, network, and computing resources) in an unknown and time-varying fashion. Moreover, these parameters are determined by the application scenarios and edge environment, which are uncertain and heterogeneous at the edge. Hence, how to partition frames to strike a balance between accuracy and latency is a nontrivial and intractable problem. To this end, we propose an online learning-based device-edge–cloud collaboration framework, ADPTD, to guide DP at the edge. We propose an optimal task dispatching algorithm (OTD) to minimize detection latency. Then, we propose a multiarmed bandit-based algorithm to pick a DP strategy and invoke OTD to dispatch tasks in each time slot. Theoretical analysis reveals that ADPTD achieves sublinear regret. Extensive experimental results show that ADPTD outperforms the state-of-the-art methods, achieving a latency reduction of up to$2.53\times $and improving accuracy by up to 49.4%.
Zhaowu Huang, Fang Dong 0001, Haopeng Zhu, Mengyang Liu, Dian Shen, Ruiting Zhou, Xiaolin Guo, Baijun Chen
IEEE Internet Things J.4
2025 Personalized similarity regression models based on maximum correntropy criterion for stock series prediction
Mengyang Liu, Xiaoyan Qiao, Shiyu Ge, Pingping Liu
Knowl. Inf. Syst.2
2024 AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data
abstract
Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that a considerable amount of the publicly available speech data is compromised by background noise, speech overlapping, lack of speech segmentation information, missing speaker labels, and incomplete transcriptions, which can largely hinder their usefulness. On the other hand, human annotation of speech data is both time-consuming and costly. To address this issue, we introduce an automatic in-the-wild speech data preprocessing framework (AutoPrep) in this paper, which is designed to enhance speech quality, generate speaker labels, and produce transcriptions automatically. The proposed AutoPrep framework comprises six components: speech enhancement, speech segmentation, speaker clustering, target speech extraction, quality filtering and automatic speech recognition. Experiments conducted on the open-sourced WenetSpeech and our self-collected AutoPrepWild corpora demonstrate that the proposed AutoPrep framework can generate preprocessed data with similar DNSMOS and PDNSMOS scores compared to several open-sourced TTS datasets. The corresponding TTS system can achieve up to 0.68 in-domain speaker similarity.1
Jianwei Yu 0001, Hangting Chen, Yanyao Bian, Yi Luo 0004, Jinchuan Tian, Mengyang Liu, Jiayi Jiang, Shuai Wang 0016
ICASSP7
2024 New reinforcement learning based on representation transfer for portfolio management
Mengyang Liu, Mingyan Xu, Shuoru Chen, Pingping Liu, Caiming Zhang 0001, Feng Zhao 0006
Knowl. Based Syst.2
2023 CoCV: Heterogeneous Processors Collaboration Mechanism for End-to-End Execution of Intelligent Computer Vision Tasks on Mobile Devices
abstract
Object detection, image classification, and various other computer vision tasks have become prevalent on mobile devices. These computer vision tasks are typically executed with three stages: pre-processing, inference, and post-processing. Mobile SoC, serving as the computing unit on mobile devices, typically consists of heterogeneous processors like CPU, GPU, and NPU. However, during the execution of a computer vision task, current available frameworks only achieve the parallelism of CPU and GPU in the inference stage. While during pre- and post-processing, only CPU is used, leaving GPU and NPU on the SoC to be idle. For mobile applications that require low latency, the overhead of pre-processing and post-processing stages often account for more than 50% of the total latency, which becoming a performance bottleneck of the entire task. To reduce latency, it is imperative to fully utilize the idle heterogeneous processors (GPU, NPU) on the SoC and achieve heterogeneous processors parallelism in all three stages during execution. In this paper, we propose CoCV, a heterogeneous processor parallel computing system for computer vision tasks on mobile devices. In CoCV, we are the first to build an image processing operator library for heterogeneous parallel computing on mobile devices. Besides, we design a task allocation scheduling algorithm to guide the partitioning of processing tasks during execution, which ensuring a relatively balanced workload between different processors. A cross-stage operator chaining technique is also proposed to reduce the data sharing overhead among different processors during execution. We build a prototype system and evaluate it with different computer vision tasks. The results show up to 33% latency reduction for end-to-end tasks and 2.32× speedup compared with the current best solution.
Ye Wan, Mengyang Liu, Guangtong Li, Fang Dong 0001
ICPADS2
2023 Adaptive Overlap Padding and Resolution Selection for Frame Split-based Edge Video Analytics
abstract
For providing accurate and fast on-device high-resolution video analytics, edge-assisted methods are widely proposed by using lower-resolution frames and splitting them with overlap padding. The use of low-resolution frames can significantly reduce the computational workload of video analytics. Dividing a frame into multiple overlapping parts simultaneously improves parallelism and maintains processing accuracy. However, accuracy and latency serve as a pair of tradeoff metrics, and prioritizing one to optimize overlap padding or resolution selection will compromise the other metric. Fortunately, we have discovered that in real-world video analytics scenarios with varying object sizes, it is not necessary to simultaneously achieve a high overlap padding size and high resolution. Hence, how to set appropriate overlap padding size and resolution to strike a balance between accuracy and latency in practical scenarios is a nontrivial and intractable problem. To this end, we propose an online learning-based method to achieve adaptive overlap padding and resolution selection, called APR. We model the problem as an integer programming and propose a Muli-armed bandit (MAB) theory-based algorithm to solve it. We discretize the continuum overlap padding size into a finite set to narrow explore space and set the frame split strategy as context information to achieve fast convergence. Theoretical analysis reveals APR achieves sub-linear regret. Extensive experimental results show APR outperforms the benchmark methods, achieving up to 2.06 × speedup in terms of latency and 0.18× increase in accuracy.
Haopeng Zhu, Zhaowu Huang, Xiaolin Guo, Mengyang Liu, Baijun Chen, Fang Dong 0001
ICPADS4
2023 SVCNet: Scribble-Based Video Colorization Network With Temporal Aggregation
abstract
In this paper, we propose a scribble-based video colorization network with temporal aggregation called SVCNet. It can colorize monochrome videos based on different user-given color scribbles. It addresses three common issues in the scribble-based video colorization area: colorization vividness, temporal consistency, and color bleeding. To improve the colorization quality and strengthen the temporal consistency, we adopt two sequential sub-networks in SVCNet for precise colorization and temporal smoothing, respectively. The first stage includes a pyramid feature encoder to incorporate color scribbles with a grayscale frame, and a semantic feature encoder to extract semantics. The second stage finetunes the output from the first stage by aggregating the information of neighboring colorized frames (as short-range connections) and the first colorized frame (as a long-range connection). To alleviate the color bleeding artifacts, we learn video colorization and segmentation simultaneously. Furthermore, we set the majority of operations on a fixed small image resolution and use a Super-resolution Module at the tail of SVCNet to recover original sizes. It allows the SVCNet to fit different image resolutions at the inference. Finally, we evaluate the proposed SVCNet on DAVIS and Videvo benchmarks. The experimental results demonstrate that SVCNet produces both higher-quality and more temporally consistent videos than other well-known video colorization approaches. The codes and models can be found at https://github.com/zhaoyuzhi/SVCNet.
Yuzhi Zhao, Lai-Man Po, Kangcheng Liu, Wing Yin Yu, Pengfei Xian, Yujia Zhang 0002, Mengyang Liu
IEEE Trans. Image Process.8
2023 VCGAN: Video Colorization With Hybrid Generative Adversarial Network
abstract
We propose a Video Colorization with Hybrid Generative Adversarial Network (VCGAN), an improved approach to video colorization using end-to-end learning and recurrent architecture. The VCGAN addresses two prevalent issues in the video colorization domain: Temporal consistency and the unification of colorization network and refinement network into a single architecture. To enhance colorization quality and spatiotemporal consistency, the mainstream of the generator in VCGAN is assisted by two additional networks,i.e.,global feature extractor and placeholder feature extractor, respectively. The global feature extractor encodes the global semantics of grayscale input to enhance colorization quality, whereas the placeholder feature extractor serves as a feedback connection to encode the semantics of the previous colorized frame in order to maintain spatiotemporal consistency. If changing the input for placeholder feature extractor as grayscale input, the hybrid VCGAN also has the potential to colorize single images. To improve the color consistency of far frames, we propose a dense long-term loss that minimizes the temporal disparity of every two remote frames. Trained with colorization and temporal losses jointly, VCGAN strikes a good balance between video color vividness and spatiotemporal continuity. Experimental results demonstrate that VCGAN produces higher-quality and temporally more consistent colorful videos than existing approaches.
Yuzhi Zhao, Lai-Man Po, Wing Yin Yu, Yasar Abbas Ur Rehman, Mengyang Liu, Yujia Zhang 0002, Weifeng Ou
IEEE Trans. Multim.5
2022 Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video Representation
abstract
Spatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignoring the intermediate state of the learned representations, which limits the overall performance. In this work, taking into account the degree of similarity of sampled instances as the intermediate state, we propose a novel pretext task - spatio-temporal overlap rate (STOR) prediction. It stems from the observation that humans are capable of discriminating the overlap rates of videos in space and time. This task encourages the model to discriminate the STOR of two generated samples to learn the representations. Moreover, we employ a joint optimization combining pretext tasks with contrastive learning to further enhance the spatio-temporal representation learning. We also study the mutual influence of each component in the proposed scheme. Extensive experiments demonstrate that our proposed STOR task can favor both contrastive learning and pretext tasks and the joint optimization scheme can significantly improve the spatio-temporal representation in video understanding. The code is available at https://github.com/Katou2/CSTP.
Yujia Zhang 0002, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing Yin Yu
AAAI4
2020 SLNet: Stereo face liveness detection via dynamic disparity-maps and convolutional neural network
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu
Expert Syst. Appl.3
2020 Data-level information enhancement: Motion-patch-based Siamese Convolutional Neural Networks for human activity recognition in videos
Yujia Zhang 0002, Lai-Man Po, Mengyang Liu, Yasar Abbas Ur Rehman, Weifeng Ou, Yuzhi Zhao
Expert Syst. Appl.3
2019 Investigating Cognitive Effects in Session-level Search User Satisfaction
abstract
User satisfaction is an important variable in Web search evaluation studies and has received more and more attention in recent years. Many studies regard user satisfaction as the ground truth for designing better evaluation metrics. However, most of the existing studies focus on designing Cranfield-like evaluation metrics to reflect user satisfaction at query-level. As information need becomes more and more complex, users often need multiple queries and multi-round search interactions to complete a search task (e.g. exploratory search). In those cases, how to characterize the user's satisfaction during a search session still remains to be investigated. In this paper, we collect a dataset through a laboratory study in which users need to complete some complex search tasks. With the help of hierarchical linear models (HLM), we try to reveal how user's query-level and session-level satisfaction are affected by different cognitive effects. A number of interesting findings are made. At query level, we found that although the relevance of top-ranked documents have important impacts (primacy effect), the average/maximum of perceived usefulness of clicked documents is a much better sign of user satisfaction. At session level, perceived satisfaction for a particular query is also affected by the other queries in the same session (anchor effect or expectation effect). We also found that session-level satisfaction correlates mostly with the last query in the session (recency effect). The findings will help us design better session-level user behavior models and corresponding evaluation metrics.
Mengyang Liu, Jiaxin Mao, Yiqun Liu 0001, Min Zhang 0006, Shaoping Ma
KDD1
2019 Deep Hashing with Triplet Labels and Unification Binary Code Selection for Fast Image Retrieval
Chang Zhou 0008, Lai-Man Po, Mengyang Liu, Wilson Y. F. Yuen, Peter Hon-Wah Wong, Hon-Tung Luk, Kin Wai Lau, Hok Kwan Cheung
MMM (1)3
2019 Face liveness detection using convolutional-features fusion of real and deep network generated face images
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu, Zijie Zou, Weifeng Ou, Yuzhi Zhao
J. Vis. Commun. Image Represent.3
2019 Video copy detection by conducting fast searching of inverted files
Mengyang Liu, Lai-Man Po, Yasar Abbas Ur Rehman, Xuyuan Xu, Litong Feng
Multim. Tools Appl.1
2019 A Novel Patch Variance Biased Convolutional Neural Network for No-Reference Image Quality Assessment
abstract
Deep convolutional neural networks (CNNs) have been successfully applied on no-reference image quality assessment (NR-IQA) with respect to human perception. Most of these methods deal with small image patches and use the average score of the test patches for predicting the whole image quality. We discovered that image patches from homogenous regions are unreliable for both neural network training and final image quality score estimation. In addition, image patches with complex structures have much higher chances of achieving better image quality prediction. Based on these findings, we enhanced the conventional CNN-based NR-IQA algorithm to avoid homogenous patches for the network training and quality score estimation. Moreover, we also use a variance-based weighting average to bias the final image quality score to the patches with complex structure. The experimental results show that this simple approach can achieve state-of-the-art performance compared with well-known NR-IQA algorithms.
Lai-Man Po, Mengyang Liu, Wilson Y. F. Yuen, Xuyuan Xu, Chang Zhou 0008, Peter Hon-Wah Wong, Kin Wai Lau, Hon-Tung Luk
IEEE Trans. Circuits Syst. Video Technol.2
2018 Towards Designing Better Session Search Evaluation Metrics
abstract
User satisfaction has been paid much attention to in recent Web search evaluation studies and regarded as the ground truth for designing better evaluation metrics. However, most existing studies are focused on the relationship between satisfaction and evaluation metrics at query-level. However, while search request becomes more and more complex, there are many scenarios in which multiple queries and multi-round search interactions are needed (e.g. exploratory search). In those cases, the relationship between session-level search satisfaction and session search evaluation metrics remain uninvestigated. In this paper, we analyze how users' perceptions of satisfaction accord with a series of session-level evaluation metrics. We conduct a laboratory study in which users are required to finish some complex search tasks and provide usefulness judgments of documents as well as session-level and query level satisfaction feedbacks. We test a number of popular session search evaluation metrics as well as different weighting functions. Experiment results show that query-level satisfaction is mainly decided by the clicked document that they think the most useful (maximum effect). While session-level satisfaction is highly correlated with the most recently issued queries (recency effect). We further propose a number of criteria for designing better session search evaluation metrics.
Mengyang Liu, Yiqun Liu 0001, Jiaxin Mao, Cheng Luo 0001, Shaoping Ma
SIGIR1
2018 "Satisfaction with Failure" or "Unsatisfied Success": Investigating the Relationship between Search Success and User Satisfaction
abstract
User satisfaction has been paid much attention to in recent Web search evaluation studies. Although satisfaction is often considered as an important symbol of search success, it doesn»t guarantee success in many cases, especially for complex search task scenarios. In this study, we investigate the differences between user satisfaction and search success, and try to adopt the findings to predict search success in complex search tasks. To achieve these research goals, we conduct a laboratory study in which search success and user satisfaction are annotated by domain expert assessors and search users, respectively. We find that both "Satisfaction with Failure" and "Unsatisfied Success" cases happen in these search tasks and together they account for as many as 40.3% of all search sessions. The factors (e.g. document readability and credibility) that lead to the inconsistency of search success and user satisfaction are also investigated and adopted to predict whether one search task is successful. Experimental results show that our proposed prediction method is effective in predicting search success.
Mengyang Liu, Yiqun Liu 0001, Jiaxin Mao, Cheng Luo 0001, Min Zhang 0006, Shaoping Ma
WWW1
2018 LiveNet: Improving features generalization for face liveness detection using convolution neural networks
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu
Expert Syst. Appl.3
2015 MonkeyDroid: Detecting Unreasonable Privacy Leakages of Android Applications
Mengyang Liu, Shanqing Guo, Tao Ban
ICONIP (3)2