EDBT 2026 Demo / reviewers in the wild / expert
Haoyu Lu
dblp:240/2720
· DBLP profile ↗
28ranked-venue papers
7as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AITQE: An Adaptive Image-Text Quality Enhancer for Scalable MLLM PretrainingabstractMultimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-text pairs within multimodal pretraining datasets. However, in the process of high-quality data curation, filter-based paradigms often discard a substantial portion of high-quality images due to inadequate semantic alignment between images and texts, leading to inefficiency in data utilization and scalability. In this paper, we propose the Adaptive Image-Text Quality Enhancer (AITQE), a model that dynamically assesses and enhances the quality of image-text pairs. AITQE employs a text rewriting mechanism for low-quality pairs and incorporates a negative sample learning strategy to improve evaluative capabilities by integrating deliberately generated low-quality samples during training. Unlike prior approaches that significantly alter text distributions, our method minimally adjusts text to preserve data volume while enhancing quality. Experimental results demonstrate that AITQE surpasses existing methods on various benchmarks, effectively leveraging raw data and scaling with increasing data volumes. Codes and model are available at https://github.com/hanhuang22/AITQE. Yuqi Huo, Zijia Zhao, Haoyu Lu, Bingning Wang, Qiang Liu 0006, Weipeng Chen, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Efficient Motion-Aware Video MLLMabstractMost current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM that utilizes compressed video structures as inputs. We propose a motion-aware GOP (Group of Pictures) encoder that fuses spatial and motion information within a GOP unit in the compressed video stream, generating compact, informative visual tokens. By integrating fewer but denser RGB frames with more but sparser motion vectors in this native slow-fast input architecture, our approach reduces redundancy and enhances motion representation. Additionally, we introduce MotionBench, a benchmark for evaluating motion understanding across four motion types: linear, curved, rotational, and contact-based. Experimental results show that EMA achieves state-of-the-art performance on both MotionBench and popular video question answering benchmarks, while reducing inference costs. Moreover, EMA demonstrates strong scalability, as evidenced by its competitive performance on long video understanding benchmarks. Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo, Haoyu Lu, Bingning Wang, Weipeng Chen, Jing Liu 0001 |
CVPR | 5 |
| 2025 | R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal Formalization
Yi Yang 0001, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu 0001, Wei Chen 0001 |
ICCV | 7 |
| 2025 | Exploring the Design Space of Visual Context Representation in Video MLLMsabstractVideo Multimodal Large Language Models~(MLLMs) have shown remarkable capability of understanding the video semantics on various downstream tasks. Despite the advancements, there is still a lack of systematic research on visual context representation, which refers to the scheme to select frames from a video and further select the tokens from a frame. In this paper, we explore the design space for visual context representation, and aim to improve the performance of video MLLMs by finding more effective representation schemes. Firstly, we formulate the task of visual context representation as a constrained optimization problem, and model the language modeling loss as a function of the number of frames and the number of embeddings (or tokens) per frame, given the maximum visual context window size. Then, we explore the scaling effects in frame selection and token selection respectively, and fit the corresponding function curve by conducting extensive empirical experiments. We examine the effectiveness of typical selection strategies and present empirical findings to determine the two factors. Furthermore, we study the joint effect of frame selection and token selection, and derive the optimal formula for determining the two factors. We demonstrate that the derived optimal settings show alignment with the best-performed results of empirical experiments. The data and code are available at: https://github.com/RUCAIBox/Opt-Visor. Yifan Du 0002, Yuqi Huo, Kun Zhou 0002, Zijia Zhao, Haoyu Lu, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen |
ICLR | 5 |
| 2025 | Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMsabstractVideo understanding is a crucial next step for multimodal large language models (MLLMs).
Various benchmarks are introduced for better evaluating the MLLMs.
Nevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills.
In this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation.
VideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos.
The framework automates the generation of query-response pairs using predefined rules, minimizing manual labor. The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths.
Utilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development. Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 0002, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, Jing Liu 0001 |
ICLR | 2 |
| 2025 | Adaptive Multi-Beamforming for Integrated Sensing and Communication SystemabstractMulti-beamforming is an effective technique to mitigate interference in the integrated sensing and communication (ISAC) system. This paper proposes an adaptive multi-beamforming approach, in which a demand factor that optimizes the phase weighting factor is developed for synthesizing sensing and communication beams. Based on this demand factor, a novel osprey sparrow optimization (OSO) algorithm is proposed to efficiently calculate the optimal sensing and communication beamforming vectors. Simulation results demonstrate that the proposed approach achieves the desired multi-beamforming, catering to the requirements of sensing and communication. OSO suppresses side lobes with a reduction of 5–10 dB and enjoys more accurate main lobe pointing for linear antenna arrays (LAAs), compared to either sparrow search algorithm (SSA) or particle swarm optimization (PSO). When the size of antenna arrays becomes larger, OSO achieves at least 3-fold convergence speed improvement for planar antenna arrays (PAAs). Jieming Xie, Haoyu Lu, Hongcheng Zhuang |
VTC2025-Spring | 2 |
| 2024 | VEMO: A Versatile Elastic Multi-modal Model for Search-Oriented Multi-task Learning
Nanyi Fei, Hao Jiang 0022, Haoyu Lu, Jinqiang Long, Yanqi Dai, Tuo Fan, Zhao Cao, Zhiwu Lu 0001 |
ECIR (1) | 3 |
| 2024 | Progressive Image Synthesis from Semantics to Details with Denoising Diffusion GANabstractAlthough denoising diffusion probabilistic models (DDPMs) have shown remarkable progress in image generation, they typically face two main challenges: the time-expensive sampling process and the semantically meaningless latent space, which are often addressed separately in previous works. In particular, the latest representative work Denoising Diffusion GAN reduces the sampling steps to as few as two but ignores the semantics of the latent space. To address the two challenges simultaneously, we propose a two-stage framework to make the latent space of Denoising Diffusion GAN more semantically meaningful while enjoying its efficiency. Extensive results on three benchmark datasets demonstrate that our proposed diffusion model achieves competitive results with only two sampling steps in unconditional image generation. More importantly, the latent space of our diffusion model trained for unconditional image generation is shown to be semantically meaningful, which can be exploited on various downstream tasks (e.g., attribute editing) without further training. Guoxing Yang, Haoyu Lu, Chongxuan Li, Guang Zhou, Zhiwu Lu 0001 |
ICASSP | 2 |
| 2024 | Multi-Level Contrastive Learning For Hybrid Cross-Modal RetrievalabstractHybrid image retrieval is a significant task for a wide range of applications. In this scenario, the hybrid query for searching images consists of a reference image and a text modifier. The reference image provides a vital visual context and displays some semantic details, while the text modifier specifies the modifications to the reference image. To address such hybrid cross-modal retrieval, we propose a multi-level contrastive learning (MLCL) method for combining the hybrid query features into a fused feature by cross-modal contrastive learning with multi-level semantic alignment. Meanwhile, we additionally consider self-supervised contrastive learning to enhance the semantic correlation of the features at different levels of the combiner network. Extensive results on three public datasets (i.e., FashionIQ, Shoes, and CIRR) demonstrate that our proposed MLCL significantly outperforms the state-of-the-art methods under the hybrid cross-modal retrieval setting. Haoyu Lu, Zhiwu Lu 0001 |
ICASSP | 2 |
| 2024 | UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingabstractLarge-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the standard full fine-tuning paradigm becomes unsustainable due to heavy computational and storage costs. This paper proposes UniAdapter, which unifies unimodal and multimodal adapters for parameter-efficient cross-modal adaptation on pre-trained vision-language models. Specifically, adapters are distributed to different modalities and their interactions, with the total number of tunable parameters reduced by partial weight sharing. The unified and knowledge-sharing design enables powerful cross-modal representations that can benefit various downstream tasks, requiring only 1.0%-2.0% tunable parameters of the pre-trained model. Extensive experiments on 7 cross-modal downstream benchmarks (including video-text retrieval, image-text retrieval, VideoQA, VQA and Caption) show that in most cases, UniAdapter not only outperforms the state-of-the-arts, but even beats the full fine-tuning strategy. Particularly, on the MSRVTT retrieval task, UniAdapter achieves 49.7% recall@1 with 2.2% model parameters, outperforming the latest competitors by 2.0%. The code and models are available at https://github.com/RERV/UniAdapter. Haoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu 0001, Masayoshi Tomizuka, Mingyu Ding |
ICLR | 1 |
| 2024 | VDT: General-purpose Video Diffusion Transformers via Mask ModelingabstractThis work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation.
It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. Additionally, we propose a unified spatial-temporal mask modeling mechanism, seamlessly integrated with the model, to cater to diverse video generation scenarios.
VDT offers several appealing benefits. (1) It excels at capturing temporal dependencies to produce temporally consistent video frames and even simulate the physics and dynamics of 3D objects over time. (2) It facilitates flexible conditioning information, e.g., simple concatenation in the token space, effectively unifying different token lengths and modalities. (3) Pairing with our proposed spatial-temporal mask modeling mechanism, it becomes a general-purpose video diffuser for harnessing a range of tasks, including unconditional generation, video prediction, interpolation, animation, and completion, etc. Extensive experiments on these tasks spanning various scenarios, including autonomous driving, natural weather, human action, and physics-based simulation, demonstrate the effectiveness of VDT. Moreover, we provide a comprehensive study on the capabilities of VDT in capturing accurate temporal dependencies, handling conditioning information, and the spatial-temporal mask modeling mechanism. Additionally, we present comprehensive studies on how VDT handles conditioning information with the mask modeling mechanism, which we believe will benefit future research and advance the field. Codes and models are available at the https://VDT-2023.github.io. Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu 0001, Ping Luo 0002, Mingyu Ding |
ICLR | 1 |
| 2024 | Gridding Based Reconfigurable Intelligent Surface-aided Wireless Network OptimizationabstractAiming at the problem of weak coverage resulting from interference and obstructions in wireless networks, we propose a gridding based Reconfigurable Intelligent Surface (RIS)-assisted optimization scheme that divides problem area and RIS deployment area into several grids for better performance and lower complexity. The objective is to maximize the signal strength received in weak coverage area, with optimization variables comprising location, angle of incidence and angle of reflection of RIS. Then Particle Swarm Optimization (PSO) and Alternating Optimization (AO) algorithms are used to solve this non-convex problem. Simulation results show that our proposed scheme can achieve more than 25% improvement in terms of average SNR than conventional ones when the number of reflection elements of RIS is more than 100. Moreover, we analyze the impact of different sizes of RIS on its optimal location and the average SNR received in the weak coverage area, finding out when RIS is closer to the weak coverage area, the performance is not always better, which is different to existed observations. Haoyu Lu, Hongcheng Zhuang, Zhaocheng Wang 0001 |
PIMRC | 1 |
| 2024 | BotCL: a social bot detection model based on graph contrastive learning
Yan Li 0163, Zhenyu Li 0004, Daofu Gong, Haoyu Lu |
Knowl. Inf. Syst. | 5 |
| 2023 | PSMiner: A Pattern-Aware Accelerator for High-Performance Streaming Graph Pattern MiningabstractStreaming Graph Pattern Mining (GPM) has been widely used in many application fields. However, the existing streaming GPM solution suffers from many unnecessary explorations and isomorphism tests, while the existing static GPM ones require many repetitive operations to compute the full graph. In this paper, we propose a pattern-aware incremental execution approach and design the first streaming GPM accelerator called PSMiner, which integrates multiple optimizations to reduce redundant computation and improve computing efficiency. We have conducted extensive experiments. The results show that compared with the state-of-the-art software and hardware solutions, PSMiner achieves the average speedups of 770.9× and 60.4×, respectively. Hao Qi 0004, Yu Zhang 0027, Ligang He, Haoyu Lu, Jin Zhao 0003, Hai Jin 0001 |
DAC | 6 |
| 2023 | Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech RecognitionabstractIn recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech as input for the backend. However, it is difficult for speech enhancement systems to directly separate speech from input due to the diverse types of noise with different intensities. Furthermore, speech distortion and residual noise are often observed in enhanced speech, and the distortion of speech and noise is different. Most existing methods focus on fusing enhanced and noisy features to address this issue. In this paper, we propose a dual-stream spectrogram refine network to simultaneously refine the speech and noise and decouple the noise from the noisy input. Our proposed method can achieve better performance with a relative 8.6% CER reduction. Haoyu Lu, Tongtong Song, Longbiao Wang, Jianwu Dang 0001, Xiaobao Wang, Shiliang Zhang |
ICASSP | 1 |
| 2023 | BotCS: A Lightweight Model for Large-Scale Twitter Bot Detection Comparable to GNN-Based ModelsabstractSocial bot detection methods using graph neural networks (GNNs) are thriving, but the structural complexity of GNN also brings more training costs on large-scale data and interpretability concerns. In this paper, we propose a social bot detection method, BotCS, which utilizes both the attribute and the structural features of the social graph at a smaller computational cost than GNN-based detection methods. BotCS makes a base prediction with a simple multilayer perceptron classifier (MLP) and then propagates the classification residuals of the training set to other nodes for further correction. Then, it smooths the corrected prediction by label propagation. With little end-to-end training, this course is low-cost and scalable. We analyze the local interaction pattern between bots and human users, and designed the corresponding residual propagation and smoothing rules from the local perspective, which ensures the interpretability of BotCS. Experimental results show that BotCS achieves similar detection results to state-of-the-art methods with one or two orders of magnitude fewer parameters. Haoyu Lu, Daofu Gong, Zhenyu Li 0004, Feng Liu 0045, Fenlin Liu |
ICC | 1 |
| 2023 | A Topology Based Denoising Approach for 2D Scalar FieldsabstractA new denoising technique based on the multi-level dissipation element (ML-DE) structure is proposed. A dissipation element (DE) is defined as collection of spatial points whose gradient trajectories share the same pair of extremal points. The entire space can be decomposed into space-filling DEs. From the concept of multi-level extremal points, such a DE structure can be extended to different scale levels. The general idea of image smoothing is realized via smoothing each decomposed DE, remaining the function values at the critical points (maximal, minimal and saddle points) unchanged, from which the key features of the original image can be preserved. The test example of spiky map topology justifies the efficiency and effectiveness of this newly proposed method. Chun Gong, Yateng Qiao, Haoyu Lu, Lipo Wang 0002 |
ICIP | 3 |
| 2023 | Shot Retrieval and Assembly with Text Script for Video Montage GenerationabstractWith the development of video sharing websites, numerous users desire to create their own attractive video montages. However, it is difficult for inexperienced users to create well-edited video montages due to the lack of professional expertise. In the meantime, it is time-consuming even for experts to create video montages of high quality, which requires effectively selecting shots from abundant candidates and assembling them together. Instead of manual creation, various automatic methods have been proposed for video montage generation, which typically take a single sentence as input for text-to-shot retrieval, and ignore the semantic cross-sentence coherence given complicated text script of multiple sentences. To overcome this drawback, we propose a novel model for video montage generation by retrieving and assembling shots with arbitrary text scripts. To this end, a sequence consistency transformer is devised for cross-sentence coherence modeling. More importantly, with this transformer, two novel sequence-level tasks are defined for sentence-shot alignment in sequence-level: Cross-Modal Sequence Matching (CMSM) task, and Chaotic Sequence Recovering (CSR) task. To facilitate the research on video montage generation, we construct a new, highly-varied dataset which collects thousands of video-script pairs in documentary. Extensive experiments on the constructed dataset demonstrate the superior performance of the proposed model. The dataset and generated video demos are available at https://github.com/RATVDemo/RATV. Guoxing Yang, Haoyu Lu, Zelong Sun, Zhiwu Lu 0001 |
ICMR | 2 |
| 2022 | COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal RetrievalabstractLarge-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recently, two-stream methods like CLIP and ALIGN with high inference efficiency have also shown promising performance, however, they only consider instance-level alignment between the two streams (thus there is still room for improvement). To overcome these limitations, we propose a novel COllaborative Two-Stream vision-language pretraining model termed COTS for image-text retrieval by enhancing cross-modal interaction. In addition to instance-level alignment via momentum contrastive learning, we leverage two extra levels of cross-modal interactions in our COTS: (1) Token-level interaction - a masked vision-language modeling (MVLM) learning objective is devised without using a cross-stream network module, where variational autoencoder is imposed on the visual encoder to generate visual tokens for each image. (2) Task-level interaction - a KL-alignment learning objective is devised between text-to-image and image-to-text retrieval tasks, where the probability distribution per task is computed with the negative queues in momentum contrastive learning. Under a fair comparison setting, our COTS achieves the highest performance among all two-stream methods and comparable performance (but with 10,800× faster in inference) w.r.t. the latest single-stream methods. Importantly, our COTS is also applicable to text-to-video retrieval, yielding new state-of-the-art on the widely-used MSR-VTT dataset. Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao 0004, Zhiwu Lu 0001, Ji-Rong Wen |
CVPR | 1 |
| 2022 | Learning Versatile Neural Architectures by Propagating Network Codes
Mingyu Ding, Yuqi Huo, Haoyu Lu, Zhe Wang 0006, Zhiwu Lu 0001, Jingdong Wang 0001, Ping Luo 0002 |
ICLR | 3 |
| 2022 | BMU-MoCo: Bidirectional Momentum Update for Continual Video-Language ModelingabstractVideo-language models suffer from forgetting old/learned knowledge when trained with streaming data. In this work, we thus propose a continual video-language modeling (CVLM) setting, where models are supposed to be sequentially trained on five widely-used video-text datasets with different data distributions. Although most of existing continual learning methods have achieved great success by exploiting extra information (e.g., memory data of past tasks) or dynamically extended networks, they cause enormous resource consumption when transferred to our CVLM setting. To overcome the challenges (i.e., catastrophic forgetting and heavy resource consumption) in CVLM, we propose a novel cross-modal MoCo-based model with bidirectional momentum update (BMU), termed BMU-MoCo. Concretely, our BMU-MoCo has two core designs: (1) Different from the conventional MoCo, we apply the momentum update to not only momentum encoders but also encoders (i.e., bidirectional) at each training step, which enables the model to review the learned knowledge retained in the momentum encoders. (2) To further enhance our BMU-MoCo by utilizing earlier knowledge, we additionally maintain a pair of global momentum encoders (only initialized at the very beginning) with the same BMU strategy. Extensive results show that our BMU-MoCo remarkably outperforms recent competitors w.r.t. video-text retrieval performance and forgetting rate, even without using any extra data or dynamic networks. Yizhao Gao 0004, Nanyi Fei, Haoyu Lu, Zhiwu Lu 0001, Hao Jiang 0022, Zhao Cao |
NeurIPS | 3 |
| 2022 | LGDN: Language-Guided Denoising Network for Video-Language ModelingabstractVideo-language modeling has attracted much attention with the rapid growth of web videos. Most existing methods assume that the video frames and text description are semantically correlated, and focus on video-language modeling at video level. However, this hypothesis often fails for two reasons: (1) With the rich semantics of video contents, it is difficult to cover all frames with a single video-level description; (2) A raw video typically has noisy/meaningless information (e.g., scenery shot, transition or teaser). Although a number of recent works deploy attention mechanism to alleviate this problem, the irrelevant/noisy information still makes it very difficult to address. To overcome such challenge, we thus propose an efficient and effective model, termed Language-Guided Denoising Network (LGDN), for video-language modeling. Different from most existing methods that utilize all extracted video frames, LGDN dynamically filters out the misaligned or redundant frames under the language supervision and obtains only 2--4 salient frames per video for cross-modal token-level alignment. Extensive experiments on five public datasets show that our LGDN outperforms the state-of-the-arts by large margins. We also provide detailed ablation study to reveal the critical importance of solving the noise issue, in hope of inspiring future video-language work. Haoyu Lu, Mingyu Ding, Nanyi Fei, Yuqi Huo, Zhiwu Lu 0001 |
NeurIPS | 1 |
| 2022 | MHCRoBERTa: pan-specific peptide-MHC class I binding prediction through transfer learning with label-agnostic protein sequencesabstractPredicting the binding of peptide and major histocompatibility complex (MHC) plays a vital role in immunotherapy for cancer. The success of Alphafold of applying natural language processing (NLP) algorithms in protein secondary struction prediction has inspired us to explore the possibility of NLP methods in predicting peptide-MHC class I binding. Based on the above motivations, we propose the MHCRoBERTa method, RoBERTa pre-training approach, for predicting the binding affinity between type I MHC and peptides. Analysis of the results on benchmark dataset demonstrates that MHCRoBERTa can outperform other state-of-art prediction methods with an increase of the Spearman rank correlation coefficient (SRCC) value. Notably, our model gave a significant improvement on IC50 value. Our method has achieved SRCC value and AUC value as 0.785 and 0.817, respectively. Our SRCC value is 14.3% higher than NetMHCpan3.0 (the second highest SRCC value on pan-specific) and is 3% higher than MHCflurry (the second highest SRCC value on all methods). The AUC value is also better than any other pan-specific methods. Moreover, we visualize the multi-head self-attention for the token representation across the layers and heads by this method. Through the analysis of the representation of each layer and head, we can show whether the model has learned the syntax and semantics necessary to perform the prediction task well. All these results demonstrate that our model can accurately predict the peptide-MHC class I binding affinity and that MHCRoBERTa is a powerful tool for screening potential neoantigens for cancer immunotherapy. MHCRoBERTa is available as an open source software at github (https://github.com/FuxuWang/MHCRoBERTa). Fuxu Wang, Haoyan Wang, Lizhuang Wang, Haoyu Lu, Shizheng Qiu, Tianyi Zang, Xinjun Zhang, Yang Hu 0008 |
Briefings Bioinform. | 4 |
| 2022 | Image fragile watermarking algorithm based on deneighbourhood mappingabstractAbstract To address the security risk caused by fixed offset mapping and the limited recoverability of random mapping used in image watermarking, a self‐embedding fragile image watermarking algorithm based on deneighbourhood mapping are proposed. First, the image is divided into several 2 × 2 blocks, and authentication watermark and recovery watermark are generated based on the average value of the image blocks. Then, the denighbourhood mapping is implemented as, for each image block, its mapping block is randomly selected outside its neighbourhood. Finally, the authentication watermark and the recovery watermark are embedded into the image block itself and its mapping block. Theoretical analysis indicates that in the case of continuous area tampering, the proposed watermarking algorithm can achieve a better recovery rate than that of the method based on the random mapping. The experimental results verify the rationality and effectiveness of the theoretical analysis. Moreover, compared with the existing embedding algorithms based on random mapping, chaos mapping, and Arnold mapping, in the case of continuous area tampering, the proposed algorithm also achieves a higher average recovery rate. Zhenyu Li 0004, Daofu Gong, Haoyu Lu, Fenlin Liu |
IET Image Process. | 4 |
| 2021 | Self-Supervised Video Representation Learning with Constrained Spatiotemporal JigsawabstractThis paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a representation learned by detecting spatiotemporal continuity/discontinuity is thus beneficial for downstream video content analysis tasks. A natural choice of such a pretext task is to construct spatiotemporal (3D) jigsaw puzzles and learn to solve them. However, as we demonstrate in the experiments, this task turns out to be intractable. We thus propose Constrained Spatiotemporal Jigsaw (CSJ) whereby the 3D jigsaws are formed in a constrained manner to ensure that large continuous spatiotemporal cuboids exist. This provides sufficient cues for the model to reason about the continuity. Instead of solving them directly, which could still be extremely hard, we carefully design four surrogate tasks that are more solvable. The four tasks aim to learn representations sensitive to spatiotemporal continuity at both the local and global levels. Extensive experiments show that our CSJ achieves state-of-the-art on various benchmarks. Yuqi Huo, Mingyu Ding, Haoyu Lu, Mingqian Tang, Zhiwu Lu 0001, Tao Xiang 0002 |
IJCAI | 3 |
| 2021 | Compressed Video Contrastive LearningabstractThis work concerns self-supervised video representation learning (SSVRL), one topic that has received much attention recently. Since videos are storage-intensive and contain a rich source of visual content, models designed for SSVRL are expected to be storage- and computation-efficient, as well as effective. However, most existing methods only focus on one of the two objectives, failing to consider both at the same time. In this work, for the first time, the seemingly contradictory goals are simultaneously achieved by exploiting compressed videos and capturing mutual information between two input streams. Specifically, a novel Motion Vector based Cross Guidance Contrastive learning approach (MVCGC) is proposed. For storage and computation efficiency, we choose to directly decode RGB frames and motion vectors (that resemble low-resolution optical flows) from compressed videos on-the-fly. To enhance the representation ability of the motion vectors, hence the effectiveness of our method, we design a cross guidance contrastive learning algorithm based on multi-instance InfoNCE loss, where motion vectors can take supervision signals from RGB frames and vice versa. Comprehensive experiments on two downstream tasks show that our MVCGC yields new state-of-the-art while being significantly more efficient than its competitors. Yuqi Huo, Mingyu Ding, Haoyu Lu, Nanyi Fei, Zhiwu Lu 0001, Ji-Rong Wen, Ping Luo 0002 |
NeurIPS | 3 |
| 2020 | PremPS: Predicting the impact of missense mutations on protein stabilityabstractComputational methods that predict protein stability changes induced by missense mutations have made a lot of progress over the past decades. Most of the available methods however have very limited accuracy in predicting stabilizing mutations because existing experimental sets are dominated by mutations reducing protein stability. Moreover, few approaches could consistently perform well across different test cases. To address these issues, we developed a new computational method PremPS to more accurately evaluate the effects of missense mutations on protein stability. The PremPS method is composed of only ten evolutionary- and structure-based features and parameterized on a balanced dataset with an equal number of stabilizing and destabilizing mutations. A comprehensive comparison of the predictive performance of PremPS with other available methods on nine benchmark datasets confirms that our approach consistently outperforms other methods and shows considerable improvement in estimating the impacts of stabilizing mutations. A protein could have multiple structures available, and if another structure of the same protein is used, the predicted change in stability for structure-based methods might be different. Thus, we further estimated the impact of using different structures on prediction accuracy, and demonstrate that our method performs well across different types of structures except for low-resolution structures and models built based on templates with low sequence identity. PremPS can be used for finding functionally important variants, revealing the molecular mechanisms of functional influences and protein design. PremPS is freely available at https://lilab.jysw.suda.edu.cn/research/PremPS/, which allows to do large-scale mutational scanning and takes about four minutes to perform calculations for a single mutation per protein with ~ 300 residues and requires ~ 0.4 seconds for each additional mutation. Haoyu Lu, Zefeng Zhu |
PLoS Comput. Biol. | 2 |
| 2019 | Affine invariant image watermarking scheme based on ASIFT and Delaunay tessellation
Liu Feng, Daofu Gong, Fenlin Liu, Haoyu Lu |
Multim. Tools Appl. | 4 |