EDBT 2026 Demo / reviewers in the wild / expert
Wenyang Liu
dblp:141/8954
· DBLP profile ↗
20ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoCo: Conformal Confidence Suppression to Optimize Search ResultsabstractModern e-commerce recommendation systems aim to improve customer experience by ranking content on search results page (SRP). However, displaying content is not always beneficial for customers across all contexts; even top-ranked content can be irrelevant, misleading, or redundant in certain scenarios. In this work, we propose a robust content suppression mechanism to selectively suppress content when necessary. Our approach leverages causal effect learning to measure the incremental value of showing versus suppressing content. Prior work uses Conditional Average Treatment Effect (CATE) to estimate treatment effect without considering the inherent uncertainty in uplift predictions, resulting in over-suppression and degraded customer experience. Our work introduces a novel Conformal Confidence (CoCo) suppressor that explicitly accounts for prediction uncertainty in uplift-based modeling. We first evaluate it on synthetic datasets with controlled noise, bias, and contexts. Results demonstrate superior performance compared to baseline approaches. Subsequently, our online traffic tests show statistically significant improvements in revenue and profit compared to existing methods. Zhou Qin 0001, Yi Liu 0033, Wenyang Liu |
SIGIR | 5 |
| 2026 | Design and Evaluation of Whole-Page Experience Optimization for E-commerce SearchabstractE-commerce Search Results Pages (SRPs) are evolving from linear lists to complex, non-linear layouts, rendering traditional position-biased ranking models insufficient. Moreover, existing optimization frameworks typically maximize short-term signals (e.g., clicks, same-day revenue) because long-term satisfaction metrics (e.g., expected two-week revenue) involve delayed feedback and challenging long-horizon credit attribution. To bridge these gaps, we propose a novel Whole-Page Experience Optimization Framework. Unlike traditional list-wise rankers, our approach explicitly models the interplay between item relevance, 2D positional layout, and visual elements. We use a causal framework to develop metrics for measuring long-term user satisfaction based on quasi-experimental data. We validate our approach through industry-scale A/B testing, where the model demonstrated a 1.86% improvement in brand relevance (our primary customer experience metric) while simultaneously achieving a statistically significant revenue uplift of +0.05%. Pratik Lahiri, Bingqing Ge, Zhou Qin 0001, Aditya Jumde, Shuning Huo, Lucas Scottini, Yi Liu 0033, Mahmoud Mamlouk, Wenyang Liu |
WSDM | 9 |
| 2026 | Corrupted bitstream semantic understanding by adaptive-modal large language models
Kejun Wu, Fangcheng Li, Wenyang Liu |
Pattern Recognit. | 3 |
| 2026 | PromptSR: Cascade Prompting for Lightweight Image Super-ResolutionabstractAlthough the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-based self-attention modeling. The quadratic computational complexity relative to window size restricts its ability to use a large window size for expanding the receptive field while maintaining low computational costs. To address this challenge, we propose PromptSR, a novel prompt-empowered lightweight image SR method. The core component is the proposed cascade prompting block (CPB), which enhances global information access and local refinement via three cascaded prompting layers: a global anchor prompting layer (GAPL) and two local prompting layers (LPLs). The GAPL leverages downscaled features as anchors to construct low-dimensional anchor prompts (APs) through cross-scale attention, significantly reducing computational costs. These APs, with enhanced global perception, are then used to provide global prompts, efficiently facilitating long-range token connections. The two LPLs subsequently combine category-based self-attention and window-based self-attention to refine the representation in a coarse-to-fine manner. They leverage attention maps from the GAPL as additional global prompts, enabling them to perceive features globally at different granularities for adaptive local refinement. In this way, the proposed CPB effectively combines global priors and local details, significantly enlarging the receptive field while maintaining the low computational costs of our PromptSR. The experimental results demonstrate the superiority of our method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations. Our code will be released at https://github.com/wenyang001/PromptSR. Wenyang Liu, Jianjun Gao 0005, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 1 |
| 2025 | From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionabstractRecent Multimodal Large Language Models (MLLMs) exhibit strong zero-shot abilities but struggle with complex Grounded Situation Recognition (GSR) and are resource-intensive for edge device deployment. Meanwhile, conventional GSR models often lack generalization ability, falling short in recognizing unseen and rare situations. In this paper, we exploit transferring knowledge from a teacher MLLM to a small GSR model to enhance its generalization and zero-shot abilities, thereby introducing the task of Open-vocabulary Grounded Situation Recognition (Ov-GSR). To achieve this, we propose Multimodal Interactive Prompt Distillation (MIPD), a novel framework that distills enriched multimodal knowledge from the foundation model, enabling the student Ov-GSR model to recognize unseen situations and be better aware of rare situations. Specifically, the MIPD framework first leverages the LLM-based Judgmental Rationales Generator (JRG) to construct positive and negative glimpse and gaze rationales enriched with contextual semantic information. The proposed scene-aware and instance-perception prompts are then introduced to align rationales with visual information from the MLLM teacher via the Negative-Guided Multimodal Prompting Alignment (NMPA) module, effectively capturing holistic and perceptual multimodal knowledge. Finally, the aligned multimodal knowledge is distilled into the student Ov-GSR model, providing a stronger foundation for generalization that enhances situation understanding, bridges the gap between seen and unseen scenarios, and mitigates prediction bias in rare cases. We evaluate MIPD on the refined Ov-SWiG dataset, achieving superior performance on seen, rare, and unseen situations, and further demonstrate improved unseen detection on the HICO-DET dataset. Jianjun Gao 0005, Wenyang Liu, Kejun Wu, Yi Wang 0068, Soo Chin Liew |
ACM Multimedia | 4 |
| 2025 | SSH-Net: A self-supervised and hybrid network for noisy image watermark removal
Wenyang Liu, Jianjun Gao 0005, Kim-Hui Yap |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | CL-HOI: Cross-level human-object interaction distillation from multimodal large language models
Jianjun Gao 0005, Wenyang Liu, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
Knowl. Based Syst. | 4 |
| 2025 | ByteNet: Rethinking Multimedia File Fragment Classification Through Visual PerspectivesabstractMultimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multimedia storage and communication. Existing MFFC methods typically treat fragments as 1D byte sequences and emphasize the relations between separate bytes (interbytes) for classification. However, the more informative relations inside bytes (intrabytes) are overlooked and seldom investigated. By looking inside bytes, the bit-level details of file fragments can be accessed, enabling a more accurate classification. Motivated by this, we first proposeByte2Image, a novel visual representation model that incorporates previously overlooked intrabyte information into file fragments and reinterprets these fragments as 2D grayscale images. This model involves a sliding byte window to reveal the intrabyte information and a rowwise stacking of intrabyte n-grams for embedding fragments into a 2D space. Thus, complex interbyte and intrabyte correlations can be mined simultaneously using powerful vision networks. Additionally, we propose an end-to-end dual-branch networkByteNetto enhance robust correlation mining and feature representation. ByteNet makes full use of the raw 1D byte sequence and the converted 2D image through a shallow byte branch feature extraction (BBFE) and a deep image branch feature extraction (IBFE) network. In particular, the BBFE, composed of a single fully-connected layer, adaptively recognizes the co-occurrence of several some specific bytes within the raw byte sequence, while the IBFE, built on a vision Transformer, effectively mines the complex interbyte and intrabyte correlations from the converted image. Experiments on the two representative benchmarks, including 14 cases, validate that our proposed method outperforms state-of-the-art approaches on different cases by up to 12.2%. Wenyang Liu, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 1 |
| 2024 | Empowering Large Language Model for Continual Video Question Answering with Collaborative PromptingabstractIn recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content.In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting.To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting.These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research.Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14% accuracy on NExT-QA and 71.24% accuracy on DramaQA, highlighting its practical relevance and effectiveness. Jianjun Gao 0005, Wenyang Liu, Runzhong Zhang, Kim-Hui Yap |
EMNLP | 4 |
| 2024 | CM2-Net: Continual Cross-Modal Mapping Network For Driver Action RecognitionabstractDriver action recognition has significantly advanced in enhancing driver-vehicle interactions and ensuring driving safety by integrating multiple modalities, such as infrared and depth. Nevertheless, compared to RGB modality only, it is always laborious and costly to collect extensive data for all types of non-RGB modalities in car cabin environments. Therefore, previous works have suggested independently learning each non-RGB modality by fine-tuning a model pretrained on RGB videos, but these methods are less effective in extracting informative features when faced with newly-incoming modalities due to large domain gaps. In contrast, we propose a Continual Cross-Modal Mapping Network (CM2Net) to continually learn each newly-incoming modality with instructive prompts from the previously-learned modalities. Specifically, we have developed Accumulative Cross-modal Mapping Prompting (ACMP), to map the discriminative and informative features learned from previous modalities into the feature space of newly-incoming modalities. Then, when faced with newly-incoming modalities, these mapped features are able to provide effective prompts for which features should be extracted and prioritized. These prompts are accumulating throughout the continual learning process, thereby boosting further recognition performances. Extensive experiments conducted on the Drive&Act dataset demonstrate the performance superiority of $\mathrm{CM}^{2}-\mathrm{Net}$ on both uni- and multi-modal driver action recognition. Jianjun Gao 0005, Dan Lin 0008, Wenyang Liu, Kim-Hui Yap |
ICIP | 6 |
| 2024 | Intra- and inter-sector contextual information fusion with joint self-attention for file fragment classificationabstractFile fragment classification (FFC) aims to identify the file type of file fragments in memory sectors, which is of great importance in memory forensics and information security. Existing works focused on processing the bytes within sectors separately and ignoring contextual information between adjacent sectors. In this paper, we introduce a joint self-attention network (JSANet) for FFC to learn intra-sector local features and inter-sector contextual features. Specifically, we propose an end-to-end network with the byte, channel, and sector self-attention modules. Byte self-attention adaptively recognizes the intra-sector significant bytes, and channel self-attention re-calibrates the features between channels. Based on the insight that adjacent memory sectors are most likely to store a file fragment, sector self-attention leverages contextual information in neighboring sectors to enhance inter-sector feature representation. Extensive experiments on seven FFC benchmarks show the superiority of our method compared with state-of-the-art methods. Moreover, we construct VFF-16, a variable-length file fragment dataset to reflect file fragmentation. Integrated with sector self-attention, our method improves accuracy by more than 16.3% against the baseline on VFF-16, and the runtime achieves 5.1 s/GB with GPU acceleration. In addition, we extend our model to malware detection and show its applicability. Yi Wang 0068, Wenyang Liu, Kejun Wu, Kim-Hui Yap, Lap-Pui Chau |
Knowl. Based Syst. | 2 |
| 2023 | Bitstream-Corrupted JPEG Images are Restorable: Two-stage Compensation and Alignment Framework for Image RestorationabstractIn this paper, we study a real-world JPEG image restoration problem with bit errors on the encrypted bitstream. The bit errors bring unpredictable color casts and block shifts on decoded image contents, which cannot be resolved by existing image restoration methods mainly relying on pre-defined degradation models in the pixel domain. To address these challenges, we propose a robust JPEG decoder, followed by a two-stage compensation and alignment framework to restore bitstream-corrupted JPEC images. Specifically, the robust JPEC decoder adopts an error-resilient mechanism to decode the corrupted JPEG bitstream. The two-stage framework is composed of the self-compensation and alignment (SCA) stage and the guided-compensation and alignment (GCA) stage. The SCA adaptively performs block-wise image color compensation and alignment based on the estimated color and block offsets via image content similarity. The GCA leverages the extracted low-resolution thumbnail from the JPEG header to guide full-resolution pixel-wise image restoration in a coarse-to-fine manner. It is achieved by a coarse-guided pix2pix network and a refine-guided bi-directional Laplacian pyramid fusion network. We conduct experiments on three benchmarks with varying degrees of bit error rates. Experimental results and ablation studies demonstrate the superiority of our proposed method. The code will be released at https://github.com/wenyang001/Two-ACIR. Wenyang Liu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
CVPR | 1 |
| 2023 | A Spatial-Focal Error Concealment Scheme for Corrupted Focal Stack VideoabstractFocal stack image sequences can be regarded as successive frames of videos, which are densely captured by focusing on a stack of focal planes. This type of data is able to provide focus cues for display technologies. Before the displays on the user side, focal stack video is possibly corrupted during compression, storage and transmission chains, generating error frames on the decoder side. The error regions are difficult to be recovered due to the focal changes among frames. Conventional error concealment methods result in sharpness inconsistency between recovered regions and their spatial adjacent regions. Motivated by this, in this paper, we propose a spatial-focal error concealment scheme specialized for focal stack videos. The spatial adjacent regions around an error region are employed to reveal the prediction relations between error frame and focal adjacent frames. Gaussian blur filtering and Lucy-Richardson deblur filtering are applied to simulate the video focal changes. In this way, the error regions can be well recovered by exploiting the spatial-focal information. Experiment results show that the proposed scheme can achieve the highest objective quality in terms of PSNR and SSIM. It can also obtain the best subjective quality with sharpness consistency in recovered regions and without block effect. Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
DCC | 3 |
| 2023 | Image Representation and Deep Inception-Attention for File-type and Malware ClassificationabstractFile-type classification aims to recognize the file types of files/fragments without file-system metadata, which is essential for memory forensics and data recovery. In this paper, we introduce an image representation and deep inception-attention manner for file-type classification. Specifically, we consider file-type classification as an image classification problem. Raw data sequences in the memory block are converted to 2D binary images, enriching the representation ability and visualization while retaining the completeness of the bitstream. With binary images as inputs, we propose a deep inception-attention network to extract discriminate horizontal features and re-calibrate the weights of feature maps, and finally, predict file types. Experiments on a large-scale benchmark show the superiority of the proposed model. Moreover, our method can be extended to a similar application, like malware classification, and achieve outstanding performance. Yi Wang 0068, Kejun Wu, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
ISCAS | 3 |
| 2023 | Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and MethodabstractThe past decade has witnessed great strides in video recovery by specialist technologies, like video inpainting, completion, and error concealment. However, they typically simulate the missing content by manual-designed error masks, thus failing to fill in the realistic video loss in video communication (e.g., telepresence, live streaming, and internet video) and multimedia forensics. To address this, we introduce the bitstream-corrupted video (BSCV) benchmark, the first benchmark dataset with more than 28,000 video clips, which can be used for bitstream-corrupted video recovery in the real world. The BSCV is a collection of 1) a proposed three-parameter corruption model for video bitstream, 2) a large-scale dataset containing rich error patterns, multiple corruption levels, and flexible dataset branches, and 3) a new video recovery framework that serves as a benchmark. We evaluate state-of-the-art video inpainting methods on the BSCV dataset, demonstrating existing approaches' limitations and our framework's advantages in solving the bitstream-corrupted video recovery problem. The benchmark and dataset are released at https://github.com/LIUTIGHE/BSCV-Dataset. Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
NeurIPS | 4 |
| 2022 | Automate Page Layout Optimization: An Offline Deep Q-Learning ApproachabstractThe modern e-commerce web pages have brought better customer experience and more profitable services by whole page optimization at different granularity, e.g., page layout optimization, item ranking optimization, etc. Generating the proper page layout per customer’s request is one of the vital tasks during the web page rendering process, which can directly impact customers’ shopping experience and their decision-making. In this paper, we formulate the request-rendering interactions as a Markov decision process (MDP) and solve it by deep reinforcement learning (RL). Specifically, we present the design and implementation of applying offline Deep Q-Learning (DQN) to the contextual page layout optimization problem. Through the offline evaluation method, we demonstrate the effectiveness of the proposed framework, i.e., the RL agent has the potential to perform better than the baseline ranker by learning from the offline data set, e.g., the RL agent can improve the average cumulative rewards up to 36.69% comparing to the baseline ranker. Zhou Qin 0001, Wenyang Liu |
RecSys | 2 |
| 2020 | MindReading: An Ultra-Low-Power Photonic Accelerator for EEG-based Human Intention RecognitionabstractA scalp-recording electroencephalography (EEG)-based brain-computer interface (BCI) system can greatly improve the quality of life for people who suffer from motor disabilities. Deep neural networks consisting of multiple convolutional, LSTM and fully-connected layers are created to decode EEG signals to maximize the human intention recognition accuracy. However, prior FPGA, ASIC, ReRAM and photonic accelerators cannot maintain sufficient battery lifetime when processing realtime intention recognition. In this paper, we propose an ultra-low-power photonic accelerator, MindReading, for human intention recognition by only low bit-width addition and shift operations. Compared to prior neural network accelerators, to maintain the real-time processing throughput, MindReading reduces the power consumption by 62.7% and improves the throughput per Watt by 168%. Qian Lou, Wenyang Liu, Weichen Liu 0001, Lei Jiang 0001 |
ASP-DAC | 2 |
| 2019 | HolyLight: A Nanophotonic Accelerator for Deep Learning in Data CentersabstractConvolutional Neural Networks (CNNs) are widely adopted in object recognition, speech processing and machine translation, due to their extremely high inference accuracy. However, it is challenging to compute massive computationally expensive convolutions of deep CNNs on traditional CPUs and GPUs. Emerging Nanophotonic technology has been employed for on-chip data communication, because of its CMOS compatibility, high bandwidth and low power consumption. In this paper, we propose a nanophotonic accelerator, HolyLight, to boost the CNN inference throughput in datacenters. Instead of an all-photonic design, HolyLight performs convolutions by photonic integrated circuits, and process the other operations in CNNs by CMOS circuits for high inference accuracy. We first build HolyLight-M by microdisk-based matrix-vector multipliers. We find analog-to-digital converters (ADCs) seriously limit its inference throughput per Watt. We further use microdisk-based adders and shifters to architect HolyLight-A without ADCs. Compared to the state-of-the-art ReRAM-based accelerator, HolyLight-A improves the CNN inference throughput per Watt by 13× with trivial accuracy degradation. Weichen Liu 0001, Wenyang Liu, Yichen Ye, Qian Lou, Yiyuan Xie, Lei Jiang 0001 |
DATE | 2 |
| 2018 | Fine-Grained Task-Level Parallel and Low Power H.264 Decoding in Multi-Core SystemsabstractIn the past few years, the extinction of Moore's Law makes people reconsider the solutions for dealing with the low computing resource utilization of applications on multicore processor systems. However, making good use of computing resources in multi-core processors systems is not easy due to the differences between single-core and multi-core architecture. Nowadays short video apps like Instagram and Tik Tok have successfully caught people's eyes by fascinating short videos, typically just 10 to 30 seconds long, uploaded by the users of apps. And almost all of these videos are recorded by their mobile devices, which are typically HD (High Definition) or FHD (Full High Definition) videos, which prefer to be encoded/decoded by H.264/AVC rather then HEVC (High Efficiency Video Coding) on mobile devices in view of the energy consumption and decoding speed. How to dive the huge potential of the computing resource on multi-core mobile devices to speed up decoding these videos while consuming low energy, is a big challenge. In our previous work [1], a relatively simple parallel framework was proposed to implement a parallel H.264/ AV C decoder. This work further proposes a more detailed systematic task-level parallel framework, together with an energy saving strategy based on this framework, to research a new H.264/AVC decoder on multi-core processor systems. The proposed parallel method is composed of a set of rules to guide parallel software programming (PSPR) and a software parallelization framework (SPF). The PSPR is applied in pre-processing steps to address the potential issues limiting the inherent parallelism, and the SPF is applied to parallelize the original serial programs. After the parallelization is successfully deployed, DVFS technique would be applied to decrease the power dissipation based on the SPF. Results show that proposed solutions make a significant improvement in decoding speed of 32% at 720p, 27% at 1080p and 29% at 2160p, and in energy savings of 25% at 720p, 25% at 1080p and 23% at 2160p on a four-core workstation running Linux, compared to the original serial H.264/ AV C decoder. The results demonstrate our methods are effective and scalable, served as a reference for future parallel software development. Wenyang Liu, Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye |
ICPADS | 1 |
| 2016 | CP-ABE with outsourced decryption and directionally hidden policyabstractAbstract Ciphertext‐policy attribute‐based encryption (CP‐ABE) is a novel cryptographic primitive for access controlling. However, the existing CP‐ABE schemes are very inefficient as the decryptions involve many expensive pairing operations. Another drawback is that the access policy itself may disclose some privacies of the users. In certain applications, access structures also should be protected. In this work, we propose a notion of CP‐ABE with outsourced decryption and directionally hidden policy, which allows a semi‐trusted proxy in the cloud to help a user decrypt a ciphertext, but the proxy cannot learn the plaintext and the access policy. We construct a concrete scheme from Waters' CP‐ABE and prove its security, verifiability, and directionally hidden policy. Copyright © 2016 John Wiley & Sons, Ltd. Zhiwei Wang 0003, Wenyang Liu |
Secur. Commun. Networks | 2 |