EDBT 2026 Demo / reviewers in the wild / expert
Lizhong Wang
dblp:70/4180
· DBLP profile ↗
17ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforced Rate Control for Neural Video Compression via Inter-Frame Rate-Distortion AwarenessabstractNeural video compression (NVC) has demonstrated superior compression efficiency, yet effective rate control remains a significant challenge due to complex temporal dependencies. Existing rate control schemes typically leverage frame content to capture distortion interactions, overlooking inter-frame rate dependencies arising from shifts in per-frame coding parameters. This often leads to suboptimal bitrate allocation and cascading parameter decisions. To address this, we propose a reinforcement‑learning (RL)‑based rate control framework that formulates the task as a frame‑by‑frame sequential decision process. At each frame, an RL agent observes a spatiotemporal state and selects coding parameters to optimize a long‑term reward that reflects rate‑distortion (R-D) performance and bitrate adherence. Unlike prior methods, our approach jointly determines bitrate allocation and coding configuration in a single step, independent of group‑of‑pictures (GOP) structure. Extensive experiments across diverse NVC architectures show that our method reduces the average relative bitrate error to 1.20 percent and achieves up to 13.45 percent bitrate savings at typical GOP sizes, outperforming existing approaches. In addition, our framework demonstrates improved robustness to content variation and bandwidth fluctuations with lower encoding/decoding overhead, making it highly suitable for practical deployment. Wuyang Cong, Junqi Shi, Lizhong Wang, Weijing Shi, Ming Lu 0003, Hao Chen 0036, Zhan Ma 0001 |
AAAI | 3 |
| 2025 | Integrating Adaptive Sampling for Optimal Learned Video CompressionabstractWe propose a novel adaptive prediction network that dynamically determines the optimal sampling factor and Lagrangian multiplier for encoding each frame, guided by sequential information. By exploiting spatio-temporal redundancy through adaptive sampling, our method reduces bitrate consumption while preserving the reconstruction quality with adjusted rate-distortion coefficients. Experimental results demonstrate significant performance gains over representative learned video compression models across various datasets, with reduced encoding and decoding latency. Wuyang Cong, Yuzhuo Kong, Ming Lu 0003, Lizhong Wang, Weijing Shi, Zhan Ma 0001 |
ICASSP | 4 |
| 2025 | SPCODEC: Split and Prediction for Neural Speech Codec
Liang Wen, Lizhong Wang, Yuxing Zheng, Weijing Shi, Kwangpyo Choi |
INTERSPEECH | 2 |
| 2025 | Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D EnvironmentsabstractAdversarial attacks in 3D environments have emerged as a critical threat to the reliability of visual perception systems, particularly in safety-sensitive applications such as identity verification and autonomous driving. These attacks employ adversarial patches and 3D objects to manipulate deep neural network (DNN) predictions by exploiting vulnerabilities within complex scenes. Existing defense mechanisms, such as adversarial training and purification, primarily employ passive strategies to enhance robustness. However, these approaches often rely on pre-defined assumptions about adversarial tactics, limiting their adaptability in dynamic 3D settings. To address these challenges, we introduce Reinforced Embodied Active Defense (Rein-EAD), a proactive defense framework that leverages adaptive exploration and interaction with the environment to improve perception robustness in 3D adversarial contexts. By implementing a multi-step objective that balances immediate prediction accuracy with predictive entropy minimization, Rein-EAD optimizes defense strategies over a multi-step horizon. Additionally, Rein-EAD involves an uncertainty-oriented reward-shaping mechanism that facilitates efficient policy updates, thereby reducing computational overhead and supporting real-world applicability without the need for differentiable environments. Comprehensive experiments validate the effectiveness of Rein-EAD, demonstrating a substantial reduction in attack success rates while preserving standard accuracy across diverse tasks. Notably, Rein-EAD exhibits robust generalization to unseen and adaptive attacks, making it suitable for real-world complex tasks, including 3D object classification, face recognition and autonomous driving. By integrating proactive policy learning with embodied scene interaction, Rein-EAD establishes a scalable and adaptable approach for securing DNN-based perception systems in dynamic and adversarial 3D environments. Xiao Yang 0028, Lingxuan Wu, Lizhong Wang, Chengyang Ying, Hang Su 0006, Jun Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | FT-CSR: Cascaded Frequency-Time Method for Coded Speech RestorationabstractLossy speech codecs often introduce coding distortions such as coding noise and constrained bandwidth, which can affect the quality of the decoded speech. This paper proposes a method called FT-CSR, which is used for coded speech restoration. FT-CSR reduces coding noise and recovers missing frequencies sequentially using a cascaded frequency-time domain model. In experiments using the Opus codec, FT-CSR was found to be effective across bitrates ranging from 8 to 16 kbps and outperformed the baseline on both objective and subjective measurements. FT-CSR achieves a MOS-POLQA score of 3.6 or higher and improves MOS-POLQA by more than 0.23 when compared to decoded speech. The results of the subjective test show that FT-CSR can improve MOS by over 0.85 for decoded speech. Liang Wen, Lizhong Wang, Yuxing Zheng, Weijing Shi, Kwangpyo Choi |
ICME | 2 |
| 2023 | Semi-Supervised Sound Event Detection with Pre-Trained ModelabstractSound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. In this paper, we focus on the semi-supervised SED task, and combine pre-trained model from other field to assist in improving the detection effect. Pre-trained models have been widely used in various tasks in the field of speech, such as automatic speech recognition, audio tagging, etc. If the training dataset is large and general enough, the embedding features extracted by the pre-trained model will cover the potential information in the original task. We use pre-trained model PANNs which is suitable for SED task and proposed two methods to fuse the features from PANNs and original model, respectively. In addition, we also propose a weight raised temporal contrastive loss to improve the model’s switching speed at event boundaries and the smoothness within events. Experimental results show that using pre-trained model features outperforms the baseline by 8.5% and 9.1% in DESED public evaluation dataset in terms of polyphonic sound detection score (PSDS). Lizhong Wang, Sijun Bi, Jing Wang 0037 |
ICASSP | 2 |
| 2023 | Distortion-Aware Convolutional Neural Network-Based Interpolation Filter for AVS3abstractMotion compensation is a key technology in video coding for removing the temporal redundancy between video frames. Considering the incompatibility between traditional interpolation filters and diversified video content, the inter prediction method still has considerable room for improvement. This paper proposed a distortion-aware convolutional neural network-based interpolation filter (DA-NNIF) to further improve the interpolation prediction accuracy of sub-pixels with one model. Distortion parameters are introduced into the proposed network to reflect the quantization noise of reference frames. The experimental result shows that the proposed method achieves on average 1.47 % BD-rate reduction on Y component for ClassB, ClassC and ClassD sequences under the random access configuration of AVS3. Liang Wen, Lizhong Wang, Yinji Piao, Weijing Shi, Kwangpyo Choi |
ICASSP | 3 |
| 2023 | NEOP: A Framework for Distributed Mobile Apps on Heterogeneous DevicesabstractToday’s apps on a mobile device, such as a smartphone and a tablet, need to access various resources to deliver quality service to users’ satisfaction. These resources may include cameras, microphones, screens, processors, various specialized sensors, and data. In today’s client-server framework, resources accessible to an app are limited to those available in the device running the app, on the cloud, and likely in a few statically connected devices. However, there can be abundant resources on devices near the app-running one with desirable functionalities that can enable or empower the app’s new features and services, but cannot be easily accessed and leveraged. The NEOP (Neutron Operation Platform) framework is an app development and execution environment that removes the barrier across the devices. Heterogeneous IoT devices make the capabilities in their hardware and service software available after security and privacy authentication. An app is developed as a composition of capabilities distributed across various end devices and the cloud. Its constituent computing tasks can be dynamically created and scheduled. Different device capabilities can be selectively and dynamically recruited into the app for the optimal user experience. In this paper we describe example scenarios that motivate the next-generation app framework, the framework’s architecture, design principles, technical challenges, and details on its design and implementation. We also compare this work with related efforts on distributed mobile computing to highlight the unique contributions made by the NEOP platform. Song Jiang 0001, Weidong Zhong, Lizhong Wang, Xiao-Feng Li |
ISADS | 4 |
| 2022 | Human Sound Classification based on Feature Fusion Method with Air and Bone Conducted Signal
Jing Wang 0037, Lizhong Wang, Sijun Bi, Jianqian Zhang, Qiuyue Ma |
INTERSPEECH | 3 |
| 2022 | Multi-Stage Progressive Audio Bandwidth ExtensionabstractAudio bandwidth extension can enhance subjective sound quality by increasing bandwidth of audio signal. This paper presents a novel multi-stage progressive method for time domain causal bandwidth extension. Each stage of the progressive model contains a light weight scale-up module to generate high frequency signal and a supervised attention module to guide features propagating between stages. Time-frequency two-step training method with weighted loss for progressive output is adopted to supervise bandwidth extension performance improves along stages. Test results show that multi-stage model can improve both objective results and perceptual quality progressively. The multi-stage progressive model makes bandwidth extension performance adjustable according to energy consumption, computing capacity and user preferences. Liang Wen, Lizhong Wang, Kwangpyo Choi |
SLT | 2 |
| 2021 | X-net: A Joint Scale Down and Scale Up Method for Voice Call
Liang Wen, Lizhong Wang, Yuxing Zheng, Youngo Park, Kwangpyo Choi |
Interspeech | 2 |
| 2014 | Embedded Stereo Coding Algorithm for Stereo Voice CodecabstractIn this paper, an embedded stereo coding algorithm is proposed. It is developed for the multiple participants video conference. The proposed algorithm is an embedded stereo tool, which can be grouped with any mono voice codec (e.g. AMR-WB, SILK etc.). Subband coding and transform coding methods are used for the proposed stereo tool. Simulation results show that the proposed the stereo voice codec algorithm has a better performance and lower delay in contrast with AMR-WB+ standard. Lizhong Wang, Jun Seong Hong |
IC2E | 2 |
| 2014 | Human-centric analysis and interpretation of time series: a perspective of granular computingabstractIn spite of the truly remarkable diversity of models of time series, there is still an evident need to develop constructs whose accuracy and interpretability are carefully identified and reconciled subsequently leading to highly interpretable (human-centric) constructs. While a great deal of research has been devoted to the design of nonlinear numeric models of time series (with an evident objective to achieve high accuracy of prediction), an issue of interpretability (transparency) of models of time series becomes an evident and ongoing challenge. The user-friendliness of models of time series comes with an ability of humans to perceive and process abstract constructs rather than dealing with plain numeric entities. In perception of time series, information granules (which are regarded as realizations of interpretable entities) play a pivotal role. This gives rise to a concept of granular models of time series or granular time series, in brief. This study revisits generic concepts of information granules and elaborates on a fundamental way of forming information granules (both sets—intervals as well as fuzzy sets) through applying a principle of justifiable granularity encountered in granular computing. Information granules are discussed with regard to the granulation of time series in a certain predefined representation space (viz. a feature space) and granulation carried out in time. The granular representation and description of time series is then presented. We elaborate on the fundamental hierarchically organized layers of processing supporting the development and interpretation of granular time series, namely (a) formation of granular descriptors used in their visualization, (b) construction of linguistic descriptors used afterwards in the generation of (c) linguistic description of time series. The layer of the linguistic prediction models of time series exploiting the linguistic descriptors is outlined as well. A number of examples are offered throughout the entire paper with intent to illustrate the main functionalities of the essential layers of the granular models of time series. Witold Pedrycz, Wei Lu 0005, Xiaodong Liu 0001, Wei Wang 0036, Lizhong Wang |
Soft Comput. | 5 |
| 2009 | Risk Management of Liability Uncertainties to Facilitate Brownfield Redevelopment: Comparing the Situation of Canada with the USabstractThe uncertainties associated with liability in brownfield redevelopment and relevant risk management tools are discussed based on a comparative review of the situation in Canada and the United States. The changes of regulations and policies, inherent uncertainties of site assessment and remediation techniques, incidents of contaminant transport and exposure, and variations of economic and financial conditions all lead to the uncertainties of environmental liability. The fear of liability especially the associated uncertainties is the key obstacle for owners or developers to undertake cleanup and redevelopment due to the subsequent unpredictability of economic profitability. Various risk management tools have been gradually developed in the last three decades to address liability uncertainties, both within Canada and the US, among which environmental insurance and innocent owner's shelter from liability are the two most viable instruments to reduce the fear of liability. Policy making and risk management have evolved more slowly in Canada than in the US and there is a trend for Canadian provinces to adopt the successful policies and tools used in the US rather than formulate their own. Lizhong Wang, Liping Fang, Keith W. Hipel |
SMC | 1 |
| 2007 | A game-theoretic approach to brownfield redevelopment: negotiation on cost and benefit allocationabstractA game-theoretic approach to brownfield redevelopment is presented to support negotiations on related cost and benefit allocation. The process of brownfield redevelopment is discussed and three types of key decision makers (DMs) are identified, namely the landowner, the developer and the government. Based on a qualitative analysis of coalitions among the key DMs, a cooperative game model is formulated for the brownfield redevelopment cost and benefit allocation problem. To initiate the negotiation on the cost and benefit allocation, a common understanding of the estimates of relevant costs and benefits in a brownfield redevelopment project is required. The proposed modeling framework utilizes various nucleolus concepts and the Shapley value to find alternative fair cost and benefit allocations from which the DMs can make the final selection. An illustrative example is used to demonstrate the application of the proposed approach. Lizhong Wang, Liping Fang, Keith W. Hipel |
SMC | 1 |
| 2006 | Modeling Water Rights Allocation in the South Saskatchewan River Basin in CanadaabstractTwo multiple objective optimization methods for water rights allocation, the priority-based maximal multiperiod network flow programming (PMMNF) method and lexicographic minimax water shortage ratios (LMWSR) method are applied to the South Saskatchewan river basin in southern Alberta, Canada. Population growth and economic expansion are straining water resources in this basin. Irrigation, municipal and industrial uses, and hydropower generation are the major water demands competing for limited water resources. The case study shows that the PMMNF method usually produces large differences in the satisfaction ratios among uses in times of water shortages since it is based on the priority concept. On the other hand, the allocations by LMWSR lead to fairer distributed satisfaction ratios because it adopts the lexicographic minimax equity concept and thus has water shortages shared among all the water demands of various types of uses. Lizhong Wang, Liping Fang, Keith W. Hipel |
SMC | 1 |
| 2003 | Cooperative water resource allocation based on equitable water rightsabstractA comprehensive methodology is developed to achieve equitable, efficient and sustainable water allocation in a river basin. Water allocation is carried out in two steps: (1) allocation of initial water rights to water stakeholders and users based on existing water rights systems or agreements; and (2) reallocation of water and net benefits through water transfers to promote equitable cooperation of all stakeholders in a river basin and to achieve efficient use of water. Three methods, prior water rights, modified riparian water rights and lexicographic minimax demand shortage ratios, are proposed for the equitable initial water rights allocation. Using cooperative game theory carries the associated water transfers and net benefit reallocation. Lizhong Wang, Liping Fang, Keith W. Hipel |
SMC | 1 |