VLDB 2026 Research / reviewers in the wild / expert
Yifan Liang
dblp:03/3688
· DBLP profile ↗
18ranked-venue papers
14as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Computer networks · 6 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation TasksabstractIn the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does SNR fail in measuring audio quality? And how to improve its reliability as an objective metric? In this paper, we identify the inadequate measurement of phase distance as a pivotal factor and propose to reformulate SNR with specially designed phase-distance terms, yielding an improved metric named GOMPSNR. We further extend the newly proposed formulation to derive two novel categories of loss function, corresponding to magnitude-guided phase refinement and joint magnitude-phase optimization, respectively. Besides, extensive experiments are conducted for an optimal combination of different loss functions. Experimental results on advanced neural vocoders demonstrate that our proposed GOMPSNR exhibits more reliable error measurement than SNR. Meanwhile, our proposed loss functions yield substantial improvements in model performance, and our well-chosen combination of different loss functions further optimizes the overall model capability. Lingling Dai, Andong Li, Yifan Liang, Xiaodong Li 0002, Chengshi Zheng |
AAAI | 4 |
| 2026 | SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisabstractAlthough lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in this task remains largely unexplored. In this paper, we introduce SLD-L2S, a novel L2S framework built upon a hierarchical subspace latent diffusion model. Our method aims to directly map visual lip movements to the continuous latent space of a pre-trained neural audio codec, thereby avoiding the information loss inherent in traditional intermediate representations. The core of our method is a hierarchical architecture that processes visual representations through multiple parallel subspaces, initiated by a subspace decomposition module. To efficiently enhance interactions within and between these subspaces, we design the diffusion convolution block (DiCB) as our network backbone. Furthermore, we employ a reparameterized flow matching technique to directly generate the target latent vectors. This enables a principled inclusion of speech language model (SLM) and semantic losses during training, moving beyond conventional flow matching objectives and improving synthesized speech quality. Our experiments show that SLD-L2S achieves state-of-the-art generation quality on multiple benchmark datasets, surpassing existing methods in both objective and subjective evaluations. Yifan Liang, Andong Li, Guochen Yu, Fangkun Liu, Lingling Dai, Xiaodong Li 0002, Chengshi Zheng |
AAAI | 1 |
| 2026 | Ground-to-air collaborative LiDAR global localization in forest environments
Yifan Liang, Jingbin Liu, Jietao Lei, Jesse Muhojoki, Antero Kukko, Harri Kaartinen, Juha Hyyppä, Dong Xu 0011 |
Expert Syst. Appl. | 1 |
| 2026 | NaturalL2S: End-to-end high-quality multispeaker lip-to-speech synthesis with differential digital signal processing
Yifan Liang, Fangkun Liu, Andong Li, Xiaodong Li 0002, Chengyou Lei, Chengshi Zheng |
Neural Networks | 1 |
| 2026 | OmniControl: Unified Audio Extraction and Elimination via Subband-Aware Separation
Yifan Liang, Andong Li, Xiaodong Li 0002, Chengshi Zheng |
IEEE Signal Process. Lett. | 1 |
| 2025 | Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution DetectionabstractVision-language models (VLMs), such as CLIP, have shown remarkable capabilities in downstream tasks. However, the coupling of semantic information between the foreground and the background in images leads to significant shortcut issues that adversely affect out-of-distribution (OOD) detection abilities. When confronted with a background OOD sample, VLMs are prone to misidentifying it as in-distribution (ID) data. In this paper, we analyze the OOD problem from the perspective of shortcuts in VLMs and propose OSPCoOp which includes background decoupling and mask-guided region regularization. We first decouple images into ID-relevant and ID-irrelevant regions and utilize the latter to generate a large number of augmented OOD background samples as pseudo-OOD supervision. We then use the masks from background decoupling to adjust the model’s attention, minimizing its focus on ID-irrelevant regions. To assess the model’s robustness against background interference, we introduce a new OOD evaluation dataset, ImageNet-Bg, which solely consists of background images with all ID-relevant regions removed. Our method demonstrates exceptional performance in few-shot scenarios, achieving strong results even in one-shot setting, and outperforms existing methods. The code and proposed ImageNet-Bg are available at https://github.com/HAIV-Lab/OSPCoOp_Imagenet-bg. Xiang Xiang 0001, Yifan Liang |
CVPR | 3 |
| 2025 | LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis for Multi-Speaker Scenarios
Yifan Liang, Fangkun Liu, Andong Li, Xiaodong Li 0002, Chengshi Zheng |
INTERSPEECH | 1 |
| 2023 | LOS Signal Identification for Passive Multi-Target Localization in Multipath EnvironmentsabstractThis letter examines line-of-sight (LOS) path identification for passive multi-target localization in multipath environments. We consider a system comprising multiple spatially distributed sensors, each transmitting a distinct waveform and using the echoes to measure the LOS and non-LOS (NLOS) delays (i.e., ranges) of the targets in the surveillance area. For simplicity, we assume a 2-D localization scenario, where each range measurement defines a circle, and measurements from different sensors create intersection points on the plane. The problem is to identify intersections that are created by LOS paths. To solve the problem, we classify the intersections into$N(N-1)/2$types, where$N$denotes the number of sensors. Then, an efficient clustering algorithm is proposed to efficiently identify the LOS intersections based on the type and other related attributes. Numerical results are presented to demonstrate the performance of the proposed technique in comparison with several peer methods. Yifan Liang, Hongbin Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Team Formation and Task Recommendation for Disabled People in Crowdsourcing SystemsabstractThe barrier-free map is a new type of map, the demand for which is increasing in modern society, and it is necessary to rely on crowdsourcing platforms to complete the labeling work of barrier-free facilities. Nowadays more people with disabilities are participating in crowdsourcing work. They tend to work with familiar people and hope to maximize their value. However, there is very little research about automatic management of organizations and tasks for the disabled to assist them to complete the work on the crowdsourcing platform. In this paper, we design a new team formation method based on the theory of socially embedded work of the disabled (TFSEW) and a bi-objective task recommendation method (BOTR), in order to ensure the efficiency of task completion and promote the participation of the disabled, thereby realizing their value in the work environment. Our method proves to be superior through experiments. At the same time, through questionnaire, people with disabilities show great anticipation for participating in the tasks in form of team and receive tasks recommended by the system. Yifan Liang, Zongzheng Zou, Peng Zhang 0060, Dongsheng Li 0002, Tun Lu, Ning Gu 0001 |
CSCWD | 1 |
| 2021 | A novel error-correcting output codes based on genetic programming and ternary digit operators
Yifan Liang, Hanrui Wang 0001, Kunhong Liu 0001, Jun-Feng Yao, Yingying She, Guiming Dai, Yuna Okina |
Pattern Recognit. | 1 |
| 2010 | Generalizing capacity: new definitions and capacity theorems for composite channelsabstractWe consider three capacity definitions for composite channels with channel side information at the receiver. A composite channel consists of a collection of different channels with a distribution characterizing the probability that each channel is in operation. TheShannon capacityof a channel is the highest rate asymptotically achievable with arbitrarily small error probability. Under this definition, the transmission strategy used to achieve the capacity must achieve arbitrarily small error probability for all channels in the collection comprising the composite channel. The resulting capacity is dominated by the worst channel in its collection, no matter how unlikely that channel is. We, therefore, broaden the definition of capacity to allow for some outage. The capacity versus outage is the highest rate asymptotically achievable with a given probability of decoder-recognized outage. Theexpected capacityis the highest average rate asymptotically achievable with a single encoder and multiple decoders, where channel side information determines the channel in use. The expected capacity is a generalization of capacity versus outage since codes designed for capacity versus outage decode at one of two rates (rate zero when the channel is in outage and the target rate otherwise) while codes designed for expected capacity can decode at many rates. Expected capacity equals Shannon capacity for channels governed by a stationary ergodic random process but is typically greater for general channels. The capacity versus outage and expected capacity definitions relax the constraint that all transmitted information must be decoded at the receiver. We derive channel coding theorems for these capacity definitions through information density and provide numerical examples to highlight their connections and differences. We also discuss the implications of these alternative capacity definitions for end-to-end distortion, source-channel coding, and separation. Michelle Effros, Andrea J. Goldsmith, Yifan Liang |
IEEE Trans. Inf. Theory | 3 |
| 2008 | Generalized Capacity and Source-Channel Coding for Packet Erasure ChannelsabstractWe study the transmission of a stationary ergodic Gaussian source over a packet erasure channel, which is a composite channel with degraded states. A broadcast channel code can be applied to a composite channel to obtain different rates in the different channel states. However, we show that a non-broadcast direct transmission strategy achieves a higher expected rate than the broadcast code, although it does not meet Shannon's definition of reliable communication since it does not guarantee which bits will be received. Each channel code has a matching source code: the multiresolution source code allows the broadcast channel code to transmit prioritized information, and the symmetric multiple description source code enables direct transmission of unprioritized information. The end-to-end expected distortions of these schemes are also compared. Yifan Liang, Andrea J. Goldsmith, Michelle Effros |
GLOBECOM | 1 |
| 2008 | Evolution of Base Stations in Cellular Networks: Denser Deployment versus CoordinationabstractIt has been demonstrated that base station cooperation can reduce co-channel interference (CCI) and increase cellular system capacity. In this work we consider another approach by dividing the system into microcells through denser base station deployment. We adopt the criterion to maximize the minimum spectral efficiency of served users with a certain user outage constraint. In a two-dimensional hexagon array with homogeneous microcell structure, under the proposed propagation model denser base station deployment outperforms suboptimal cooperation schemes (zero-forcing) when the density increases beyond 3 - 12 base stations per km2, the exact value depending on the rules of outage user selection. However, close- to-optimal cooperation schemes (zero-forcing with dirty-paper- coding) are always superior to denser deployment. Performance of a hierarchial cellular structure mixed with both macrocells and microcells is also evaluated. Yifan Liang, Andrea J. Goldsmith, Gerard J. Foschini, Reinaldo A. Valenzuela, Dmitry Chizhik |
ICC | 1 |
| 2007 | Adaptive Channel Reuse in Cellular SystemsabstractIn cellular systems a large reuse distance reduces co-channel interference while a small reuse distance increases bandwidth allocated to each cell. The optimal reuse distance is chosen to balance these two factors. Instead of applying a fixed reuse distance to the entire system, in this paper we study the effect of adaptive channel reuse based on channel strength. Assuming the traditional single base station transmission, adaptive channel reuse under different propagation models, with or without fading, is analyzed for the Wyner linear cellular model. A new approach where base stations collaborate in transmission is also considered. We observe that adjacent base cooperation does not show much advantage over traditional single base transmission under intra-cell orthogonal schemes and AWGN channel models. In order to fully exploit the benefit of base station cooperation, more sophisticated transmission schemes need to be investigated. Yifan Liang, Andrea J. Goldsmith |
ICC | 1 |
| 2007 | Capacity Definitions of General Channels with Receiver Side InformationabstractWe consider three capacity definitions for general channels with channel side information at the receiver, where the channel is modeled as a sequence of finite dimensional conditional distributions not necessarily stationary, ergodic, or information stable. The Shannon capacity is the highest rate asymptotically achievable with arbitrarily small error probability. The outage capacity is the highest rate asymptotically achievable with a given probability of decoder-recognized outage. The expected capacity is the highest expected rate asymptotically achievable with a single encoder and multiple decoders, where the channel side information determines the decoder in use. Expected capacity equals Shannon capacity for channels governed by a stationary ergodic random process but is typically greater for general channels. These alternative definitions essentially relax the constraint that all transmitted information must be decoded at the receiver. We derive equations for these capacity definitions through information density. Examples are also provided to demonstrate their implications. Michelle Effros, Andrea J. Goldsmith, Yifan Liang |
ISIT | 3 |
| 2006 | Symmetric Rate Capacity of Cellular Systems with Cooperative Base StationsabstractCooperation among base stations has demonstrated substantial capacity gain in cellular systems. The optimal transmission scheme under full base station cooperation requires all users to transmit simultaneously and a central joint receiver for multi-user detection. We consider some sub-optimal but more practical schemes of orthogonal channel access either within a cell (intra-cell TDMA) or among cells (inter-cell time sharing), which correspond to a partitioning of overall channel resources. The effects of various schemes on the uplink capacity of a cellular system are then compared for a modified Wyner model. Yifan Liang, Andrea J. Goldsmith |
GLOBECOM | 1 |
| 2006 | Coverage Spectral Efficiency of Cellular Systems with Cooperative Base StationsabstractCoverage spectral efficiency (CSE) characterizes the tradeoff between efficient channel reuse and the achievable rates per cell, under the assumption of detection by a single base station and intra-cell FDMA. It is well known that intra-cell FDMA is not in general optimal. In this paper we study an alternative intra- cell wide-band scheme as well as the base station cooperation in detection, which has demonstrated potential capacity gain. The effect on CSE of different schemes are then compared and the optimal reuse distance is determined for each scheme. Yifan Liang, Taesang Yoo, Andrea J. Goldsmith |
GLOBECOM | 1 |
| 2004 | Adaptive spatial spectrum estimation via conjugate gradients and dominant mode rejection beamformingabstractUnder the condition of low sample support, there is deviation between the estimated and true correlation matrices. In this case, the full-rank sample matrix inversion (SMI) beamformer is no longer optimal. Instead, a low-rank adaptive beamformer could yield better performance. When extending from a couple of signal directions to the entire spatial spectrum, the output power is a more applicable touchstone than SINR, which is only defined for signal directions. Adaptive algorithms (CG/DMR), if operating at the best rank for each direction, could effectively reveal the spatial spectrum. The best rank is obtained by investigating an ideal evaluation from the true correlation matrix and adaptive beamformer. Yifan Liang, Michael D. Zoltowski |
GLOBECOM | 1 |