VLDB 2026 Research / reviewers in the wild / expert
Yiqing Cao
dblp:59/2468
· DBLP profile ↗
4ranked-venue papers
2as first author
1since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 61% Reinforcement learning · 30% Question answering and dialogue systems · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict |
0.8 | 1 | 2024 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024 |
Machine learning › Reinforcement learning
preference learning |
0.8 | 1 | 2024 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model training
pretraining data quality |
0.8 | 1 | 2024 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks |
0.2 | 1 | 2024 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
statistical co-occurrence analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting KnowledgeabstractHaving been trained on massive pretraining data, large language models have shown excellent performance on many knowledge-intensive tasks.However, pretraining data tends to contain misleading and even conflicting information, and it is intriguing to understand how LLMs handle these noisy data during training.In this study, we systematically analyze LLMs' learning preferences for data with conflicting knowledge.We find that pretrained LLMs establish learning preferences similar to humans, i.e., preferences towards formal texts and texts with fewer spelling errors, resulting in faster learning and more favorable treatment of knowledge in data with such features when facing conflicts.This finding is generalizable across models and languages and is more evident in larger models.An in-depth analysis reveals that LLMs tend to trust data with features that signify consistency with the majority of data, and it is possible to instill new preferences and erase old ones by manipulating the degree of consistency with the majority data. Jiahuan Li, Yiqing Cao, Shujian Huang, Jiajun Chen 0001 |
EMNLP | 2 |
| 2017 | Resource Spread Multiple Access - A Novel Transmission Scheme for 5G UplinkabstractAmong the three 5G scenarios identified by ITU-R, massive MTC poses very challenging uplink design target which requires the simultaneous support of enhanced coverage, long battery life and massive connection density. During the 5G Release 14 study phase in 3GPP, the working group studied the Non-Orthogonal Multiple Access (NOMA) scheme which can improve link budget and increase user density with the cost of complex user detection scheme. Among all the candidates, Resource Spread Multiple Access (RSMA) scheme shows promising performance gain without introducing too much complexity on the receiver design. In this paper, we briefly review the RSMA scheme. For the link level performance analysis, we firstly analyze and compare the link budget among RSMA, TDMA, and FDMA to find where the gain for RSMA comes from. Meanwhile, we provide preliminary results to justify the gains compared with OMA scheme. On system design, we review the whole procedure and propose to use user centric mobility and grant free to improve the system efficiency along with NOMA. Yiqing Cao, Haitong Sun, Joseph B. Soriaga, Tingfang Ji |
VTC Fall | 1 |
| 2007 | Degree Distribution Based MIMO - LDPC SchemeabstractIn this paper, an LDPC based MIMO scheme is proposed in which matrix modulation scheme of MIMO and unequal protection ability of irregular LDPC codes are employed. The new scheme named degree distribution based-MIMO (DDB-MIMO for short) combines the two schemes, in which the important bits with large degrees are transmitted in transmission diversity (TD) symbols and bits with small degrees are mapped to the spatial multiplexing (SM) mode symbols. Since the TD mode could offer more protection to the data streams by offering more transmission power, the important bits with large degrees which could contribute more to the decoding process would have more protection from TD mode and hence the system performance would be improved namely. The performance of the new scheme has also been evaluated by the Gaussian approximation (GA) scheme which is used to analyze the asymptotic performance of LDPC coded system. The analysis results provide us the optimal mapping fractions between the bits with different degrees and transmitted symbols in different modes which is almost the same as the DDB-MIMO scheme offering. The simulation results also proved the new scheme. Yiqing Cao, Xuehua Li, Zhensong Li, Dacheng Yang |
VTC Fall | 1 |
| 2007 | Improved sRB-HARQ for OFDM SystemabstractIn the sub-channel reliability based HARQ (sRB-HARQ) scheme the channel qualities are exploited in retransmission to enhance the system throughput. The sub-channels with poor channel qualities measured by signal to noise ratio (SNR) are considered to be more unreliable, and the data on such sub-carriers are prior to be retransmitted when necessary. In this paper, an improved sRB-HARQ scheme for OFDM system is proposed, and the retransmission amount of sub-channels is determined adaptively by the target reliability of the received information which is related with the error ratio of received data (e. g. BLER in this paper). The target reliability for each (re)transmission could be same or different which are categorized as single target and multi target respectively. The sub-channels with least reliabilities are retransmitted on better sub-channels in order to achieve higher diversity gain. Thus the system throughput could be improved significantly and high power efficiency is achieved which are also proved by simulation results. Yiqing Cao, Liangang Chi, Dacheng Yang |
VTC Spring | 2 |