VLDB 2026 Research / reviewers in the wild / expert
Shaoliang Nie
dblp:213/7860
· DBLP profile ↗
14ranked-venue papers
2as first author
11since 2021 · last 2024
0000-0002-5513-3439ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 31% Information extraction and text analysis · 18% Language models and text generation · 17% | |
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 100% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.2 | 2 | 2023 | Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-text Rationales · ACL (1) 2023 UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Computer vision › Vision and language
multimodal prompt learning |
0.8 | 1 | 2024 | M²PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning · EMNLP 2024 |
Algorithmic game theory and mechanism design
mechanism design |
0.8 | 1 | 2024 | Unveiling User Satisfaction and Creator Productivity Trade-Offs in Recommendation Platforms · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › transformer
multimodal transformer |
0.7 | 1 | 2023 | MUSTIE: Multimodal Structural Transformer for Web Information Extraction · ACL (1) 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.7 | 1 | 2023 | APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models · EMNLP 2023 |
Natural language and speech › Language models and text generation
prompt tuning |
0.7 | 1 | 2023 | APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.7 | 1 | 2023 | MUSTIE: Multimodal Structural Transformer for Web Information Extraction · ACL (1) 2023 |
Recommender systems
explainable recommendation |
0.7 | 1 | 2023 | COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable Recommendation · EMNLP 2023 |
Machine learning › Deep learning architectures and training › regularization
dropout |
0.6 | 1 | 2022 | AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-Tuning · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation |
0.6 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.6 | 1 | 2022 | AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-Tuning · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › interpretability › rationalization
rationale extraction |
0.6 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.2 | 1 | 2023 | COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable Recommendation · EMNLP 2023 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning |
0.2 | 1 | 2023 | APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
text classification |
0.2 | 1 | 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale Extraction · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
optimization · 1.5guided signal learning · 1.3counterfactual fairness · 1.3prompt tuning · 0.8structural modeling · 0.7multimodal transformer · 0.7attention prompt tuning · 0.7select-predict pipeline · 0.6joint training · 0.6cross-tuning · 0.6attribution algorithm · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningabstractTaowen Wang, Yiyang Liu, James Chenhao Liang, Junhan Zhao, Yiming Cui, Yuning Mao, Shaoliang Nie, Jiahao Liu, Fuli Feng, Zenglin Xu, Cheng Han, Lifu Huang, Qifan Wang, Dongfang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Taowen Wang, Yiyang Liu 0003, James Liang, Junhan Zhao, Yiming Cui 0002, Yuning Mao, Shaoliang Nie, Fuli Feng, Zenglin Xu, Cheng Han 0001, Lifu Huang, Qifan Wang 0001, Dongfang Liu |
EMNLP | 7 |
| 2024 | Unveiling User Satisfaction and Creator Productivity Trade-Offs in Recommendation PlatformsabstractOn User-Generated Content (UGC) platforms, recommendation algorithms significantly impact creators' motivation to produce content as they compete for algorithmically allocated user traffic. This phenomenon subtly shapes the volume and diversity of the content pool, which is crucial for the platform's sustainability. In this work, we demonstrate, both theoretically and empirically, that a purely relevance-driven policy with low exploration strength boosts short-term user satisfaction but undermines the long-term richness of the content pool. In contrast, a more aggressive exploration policy may slightly compromise user satisfaction but promote higher content creation volume. Our findings reveal a fundamental trade-off between immediate user satisfaction and overall content production on UGC platforms. Building on this finding, we propose an efficient optimization method to identify the optimal exploration strength, balancing user and creator engagement. Our model can serve as a pre-deployment audit tool for recommendation algorithms on UGC platforms, helping to align their immediate objectives with sustainable, long-term goals. Fan Yao 0002, Yiming Liao, Jingzhou Liu, Shaoliang Nie, Qifan Wang 0001, Hongning Wang |
NeurIPS | 4 |
| 2024 | AK-GPSR: An Adaptive K-Medoids-Based Greedy Perimeter Stateless Routing Algorithm for Multi-Channel Vehicular Network CommunicationabstractAs a direct application of 5G communications and computer technology, Vehicular Ad-Hoc Networks (VANETs) are already having a profound impact on all sectors of society. However, frequent changes in the topology of VANETs have resulted in poor vehicle communication quality, highly susceptible to communication link breaks and data transmission reliability decreases, and the cost of vehicular communication increases continuously. In this paper, an Adaptive K-medoids based on Greedy Perimeter Stateless Routing (AK-GPSR) algorithm is proposed in multi-channel vehicular network communication of urban scenario. It is an unsupervised learning algorithm, aiming to form high-quality link communication, more stable network topology, and improved data information transmission reliability. First, the proposed AK-GPSR algorithm applies Gap statistic to evaluate the K-medoids algorithm and select the best K value. Further, the K-medoids algorithm clusters the vehicles in the simulation area by the optimal K-value to divide the K clusters. Finally, the packets are forwarded from the source vehicle to the destination vehicle or Road Side Unit (RSU) using the forwarding method of the Greedy Perimeter Stateless Routing (GPSR) algorithm. The experimental results show that our proposed AK-GPSR algorithm has good performance and applicability in multi-channel vehicular network communication. Wanneng Shu, Shaoliang Nie, Fengjun Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-text RationalesabstractBrihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Brihi Joshi, Ziyi Liu 0007, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang 0001, Yejin Choi 0001, Xiang Ren 0001 |
ACL (1) | 6 |
| 2023 | MUSTIE: Multimodal Structural Transformer for Web Information ExtractionabstractQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qifan Wang 0001, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu |
ACL (1) | 6 |
| 2023 | Generating Hashtags for Short-form Videos with Guided SignalsabstractTiezheng Yu, Hanchao Yu, Davis Liang, Yuning Mao, Shaoliang Nie, Po-Yao Huang, Madian Khabsa, Pascale Fung, Yi-Chia Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tiezheng Yu, Hanchao Yu, Davis Liang, Yuning Mao, Shaoliang Nie, Po-Yao Huang 0001, Madian Khabsa, Pascale Fung, Yi-Chia Wang |
ACL (1) | 5 |
| 2023 | APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsabstractQifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qifan Wang 0001, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu |
EMNLP | 5 |
| 2023 | COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable RecommendationabstractNan Wang, Qifan Wang, Yi-Chia Wang, Maziar Sanjabi, Jingzhou Liu, Hamed Firooz, Hongning Wang, Shaoliang Nie. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qifan Wang 0001, Yi-Chia Wang, Maziar Sanjabi, Jingzhou Liu, Hamed Firooz, Hongning Wang, Shaoliang Nie |
EMNLP | 8 |
| 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionabstractAn extractive rationale explains a language model’s (LM’s) prediction on a given task instance by highlighting the text inputs that most influenced the prediction. Ideally, rationale extraction should be faithful (reflective of LM’s actual behavior) and plausible (convincing to humans), without compromising the LM’s (i.e., task model’s) task performance. Although attribution algorithms and select-predict pipelines are commonly used in rationale extraction, they both rely on certain heuristics that hinder them from satisfying all three desiderata. In light of this, we propose UNIREX, a flexible learning framework which generalizes rationale extractor optimization as follows: (1) specify architecture for a learned rationale extractor; (2) select explainability objectives (\ie faithfulness and plausibility criteria); and (3) jointly train the task model and rationale extractor on the task using selected objectives. UNIREX enables replacing prior works’ heuristic design choices with a generic learned rationale extractor in (1) and optimizing it for all three desiderata in (2)-(3). To facilitate comparison between methods w.r.t. multiple desiderata, we introduce the Normalized Relative Gain (NRG) metric. On five English text classification datasets, our best UNIREX configuration outperforms baselines by an average of 32.9% NRG. Plus, UNIREX rationale extractors’ faithfulness can even generalize to unseen datasets and tasks. Aaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan 0005, Shaoliang Nie, Xiaochang Peng, Xiang Ren 0001, Hamed Firooz |
ICML | 5 |
| 2022 | AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningabstractFine-tuning large pre-trained language models on downstream tasks is apt to suffer from overfitting when limited training data is available. While dropout proves to be an effective antidote by randomly dropping a proportion of units, existing research has not examined its effect on the self-attention mechanism. In this paper, we investigate this problem through self-attention attribution and find that dropping attention positions with low attribution scores can accelerate training and increase the risk of overfitting. Motivated by this observation, we propose Attribution-Driven Dropout (AD-DROP), which randomly discards some high-attribution positions to encourage the model to make predictions by relying more on low-attribution positions to reduce overfitting. We also develop a cross-tuning strategy to alternate fine-tuning and AD-DROP to avoid dropping high-attribution positions excessively. Extensive experiments on various benchmarks show that AD-DROP yields consistent improvements over baselines. Analysis further confirms that AD-DROP serves as a strategic regularizer to prevent overfitting during fine-tuning. Tao Yang 0033, Jinghao Deng, Xiaojun Quan, Qifan Wang 0001, Shaoliang Nie |
NeurIPS | 5 |
| 2021 | Visual Analytics of Text Conversation Sentiment and SemanticsabstractAbstract This paper describes the design and implementation of a web‐based system to visualize large collections of text conversations integrated into a hierarchical four‐level‐of‐detail design. Viewers can visualize conversations: (1) in a streamgraph topic overview for a user‐specified time period; (2) as emotion patterns for a topic chosen from the streamgraph; (3) as semantic sequences for a user‐selected emotion pattern, and (4) as an emotion‐driven conversation graph for a single conversation. We collaborated with the Live Chatcustomer service group at SAS Institute to design and evaluate our system's strengths and limitations. Christopher G. Healey, Gowtham Dinakaran, Kalpesh Padia, Shaoliang Nie, J. Riley Benson, Dave Caira, Dean Shaw, Gary Catalfu, Ravi Devarajan |
Comput. Graph. Forum | 4 |
| 2019 | Feature Selection for Facebook Feed Ranking System via a Group-Sparsity-Regularized Training AlgorithmabstractIn modern production platforms, large scale online learning models are applied to data of very high dimension. To save computational resource, it is important to have an efficient algorithm to select the most significant features from an enormous feature pool. In this paper, we propose a novel neural-network-suitable feature selection algorithm, which selects important features from the input layer during training. Instead of directly regularizing the training loss, we inject group-sparsity regularization into the (stochastic) training algorithm. In particular, we introduce a group sparsity norm into the proximally regularized stochastical gradient descent algorithm. To fully evaluate the practical performance, we apply our method to Facebook News Feed dataset, and achieve favorable performance compared with state-of-the-arts using traditional regularizers. Xiuyan Ni, Peng Wu 0017, Youlin Li, Shaoliang Nie, Qichao Que, Chao Chen 0012 |
CIKM | 5 |
| 2019 | Rapid Sequence Matching for Visualization Recommender Systems
Shaoliang Nie, Christopher G. Healey, Rada Chirkova, Juan L. Reutter |
Graphics Interface | 1 |
| 2018 | Visualizing Deep Neural Networks for Text AnalyticsabstractDeep neural networks (DNNs) have made tremendous progress in many different areas in recent years. How these networks function internally, however, is often not well understood. Advances in under-standing DNNs will benefit and accelerate the development of the field. We present TNNVis, a visualization system that supports un-derstanding of deep neural networks specifically designed to analyze text. TNNVis focuses on DNNs composed of fully connected and convolutional layers. It integrates visual encodings and interaction techniques chosen specifically for our tasks. The tool allows users to: (1) visually explore DNN models with arbitrary input using a combination of node-link diagrams and matrix representation; (2) quickly identify activation values, weights, and feature map patterns within a network; (3) flexibly focus on visual information of interest with threshold, inspection, insight query, and tooltip operations; (4) discover network activation and training patterns through animation; and (5) compare differences between internal activation patterns for different inputs to the DNN. These functions allow neural network researchers to examine their DNN models from new perspectives, producing insights on how these models function. Clustering and summarization techniques are employed to support large convolutional and fully connected layers. Based on several part of speech models with different structure and size, we present multiple use cases where visualization facilitates an understanding of the models. Shaoliang Nie, Christopher G. Healey, Kalpesh Padia, Samuel P. Leeman-Munk, Jordan Benson, Dave Caira, Saratendu Sethi, Ravi Devarajan |
PacificVis | 1 |