Ruolin Su

dblp:209/4123 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Raster-to-Graph: Floorplan Recognition via Autoregressive Graph Prediction with an Attention Transformer
abstract
Abstract Recognizing the detailed information embedded in rasterized floorplans is at the research forefront in the community of computer graphics and vision. With the advent of deep neural networks, automatic floorplan recognition has made tremendous breakthroughs. However, co‐recognizing both the structures and semantics of floorplans through one neural network remains a significant challenge. In this paper, we introduce a novel framework Raster‐to‐Graph, which automatically achieves structural and semantic recognition of floorplans. We represent vectorized floorplans as structural graphs embedded with floorplan semantics, thus transforming the floorplan recognition task into a structural graph prediction problem. We design an autoregressive prediction framework using the neural network architecture of the visual attention Transformer, iteratively predicting the wall junctions and wall segments of floorplans in the order of graph traversal. Additionally, we propose a large‐scale floorplan dataset containing over 10,000 real‐world residential floorplans. Our autoregressive framework can automatically recognize the structures and semantics of floorplans. Extensive experiments demonstrate the effectiveness of our framework, showing significant improvements on all metrics. Qualitative and quantitative evaluations indicate that our framework outperforms existing state‐of‐the‐art methods. Code and dataset for this paper are available at: https://github.com/HSZVIS/Raster-to-Graph .
Sizhe Hu, Wenming Wu 0001, Ruolin Su, Wanni Hou, Liping Zheng, Benzhu Xu
Comput. Graph. Forum3
2023 Clicker: Attention-Based Cross-Lingual Commonsense Knowledge Transfer
abstract
Recent advances in cross-lingual commonsense reasoning (CSR) are facilitated by the development of multilingual pre-trained models (mPTMs). While mPTMs show the potential to encode commonsense knowledge for different languages, transferring commonsense knowledge learned in large-scale English corpus to other languages is challenging. To address this problem, we propose the attention-based Cross-LIngual Commonsense Knowledge transfER (CLICKER) framework, which minimizes the performance gaps between English and non-English languages in commonsense question-answering tasks. CLICKER effectively improves commonsense reasoning for non-English languages by differentiating non-commonsense knowledge from commonsense knowledge. Experimental results on public benchmarks demonstrate that CLICKER achieves remarkable improvements in the cross-lingual CSR task for languages other than English.
Ruolin Su, Zhongkai Sun, Sixing Lu, Chengyuan Ma, Chenlei Guo
ICASSP1
2023 Choice Fusion As Knowledge For Zero-Shot Dialogue State Tracking
abstract
With the demanding need for deploying dialogue systems in new domains with less cost, zero-shot dialogue state tracking (DST), which tracks user’s requirements in task-oriented dialogues without training on desired domains, draws attention increasingly. Although prior works have leveraged question-answering (QA) data to reduce the need for in-domain training in DST, they fail to explicitly model knowledge transfer and fusion for tracking dialogue states. To address this issue, we propose CoFunDST, which is trained on domain-agnostic QA datasets and directly uses candidate choices of slot-values as knowledge for zero-shot dialogue-state generation, based on a T5 pre-trained language model. Specifically, CoFunDST selects highly-relevant choices to the reference context and fuses them to initialize the decoder to constrain the model outputs. Our experimental results show that our proposed model achieves outperformed joint goal accuracy compared to existing zero-shot DST approaches in most domains on the MultiWOZ 2.1. Extensive analyses demonstrate the effectiveness of our proposed approach for improving zero-shot DST learning from QA.
Ruolin Su, Jingfeng Yang 0001, Ting-Wei Wu, Biing-Hwang Juang
ICASSP1
2022 Why Patient Data Cannot Be Easily Forgotten?
Ruolin Su, Xiao Liu 0037, Sotirios A. Tsaftaris
MICCAI (8)1
2021 A Label-Aware BERT Attention Network for Zero-Shot Multi-Intent Detection in Spoken Language Understanding
abstract
With the early success of query-answer assistants such as Alexa and Siri, research attempts to expand system capabilities of handling service automation are now abundant.However, preliminary systems have quickly found the inadequacy in relying on simple classification techniques to effectively accomplish the automation task.The main challenge is that the dialogue often involves complexity in user's intents (or purposes) which are multiproned, subject to spontaneous change, and difficult to track.Furthermore, public datasets have not considered these complications and the general semantic annotations are lacking which may result in zero-shot problem.Motivated by the above, we propose a Label-Aware BERT Attention Network (LABAN) for zeroshot multi-intent detection.We first encode input utterances with BERT and construct a label embedded space by considering embedded semantics in intent labels.An input utterance is then classified based on its projection weights on each intent embedding in this embedded space.We show that it successfully extends to few/zero-shot setting where part of intent labels are unseen in training data, by also taking account of semantics in these unseen intent labels.Experimental results show that our approach is capable of detecting many unseen intent labels correctly.It also achieves the state-of-the-art performance on five multiintent datasets in normal cases.
Ting-Wei Wu, Ruolin Su, Biing-Hwang Juang
EMNLP (1)2
2021 Act-Aware Slot-Value Predicting in Multi-Domain Dialogue State Tracking
abstract
As an essential component in task-oriented dialogue systems, dialogue state tracking (DST) aims to track human-machine interactions and generate state representations for managing the dialogue. Representations of dialogue states are dependent on the domain ontology and the user's goals. In several task-oriented dialogues with a limited scope of objectives, dialogue states can be represented as a set of slot-value pairs. As the capabilities of dialogue systems expand to support increasing naturalness in communication, incorporating dialogue act processing into dialogue model design becomes essential. The lack of such consideration limits the scalability of dialogue state tracking models for dialogues having specific objectives and ontology. To address this issue, we formulate and incorporate dialogue acts, and leverage recent advances in machine reading comprehension to predict both categorical and non-categorical types of slots for multi-domain dialogue state tracking. Experimental results show that our models can improve the overall accuracy of dialogue state tracking on the MultiWOZ 2.1 dataset, and demonstrate that incorporating dialogue acts can guide dialogue state design for future task-oriented dialogue systems.
Ruolin Su, Ting-Wei Wu, Biing-Hwang Juang
Interspeech1
2021 A Context-Aware Hierarchical BERT Fusion Network for Multi-Turn Dialog Act Detection
abstract
The success of interactive dialog systems is usually associated with the quality of the spoken language understanding (SLU) task, which mainly identifies the corresponding dialog acts and slot values in each turn.By treating utterances in isolation, most SLU systems often overlook the semantic context in which a dialog act is expected.The act dependency between turns is nontrivial and yet critical to the identification of the correct semantic representations.Previous works with limited context awareness have exposed the inadequacy of dealing with complexity in multiproned user intents, which are subject to spontaneous change during turn transitions.In this work, we propose to enhance SLU in multi-turn dialogs, employing a context-aware hierarchical BERT fusion Network (CaBERT-SLU) to not only discern context information within a dialog but also jointly identify multiple dialog acts and slots in each utterance.Experimental results show that our approach reaches new state-of-the-art (SOTA) performances in two complicated multi-turn dialogue datasets with considerable improvements compared with previous methods, which only consider single utterances for multiple intents and slot filling.
Ting-Wei Wu, Ruolin Su, Biing-Hwang Juang
Interspeech2
2017 Energy-Efficient Scheduling and Power Allocation for Energy Harvesting-Based D2D Communication
abstract
Energy Harvesting (EH)-based Device-to-Device (D2D) communication brings some challenges in resources management due to the joint influence of the volatility of available energy and the interference between cellular and D2D users. In this paper, we focus on improving the energy efficiency of EH-based D2D communication for the scenario where multiple EH- based D2D communication links multiplex the uplink channel resource of one cellular user (CU). Considering the variation of transmission requests based on available energy in different time slots, a short-term sum energy efficiency maximization problem for EH-based D2D communication is formulated to integrate the transmission scheduling and power allocation while maintaining a given transmission rate requirement for both CU and D2D links. The modeled problem is a non-convex mixed integer non- linear programming (MINLP) problem. In view of the NP-hardness property of the optimization problem, we develop a two-layer convex approximation iteration algorithm (CAIA) to obtain a feasible suboptimal solution. Finally, numerical simulation results indicate the performance of CAIA in aspects of average energy efficiency and transmission rate of D2D communication.
Ying Luo 0002, Peilin Hong, Ruolin Su
GLOBECOM3