Yun Han

dblp:51/3448 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
1since 2021 · last 2021
0000-0003-1410-8090ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 5 first-authorComputer networks · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visualization and visual analytics · 87% Multimedia analysis and retrieval · 13%

Topics — the 1 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › data visualization
animated visualization
0.412020
Automatic Annotation Synchronizing with Textual Description for Visualization · CHI 2020

Methods — techniques the papers use, named apart from their topics

natural language parsing · 0.4Mask R-CNN · 0.4
YearPublicationVenuePosition
2021 ADVISor: Automatic Visualization Answer for Natural-Language Question on Tabular Data
abstract
We propose an automatic pipeline to generate visualization with annotations to answer natural-language questions raised by the public on tabular data. With a pre-trained language representation model, the input natural language questions and table headers are first encoded into vectors. According to these vectors, a multi-task end-to-end deep neural network extracts related data areas and corresponding aggregation type. We present the result with carefully designed visualization and annotations for different attribute types and tasks. We conducted a comparison experiment with state-of-the-art works and the best commercial tools. The results show that our method outperforms those works with higher accuracy and more effective visualization.
Can Liu 0004, Yun Han, Ruike Jiang, Xiaoru Yuan
PacificVis2
2020 Interactive Assigning of Conference Sessions with Visualization and Topic Modeling
abstract
Creating thematic sessions based on accepted papers is important to the success of a conference. Facing a large number of papers from multiple topics, conference organizers need to identify the topics of papers and group them into sessions by considering the constraints on session numbers and paper numbers in individual sessions. In this paper, we present a system using visualization and topic modeling to help the construction of conference sessions. The system provides multiple automatically generated session schemes and allows users to create, evaluate, and manipulate paper sessions with given constraints. A case study based on our system on the VAST papers shows that our method can help users successfully construct coherent conference sessions. In addition to conference session management, our method can be extended to other tasks, such as event and class schedule.
Yun Han, Zhenhuang Wang, Siming Chen 0001, Guozheng Li 0002, Xiaolong Zhang 0001, Xiaoru Yuan
PacificVis1
2020 AutoCaption: An Approach to Generate Natural Language Description from Visualization Automatically
abstract
In this paper, we propose a novel approach to generate captions for visualization charts automatically. In the proposed method, visual marks and visual channels, together with the associated text information in the original charts, are first extracted and identified with a multilayer perceptron classifier. Meanwhile, data information can also be retrieved by parsing visual marks with extracted mapping relationships. Then a 1-D convolutional residual network is employed to analyze the relationship between visual elements, and recognize significant features of the visualization charts, with both data and visual information as input. In the final step, the full description of the visual charts can be generated through a template-based approach. The generated captions can effectively cover the main visual features of the visual charts and support major feature types in commons charts. We further demonstrate the effectiveness of our approach through several cases.
Can Liu 0004, Liwenhan Xie, Yun Han, Datong Wei, Xiaoru Yuan
PacificVis3
2020 Automatic Annotation Synchronizing with Textual Description for Visualization
abstract
In this paper, we propose a technique for automatically annotating visualizations according to the textual description. In our approach, visual elements in the target visualization, along with their visual properties, are identified and extracted with a Mask R-CNN model. Meanwhile, the description is parsed to generate visual search requests. Based on the identification results and search requests, each descriptive sentence is displayed beside the described focal areas as annotations. Different sentences are presented in various scenes of the generated animation to promote a vivid step-by-step presentation. With a user-customized style, the animation can guide the audience's attention via proper highlighting such as emphasizing specific features or isolating part of the data. We demonstrate the utility and usability of our method through a user study with use cases.
Chufan Lai, Zhixian Lin, Ruike Jiang, Yun Han, Can Liu 0004, Xiaoru Yuan
CHI4
2020 Overview of Network Outage Detection Technology in 5G
abstract
Understanding the network outage is beneficial to improve and enhance the security of the network. Detecting the outage can easily locate the event source and estimate the spread of the outage. Large-scale network outages will seriously affect the normal operation of network services and infrastructure. This study summarizes and compares the existing methods of network outage detection, and looks forward to future research on network interruptions, to provide a reference for the subsequent research on network security in 5G.
Yun Han, Yun Zhong
IWCMC1
2019 Learning Subgraph Structure with LSTM for Complex Network Link Prediction
Yun Han, Donghai Guan, Weiwei Yuan
ADMA1
2018 Highlighted Deep Learning based Identification of Pharmaceutical Blister Packages
abstract
Precise identification of blister packages carries utmost importance at dispensing stations, where numerous prescriptions are to be efficiently dispensed by pharmacists. However, the usual presence of several hundreds of similarly looking, but completely different types of blister packages, in a crowded dispensing station makes it prone to human error, posing serious safety and health concerns for a patients life. In this work, we propose a highlighted deep learning (HDL) based approach for accurate identification of blister packages. HDL allows smart manipulation on raw data in order to better segment the identified targets and to expose inherent descriptive features, thus facilitating an accurate deep learning based classification process. Specifically, HDL uses automatic detection and then segments and processes the raw blister pack images irrespective of position, lighting variations, etc., making them suitable for CNN to classify the correct blister pack types. A ResNet CNN classifier has been trained using the processed images, and the resultant CNN model is finally deployed for classification. We have conducted an extensive experiment at the adult lozenge dispensing station of MacKay Memorial Hospital. The database that will be released along with this paper, consists of 272 types of blister packages, with 65 images belonging to each type for training and additional 7 images for validation. On real testing scenario, the proposed approach yielded almost 100% accuracy, consistently over the entire testing set. The proposed solution can also be extended to any dispensing station for automatic detection and classification of blister packages, thereby offering an highly effective and efficient delivery system.
Jing-Syuan Wang, Arul-Murugan Ambikapathi, Yun Han, Sheng-Luen Chung, Hsien-Wei Ting, Chih-Fang Chen
ETFA3
2018 Robust Human Action Recognition Using Global Spatial-Temporal Attention for Human Skeleton Data
abstract
Human action recognition from video sequences is one of the most challenging computer vision applications, primarily owing to intrinsic variations in lighting, pose, occlusions, and other factors. The human skeleton joints extracted by the depth camera Kinect have the advantages of simplified structures and rich contents, and are therefore widely used for capturing human actions. However, at present, most of the skeletal joint and Deep learning based action recognition methods treat all skeletal joints equally in both spatial and temporal dimensions. Logically, this is not in accordance with the fact that for different human actions the contributions from skeletal joints could significantly vary spatially and temporally. Incorporating information pertaining to such natural variations will certainly aid in designing a robust human action recognitions system. Hence, in this work, we endeavor to propose a global spatial attention (GSA) model to suitably express the different skeletal joints with different weights so as to provide precise spatial information for human action recognition. Further, we will introduce the notion of accumulative learning curve (ALC) model that can highlight which frames contribute most to the final decision by giving varying temporal weights to each intermediate accumulated learning results provided by an LSTM upon input frames. The proposed GSA (for spatial information) and ALC (for temporal processing) models are integrated into the LSTM framework to construct a robust action recognition framework that takes the human skeletal joints as input and predicts the human action using the enhanced spatial-temporal attention model. Rigorous experiments on NTU datasets (by-far the largest benchmark RGB-D dataset) show that the proposed framework offers the best performance accuracy, least algorithmic complexity and training overheads, when compared with other state-of-the-art human action recognition models.
Yun Han, Sheng-Luen Chung, Arul-Murugan Ambikapathi, Jui-Shan Chan, Wei-You Lin, Shun-Feng Su
IJCNN1
2018 Two-Stream LSTM for Action Recognition with RGB-D-Based Hand-Crafted Features and Feature Combination
abstract
Good action recognition relies on correct interpretation of two critical attributes related to action: the spatial attribute on the detected person's posture, and the temporal attribute on the detected person's body movement. Whereas deep learning has greatly improved image recognition, we have not found a similar progress for action recognition. One of the main reasons is due to the complexity caused by the additional temporal dimension; another, to the fact that there are less annotated training data samples for action recognition than that for image recognition. In this regard, this paper proposes a handcrafted cued LSTM model for human action recognition based on RGB-D data, as a collection of 25 skeleton joints in 3D coordinates, found in NTU-RGB-D, currently the most comprehensive dataset for action recognition. As opposed to the raw data of skeleton joints, handcrafted cues, pre-processed results geared to facilitate focused learning, are proposed as input to the LSTM structure. In particular, pertaining to the spatial cue, the SVIT cue derived by Skeleton View-invariant Transformation is adopted; pertaining to the temporal cue, the Diff cue computed by taking the displacements of all joint across down-sampled raw data is utilized. Based on the train/test protocol, the experiment we conducted on NTU-RGB-D shows that the recognition result based on either of the proposed handcrafted cues is better than that based on the raw data. In addition, by our proposed techniques of feature fusion and/or decision fusion of these two handcrafted cues, the recognition performance is better than that of the state-of-the-art approaches conducting on the same dataset by same train/test protocol.
Yun Han, Sheng-Luen Chung, Sheng-Fang Chen, Shun-Feng Su
SMC1
2017 Automatic action segmentation and continuous recognition for basic indoor actions based on kinect pose streams
abstract
One main difficulty in applying action recognition to practical applications is the need to segment beginnings and ends of actions in a continuous online monitoring process. This paper proposed a finite state machine (FSM) model for automatic action segmentation and recognition solution, based on pose streams in the form of skeleton joint data provided by Kinect. With the action recognition problem reframed as a state identification problem, the key solution to state identification hinges on detection of changing events, which signify the start of new action and the recognition of the underlying action. In that regard, a decision tree is constructed to detect these events based on the spatial positions and the changes of the skeleton data. In addition, to identify the current state or equivalently the action of a detected person, a fault observer is derived from the modeled action FSM. The fault observer does not only identify the initial state of action when the recognition process starts, but also serve the purpose of error recovery when the system loses track of ongoing events at times of intermittent sensor faults. Additionally, an AutoCorrect mechanism is presented to further enhance the accuracy of action recognition. To evaluate the proposed approach, an experiment with 300 participating subjects has been conducted for a total of 900 test sequences. The 98.63% correct identification result ensures the proposed approach a promising solution to constant action monitoring solution.
Yun Han, Sheng-Luen Chung, Shun-Feng Su
SMC1
2015 Activity Recognition Based on Relative Positional Relationship of Human Joints
abstract
Kinect have been used as a revolutionary sensor for recent human activity recognition research, mainly due to its ready skeletal joint information that facilitates activity analysis. However, the sensor's unstable and imprecise measurement may impair the analysis results. To alleviate this impact due to the unavoidable noisy measurement, this paper presents a recognition approach based on the relative positional relationship among measured joints data. Relative positional relationship between two joints refers to the relation either one joint's position in the x, y, z coordinates is to the left or right, above or below, in front or behind the other. Capturing these relative positional relationship among all measured joints data as a {-1, 0, 1} sequence throughout the time span of an activity, or more succinctly represented by a summary of the change frequency of these relationship, proves to be an effective feature for comparison and classification. In doing so, measured joint data are first pre-processed by SVIT (skeleton-based viewpoint invariant transformation) to eliminate viewpoint variation. Then, as the first feature, change frequencies of relative positional relationship among joints are constructed. As a complimentary feature, a HOG3D-like feature that summarizes the joints spatial distribution throughout an activity is also constructed. Classification credibility of feature is then applied to fuse the two features with ELM as machine learning for activity classification. The proposed approach with its efficient implementation has been applied to MSRDailyActivity3D and demonstrated favorable performance than results from literature.
Yun Han, Sheng-Luen Chung
SMC1
2014 SAR target segmentation based on shape prior
abstract
It's very challenging to interpret the SAR image due to the speckle noise and target's heterogeneity, etc. And the use of image information alone often leads to poor segmentation. However, the introduction of shape prior into the active contours has proved to be very effective to segment the target. In this study, we define the shape of SAR target and introduce it into our active contours. To avoid the irregularity of level set evolution, a novel potential function is introduced to regularize the level set function. And the energy of gray intensity helps to locate the target while the energy of target's shape regularizes the target's contour. Experiments on both ship and building segmentation prove the validity of the proposal algorithm.
Yun Han, Wenxian Yu
IGARSS1
2013 Localization of RGB-D Camera Networks by Skeleton-Based Viewpoint Invariance Transformation
abstract
Combining depth information and color image, RGB-D cameras provide ready detection of humans and the associated 3D skeleton joints data, facilitating if not revolutionizing conventional image centric researches in, among others, computer vision, surveillance, and human activity analysis. Applicability of a D-RBG camera, however, is restricted by its limited range of frustum of depth in the range of 0.8 to 4 meters. Although a RGB-D camera network, constructed by deployment of several RGB-D cameras at various locations, could extend the range of coverage, it requires precise localization of the camera network: relative location and orientation of neighboring cameras. By introducing a skeleton-based viewpoint invariant transformation (SVIT) that derives relative location and orientation of a detected human's upper torso to a RGB-D camera, this paper presents a reliable automatic localization technique without the need for additional instrument or human intervention. By respectively applying SVIT to two neighboring RGB-D cameras on a commonly observed skeleton, the respective relative position and orientation of the detected human's skeleton to these two respective cameras can be obtained before being combined to yield the relative position and orientation of these two cameras, thus solving the localization problem. Experiments have been conducted where two Kinects are situated with bearing differences of about 45 degrees and 90 degrees when coverage extended by up to 70% with the installment of an additional Kinect. The same localization technique can be applied repeatedly to a larger number of RGB-D cameras, thus extending the applicability of RGB-D cameras to camera networks in making human behavior analysis and context-aware service in a lager surveillance area.
Yun Han, Sheng-Luen Chung, Jeng-Sheng Yeh
SMC1
2009 APS-FeW: Improving TCP throughput over multihop adhoc networks
Xinbing Wang, Yun Han, Youyun Xu
Comput. Commun.2
2008 Performance Improvement of Voice over Multihop 802.11 Networks
abstract
Extensive studies on supporting voice traffic over wireless 802.11 networks have been carried out in the literature. Most of them were focused only on one hop infrastructure mode. This paper addresses the issue of voice over multihop wireless networks. Through simulation considering a simple topology, we identify the voice traffic bottleneck in 802.11 MAC and thus propose a burst queue (BQ) scheme to cooperate with 802.11b and 802.11e MAC protocols. Extensive simulation results under variety of circumstances show that BQ scheme can support 50%~100% more voice calls and obtain 20%~50% lower delay at the cost of a little higher loss ratio which is tolerable to voice flows. Even for the grid and random topology, our BQ scheme can provide smaller packet loss ratio, and shorter end-to-end delay.
Chenhui Hu, Youyun Xu, Wen Chen 0001, Xinbing Wang, Yun Han, Hsiao-Hwa Chen
GLOBECOM5
2007 APS-FeW: The Second Order Enhancement for TCP over Multihop 802.11 Networks
abstract
TCP is a reliable transport protocol tuned to perform well over traditional wired networks. Although it performs well for wired networks, TCP's implicit assumption that any packet loss is due to congestion is not valid any longer in mobile ad hoc networks. It is observed that TCP induces the over-action of routing protocol and reduces the performance of the connection. Fraction window increment (FeW) scheme for TCP improves the connection performance by limiting TCP's aggressiveness. But to some extent, this limitation is too strict in that it eliminates the possibility to deliver more bytes under the same congestion window. To solve this problem, we propose an adaptive packet size (APS) scheme to work on top of FeW for TCP. The proposed scheme utilizes the advantages of both legacy TCP and FeW to achieve high performance over multihop 802.11 networks. Extensive simulation results demonstrate that APS over FeW outperforms FeW alone by 10 -25 % according to different scenarios, e.g., chain-topology, grid-topology, and random-topology with mobility.
Yun Han, Xinbing Wang, Youyun Xu, Ruhai Wang
GLOBECOM1