Jianping Fan 0002

dblp:69/2360-2 · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-7389-9112ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 since 2021Artificial intelligence and machine learning · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Computer networks · 5Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 61% Electronic design automation · 26% Performance modeling and evaluation · 13%
Artificial intelligence
1 paper
3D vision · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Computer networks
1 paper
Wireless networking · 77% Internet of things and sensor networks · 23%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
human mesh recovery
0.912025
A Twist Representation and Shape Refinement Method for Human Mesh Recovery · IEEE Trans. Multim. 2025
Computer vision › 3D vision › 3d shape modeling
shape refinement
0.912025
A Twist Representation and Shape Refinement Method for Human Mesh Recovery · IEEE Trans. Multim. 2025
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.512021
GML: Efficiently Auto-Tuning Flink's Configurations Via Guided Machine Learning · IEEE Trans. Parallel Distributed Syst. 2021
Cloud and datacenter computing
configuration tuning
0.512021
GML: Efficiently Auto-Tuning Flink's Configurations Via Guided Machine Learning · IEEE Trans. Parallel Distributed Syst. 2021
Electronic design automation › machine learning for EDA
machine learning-based tuning
0.512021
GML: Efficiently Auto-Tuning Flink's Configurations Via Guided Machine Learning · IEEE Trans. Parallel Distributed Syst. 2021
Cloud and datacenter computing › big data platform
big data frameworks
0.112021
GML: Efficiently Auto-Tuning Flink's Configurations Via Guided Machine Learning · IEEE Trans. Parallel Distributed Syst. 2021
Wireless networking
cognitive radio
0.112012
Connectivity of large-scale Cognitive Radio Ad Hoc Networks · INFOCOM 2012
Data mining
clustering
0.112010
Towards mobility-based clustering · KDD 2010
Data mining › clustering
density-based clustering
0.112010
Towards mobility-based clustering · KDD 2010
Performance modeling and evaluation › workload characterization › memory characterization
memory access characterization
0.112008
HMTT: a platform independent full-system memory trace monitoring system · SIGMETRICS 2008
Performance modeling and evaluation › trace analysis
memory trace analysis
0.112008
HMTT: a platform independent full-system memory trace monitoring system · SIGMETRICS 2008
Performance modeling and evaluation
workload characterization
0.112008
HMTT: a platform independent full-system memory trace monitoring system · SIGMETRICS 2008

Methods — techniques the papers use, named apart from their topics

inverse kinematics · 0.9SMPL · 0.9guided machine learning · 0.5generative adversarial network · 0.5mobility-based clustering · 0.2simulation · 0.1graph modeling · 0.1continuum percolation · 0.1hardware monitoring · 0.1DIMM snooping · 0.1
YearPublicationVenuePosition
2026 CLIP-based dual temporal decoupling network for video action recognition
Yinbin Zhang, Jing Sun 0010, Jianping Fan 0002
Neurocomputing7
2025 A Twist Representation and Shape Refinement Method for Human Mesh Recovery
abstract
3D human mesh recovery from single RGB images or monocular videos is a challenging task. The twist representation utilized in existing inverse kinematics-based methods fails to accurately describe the twisting posture when the estimated bone direction is imprecise. Additionally, supervising SMPL shape parameters has the issue of shape estimation overfitting due to limited training data. This often results in compromised bone lengths that subsequently impair the precision of joint positions. To address these issues, we propose a framework that breaks down both human pose and shape into finer components, effectively managing and minimizing errors within each component. The proposed framework integrates two key advancements: the advanced Ortho-Twist and Swing Representation (OTSR) and the Skeleton-Focused Shape Refinement (SFSR). OTSR offers a more sophisticated representation for limb rotations compared to the traditional twist angle and swing representation to enhance the accuracy of twisting posture estimation. SFSR refines the estimated SMPL shape parameters by fitting bone lengths using the estimated joint positions, thereby significantly mitigating shape overfitting and enhancing joint position accuracy in the recovered mesh. We conduct experiments on the Human3.6 M and 3DPW datasets. The results demonstrate the superiority of the proposed framework in both single-image and video scenarios. Additionally, the ablation studies confirm the effectiveness of our proposed modules, and further generalizability experiments demonstrate that our two key advancements can serve as plug-and-play modules to enhance existing methods.
Xiaoyang Hao, Jing Sun 0010, Lei Wang 0018, Jianping Fan 0002
IEEE Trans. Multim.5
2024 A CNN_LSTM_KAN Based Genetic Algorithm for Photovoltaic Power Generation Revenue Prediction
Ruikang Ma, Guanyu Lin, De Dong, Jianping Fan 0002, Keliang Duan
PDCAT5
2022 nGIA: A novel Greedy Incremental Alignment based algorithm for gene sequence clustering
Zhen Ju, Jintao Meng 0001, Jianping Fan 0002, Yi Pan 0001, Xuelei Li, Yanjie Wei
Future Gener. Comput. Syst.5
2022 SOCA-DOM: A Mobile System-on-Chip Array System for Analyzing Big Data on the Move
Le-Le Li, Jiang-Yi Liu, Jianping Fan 0002, Xuehai Qian, Kai Hwang 0001, Yeh-Ching Chung, Zhibin Yu 0001
J. Comput. Sci. Technol.3
2022 Boundary-aware context neural network for medical image segmentation
Ruxin Wang 0001, Shuyuan Chen, Chaojie Ji, Jianping Fan 0002, Ye Li 0002
Medical Image Anal.4
2021 An Efficient Greedy Incremental Sequence Clustering Algorithm
Zhen Ju, Jingtao Meng, Xuelei Li, Jianping Fan 0002, Yi Pan 0001, Yanjie Wei
ISBRA6
2021 A MVCC Approach to Parallelizing Interoperability of Consortium Blockchain
Weiyi Lin, Qiang Qu 0001, Li Ning 0001, Jianping Fan 0002, Qingshan Jiang
PDCAT4
2021 An Effective and Reliable Cross-Blockchain Data Migration Approach
Mengqiu Zhang, Qiang Qu 0001, Li Ning 0001, Jianping Fan 0002, Ruijie Yang
PDCAT4
2021 GML: Efficiently Auto-Tuning Flink's Configurations Via Guided Machine Learning
abstract
The increasingly popular fused batch-streaming big data framework, Apache Flink, has many performance-critical as well as untamed configuration parameters. However, how to tune them for optimal performance has not yet been explored. Machine learning (ML) has been chosen to tune the configurations for other big data frameworks (e.g., Apache Spark), showing significant performance improvements. However, it needs a long time to collect a large amount of training data by nature. In this article, we propose a guided machine learning (GML) approach to tune the configurations of Flink with significantly shorter time for collecting training data compared to traditional ML approaches. GML innovates two techniques. First, it leverages generative adversarial networks (GANs) to generate a part of training data, reducing the time needed for training data collection. Second, GML guides a ML algorithm to select configurations that the corresponding performance is higher than the average performance of random configurations. We evaluate GML on a lab cluster with 4 servers and a real production cluster in an internet company. The results show that GML significantly outperforms the state-of-the-art, DAC (Datasize-Aware-Configuration) (Z. Yu et al. 2018) for tuning the configurations of Spark, with 2.4× of reduced data collection time but with 30 percent reduced 99th percentile latency. When GML is used in the internet company, it reduces the latency by up to 57.8× compared to the configurations made by the company.
Yijin Guo, Huasong Shan, Shixin Huang, Kai Hwang 0001, Jianping Fan 0002, Zhibin Yu 0001
IEEE Trans. Parallel Distributed Syst.5
2020 A decentralised approach for link inference in large signed graphs
Muhammad Muzammal, Faima Abbasi, Qiang Qu 0001, Romana Talat, Jianping Fan 0002
Future Gener. Comput. Syst.5
2020 AirCargoChain: A Distributed and Scalable Data Sharing Approach of Blockchain for Air Cargo
Gejun Le, Qifeng Gu, Qiang Qu 0001, Qingshan Jiang, Jianping Fan 0002
J. Grid Comput.5
2020 Deep Multi-Scale Fusion Neural Network for Multi-Class Arrhythmia Detection
abstract
Automated electrocardiogram (ECG) analysis for arrhythmia detection plays a critical role in early prevention and diagnosis of cardiovascular diseases. Extracting powerful features from raw ECG signals for fine-grained diseases classification is still a challenging problem today due to variable abnormal rhythms and noise distribution. For ECG analysis, the previous research works depend mostly on heartbeat or single scale signal segments, which ignores underlying complementary information of different scales. In this paper, we formulate a novel end-to-end Deep Multi-Scale Fusion convolutional neural network (DMSFNet) architecture for multi-class arrhythmia detection. Our proposed approach can effectively capture abnormal patterns of diseases and suppress noise interference by multi-scale feature extraction and cross-scale information complementarity of ECG signals. The proposed method implements feature extraction for signal segments with different sizes by integrating multiple convolution kernels with different receptive fields. Meanwhile, joint optimization strategy with multiple losses of different scales is designed, which not only learns scale-specific features, but also realizes cumulatively multi-scale complementary feature learning during the learning process. In our work, we demonstrate our DMSFNet on two open datasets (CPSC_2018 and PhysioNet/CinC_2017) and deliver the state-of-art performance on them. Among them, CPSC_2018 is a 12-lead ECG dataset and CinC_2017 is a single-lead dataset. For these two datasets, we achieve the F1 score [Formula: see text] and [Formula: see text] which are higher than previous state-of-art approaches respectively. The results demonstrate that our end-to-end DMSFNet has outstanding performance for feature extraction from a broad range of distinct arrhythmias and elegant generalization ability for effectively handling ECG signals with different leads.
Ruxin Wang 0001, Jianping Fan 0002, Ye Li 0002
IEEE J. Biomed. Health Informatics2
2015 A restricted Boltzmann machine based two-lead electrocardiography classification
abstract
An restricted Boltzmann machine learning algorithm were proposed in the two-lead heart beat classification problem. ECG classification is a complex pattern recognition problem. The unsupervised learning algorithm of restricted Boltzmann machine is ideal in mining the massive unlabelled ECG wave beats collected in the heart healthcare monitoring applications. A restricted Boltzmann machine (RBM) is a generative stochastic artificial neural network that can learn a probability distribution over its set of inputs. In this paper a deep belief network was constructed and the RBM based algorithm was used in the classification problem. Under the recommended twelve classes by the ANSI/AAMI EC57: 1998/(R)2008 standard as the waveform labels, the algorithm was evaluated on the two-lead ECG dataset of MIT-BIH and gets the performance with accuracy of 98.829%. The proposed algorithm performed well in the two-lead ECG classification problem, which could be generalized to multi-lead unsupervised ECG classification or detection problems.
Yan Yan 0022, Xinbing Qin, Yige Wu, Jianping Fan 0002, Lei Wang 0029
BSN5
2014 Bandwidth-Availability-Based Replication Strategy for P2P VoD Systems
abstract
In a peer-to-peer (P2P) video-on-demand (VoD) system, each peer contributes a limited disc storage and stores some watched movies to offload the servers when these movies are requested. When the contributed disc storage of a peer is full, to minimize the server load, which movie should be replaced is a key design problem for P2P VoD systems. This problem is a P2P replication problem. Previous studies on this problem mainly consider content availability, but fail to consider bandwidth availability of peers on the system level. In this paper, assuming that movie popularity is known, we first analyze bandwidth availability of peers on the system level, and then formulate the replication problem as a minimization problem. Based on the minimization problem, we derive two design guidelines. According to the two design guidelines, we propose a bandwidth-availability-based replication algorithm aiming at minimizing the server load, called BAB algorithm. BAB algorithm can make the replicas’ distribution towards the optimal distribution in a distributed way. Furthermore, we consider some practical implementation issues of BAB algorithm. Through extensive simulations, we demonstrate that BAB algorithm outperforms the previously proposed algorithms in terms of reducing the server load and improving the streaming quality, in stable environment and dynamic environment, respectively.
Pingshan Liu, Shengzhong Feng, Guimin Huang, Jianping Fan 0002
Comput. J.4
2014 MR-DBSCAN: a scalable MapReduce-based DBSCAN algorithm for heavily skewed data
Yaobin He, Haoyu Tan, Wuman Luo, Shengzhong Feng, Jianping Fan 0002
Frontiers Comput. Sci.5
2014 Quick attribute reduction in inconsistent decision tables
Min Li 0020, Changxing Shang, Shengzhong Feng, Jianping Fan 0002
Inf. Sci.4
2014 Hierarchical clustering algorithm for categorical data using a probabilistic rough set model
Min Li 0020, Shaobo Deng, Lei Wang 0191, Shengzhong Feng, Jianping Fan 0002
Knowl. Based Syst.5
2014 Interference-aware spectrum handover for cognitive radio networks
abstract
ABSTRACT Cognitive radio (CR) is a promising technique for future wireless networks, which significantly improves spectrum utilization. In CR networks, when the primary users (PUs) appear, the secondary users (SUs) have to switch to other available channels to avoid the interference to PUs. However, in the multi‐SU scenario, it is still a challenging problem to make an optimal decision on spectrum handover because of the the accumulated interference constraint of PUs and SUs. In this paper, we propose an interference‐aware spectrum handover scheme that aims to maximize the CR network capacity and minimize the spectrum handover overhead by coordinating SUs’ handover decision optimally in the PU–SU coexisted CR networks. On the basis of the interference temperature model, the spectrum handover problem is formulated as a constrained optimization problem, which is in general a non‐deterministic polynomial‐time hard problem. To address the problem in a feasible way, we design a heuristic algorithm by using the technique of Branch and Bound. Finally, we combine our spectrum handover scheme with power control and give a convenient solution in a single‐SU scenario. Experimental results show that our algorithm can improve the network performance efficiently.Copyright © 2012 John Wiley & Sons, Ltd.
Dianjie Lu, Xiaoxia Huang 0004, Weile Zhang, Jianping Fan 0002
Wirel. Commun. Mob. Comput.4
2013 Event-Driven High-Priority First Data Scheduling Scheme for P2P VoD Streaming
abstract
The peer churn rate in the peer-to-peer (P2P) video-on-demand streaming service is much higher than in the P2P live streaming service, which makes the data scheduling problem more challenging. First, the available upload bandwidth information used in the data scheduling scheme is often inaccurate due to the peer churn, which lets a peer make bad scheduling decisions and leads to load imbalance. The higher peer churn makes this problem worse. Secondly, the higher peer churn exacerbates the bandwidth contention problem which occurs between a newly joined peer and some already-existing peers. To tackle the above two challenges, we propose an event-driven high-priority first data scheduling scheme, called EHPF scheme. To tackle the first challenge, we design a piggyback mechanism based on the event-driven mechanism. To tackle the second challenge, we design a priority calculation strategy to differentiate the requests from the newly joined peers and those from the already-existing peers, and use the high-priority first policy to allocate the upload bandwidths of peers. Through simulations and a real environment experiment, we demonstrate that the EHPF scheme outperforms the periodical data scheduling scheme in terms of startup delay, streaming quality and load balancing.
Pingshan Liu, Guimin Huang, Shengzhong Feng, Jianping Fan 0002
Comput. J.4
2013 Feature selection via maximizing global information gain for text classification
Changxing Shang, Min Li 0020, Shengzhong Feng, Qingshan Jiang, Jianping Fan 0002
Knowl. Based Syst.5
2012 Connectivity of large-scale Cognitive Radio Ad Hoc Networks
abstract
Connectivity of large-scale wireless networks has received considerable attention in the past several years. Different from traditional wireless networks, in Cognitive Radio Ad-hoc Networks (CRAHNs), primary users have spectrum access priority of the licensed bands over secondary users. Therefore, the connectivity of the secondary network is affected by not only the density and transmission power of secondary users, but also the activities of primary users. In addition, the number of licensed bands also has impact on the connectivity of CRAHNs. To capture the dynamic characteristics of opportunistic spectrum access, we introduce the Cognitive Radio Graph Model (CRGM) which takes into account the impact of the number of channels and the activities of primary users. Furthermore, we combine the CRGM with continuum percolation model to study the connectivity in the secondary network. We prove that secondary users can form the percolated network when the density of primary users is below the critical density. Then, the upper bound of the critical density of the primary users in the percolated CRAHNs is derived. Simulation results show that both the number of channels and the activities of primary users greatly impact the connectivity of CRAHNs.
Dianjie Lu, Xiaoxia Huang 0004, Pan Li 0001, Jianping Fan 0002
INFOCOM4
2012 Topic oriented community detection through social objects and link analysis in social networks
Zhongying Zhao 0001, Shengzhong Feng, Qiang Wang 0053, Joshua Zhexue Huang, Graham J. Williams, Jianping Fan 0002
Knowl. Based Syst.6
2012 Performance analysis and optimization of MPI collective operations on multi-core clusters
Bibo Tu, Jianping Fan 0002, Jianfeng Zhan
J. Supercomput.2
2011 MR-DBSCAN: An Efficient Parallel Density-Based Clustering Algorithm Using MapReduce
abstract
Data clustering is an important data mining technology that plays a crucial role in numerous scientific applications. However, it is challenging due to the size of datasets has been growing rapidly to extra-large scale in the real world. Meanwhile, MapReduce is a desirable parallel programming platform that is widely applied in kinds of data process fields. In this paper, we propose an efficient parallel density-based clustering algorithm and implement it by a 4-stages MapReduce paradigm. Furthermore, we adopt a quick partitioning strategy for large scale non-indexed data. We study the metric of merge among bordering partitions and make optimizations on it. At last, we evaluate our work on real large scale datasets using Hadoop platform. Results reveal that the speedup and scale up of our work are very efficient.
Yaobin He, Haoyu Tan, Wuman Luo, Huajian Mao, Shengzhong Feng, Jianping Fan 0002
ICPADS7
2011 Improving Data Locality of MapReduce by Scheduling in Homogeneous Computing Environments
abstract
Data Locality is one of the critical factors to affect performance. This paper proposes a next-k-node scheduling (NKS) method to improve the data locality of map tasks. The method first calculates the probabilities of each map task, and then preferentially schedules the one with the highest probability. It generates low probabilities for the tasks which satisfy node locality with the nodes to issue requests, so it can reserve these tasks to these nodes. We have implemented the NKS method in hadoop-0.20.2. The experiment results have shown that the NKS method reduced 78% of the map tasks processed without node locality, reduced 77%of the network load caused by the tasks, and improved the performance of Hadoop MapReduce when comparing with the default task scheduling method in Hadoop. Obviously, the NKS method is very suitable for the homogeneous environment with network overload.
Zhiyong Zhong, Shengzhong Feng, Bibo Tu, Jianping Fan 0002
ISPA5
2011 Info-Cluster Based Regional Influence Analysis in Social Networks
Chao Li 0022, Zhongying Zhao 0001, Jun Luo 0008, Jianping Fan 0002
PAKDD (2)4
2011 Channel capacity optimization via exploiting multi-SU coexistence in Cognitive Radio Networks
abstract
In Cognitive Radio Networks (CRNs), when the Primary Users (PUs) appear, the SUs have to evacuate the licensed spectrum in use or reduce the transmit power so that no harmful interference is introduced to the PUs. In this paper, we explore the multiple Secondary Users (SUs) coexistence system in CRNs based on power control mechanism and interference temperature model. We propose an optimal solution that can maximize the channel capacity and minimize the spectrum handover overhead, constrained by the accumulated interference of both the SUs-to-PU and SUs-to-SUs. We formulate this problem as a non-linear optimization problem and propose a heuristic algorithm to solve it efficiently. Experimental results show that compared with two alternative approaches, our algorithm can improve the usage of the spectrum by up to 51% (with a random approach) and up to 278% (with a conservative approach).
Dianjie Lu, Xiaoxia Huang 0004, Jianping Fan 0002
WCNC4
2011 An effective discretization based on Class-Attribute Coherence Maximization
Min Li 0020, Shaobo Deng, Shengzhong Feng, Jianping Fan 0002
Pattern Recognit. Lett.4
2010 QTL: An efficient scheduling policy for 10Gbps network intrusion detection system
abstract
Broad network bandwidth and deep inspection impose great challenge for the capability of 10Gpbs network security monitoring. Proper scheduling policies can improve system capability without requiring additional resources. LAS, a size-based scheduling policy which can achieve optimal mean response time by giving preferential analysis to short flows, is widely used in various aspects of network field. Due to the high variability property of Internet traffic, LAS favors short flows without penalizing large flows very much. Unfortunately, the inspection of large flows can not be guaranteed in those network intrusion detection systems on 10Gbps links, which are usually heavily loaded, or even overloaded. Although tiny in percentage, large flows comprise more than 50% of the total load, and therefore can not be ignored, especially when specified by users as critical. How to avoid starving large flows while still giving higher priority to short flows is a dilemma we have to face in practice. In this paper, we propose a QoS-supported three-level scheduling policy (QTL), which can remedy LAS' defect. The experimental results show that our QTL scheduling policy has approximately the same performance as LAS for short flows, and meanwhile exhibits greatly enhanced processing capability for large flows.
Weibing Yang, Mingyu Chen 0001, Jianping Fan 0002
ISCC5
2010 Towards mobility-based clustering
abstract
Identifying hot spots of moving vehicles in an urban area is essential to many smart city applications. The practical research on hot spots in smart city presents many unique features, such as highly mobile environments, supremely limited size of sample objects, and the non-uniform, biased samples. All these features have raised new challenges that make the traditional density-based clustering algorithms fail to capture the real clustering property of objects, making the results less meaningful. In this paper we propose a novel, non-density-based approach called mobility-based clustering. The key idea is that sample objects are employed as "sensors" to perceive the vehicle crowdedness in nearby areas using their instant mobility, rather than the "object representatives". As such the mobility of samples is naturally incorporated. Several key factors beyond the vehicle crowdedness have been identified and techniques to compensate these effects are proposed. We evaluate the performance of mobility-based clustering based on real traffic situations. Experimental results show that using 0.3% of vehicles as the samples, mobility-based clustering can accurately identify hot spots which can hardly be obtained by the latest representative algorithm UMicro.
Siyuan Liu 0001, Yunhuai Liu, Lionel M. Ni, Jianping Fan 0002, Minglu Li 0001
KDD4
2010 Robust TCP Reassembly with a Hardware-Based Solution for Backbone Traffic
abstract
There is a growing interest in designing high-speed network devices to perform packet processing at stream layer. However, variety kinds of out-of- sequence packets in real traffic will make trouble for hardware-based TCP reassembly system which is less flexible for exceptional processing. In this paper, we present a detailed analysis of behavior characteristic of out-of-sequence packets in real backbone traffic and propose a hardware-based solution for TCP reassembly based on the results of analysis. This solution could reassemble a TCP stream with concurrent multi discontinuous out-of-order data segments and provide a robust buffer management strategy for out-of-order data. We have also assessed the memory size and bandwidth required for reassembling real 10G traffic. The simulation result shows that the system can process over 99% of the 10G backbone traffic using reasonable storage resources. A FPGA-based prototype is also implemented for evaluation.
Yuan Ruan, Weibing Yang, Mingyu Chen 0001, Jianping Fan 0002
NAS5
2010 Achieving Flow-Level Controllability in Network Intrusion Detection System
abstract
Current network intrusion detection systems are lack of controllability, manifested as significant packet loss due to the long-term resources occupation by a single flow. The reasons can be classified into two kinds. The first kind is known as normal reasons, that is, the processing of mass arriving packets of a large flow can not be limited to a determinable period of time and thus makes other flows starved. The second kind, in which the CPU is trapped in a dead-loop like state due to processing some packets with particular content of a flow, is considered as abnormal reasons. In fact, it is a kind of software crashes. In this paper, we discuss the innate defects of traditional packet-driven NIDS, and implement a flow-driven framework which can achieve fine-grained controllability. An Active Two-threshold scheme based on ideal Exit-Point (ATEP) is proposed in order to diminish data preserving overhead during flow switches and to detect crash in time. A quick crash recovery mechanism is also given which can recover the trapped thread from 90% crashes in 0.2 ms. The experimental results show that our flow-driven framework with ATEP scheme can achieve higher throughput and less packet loss ratio than the uncontrollable packet-driven systems with less than 1% of extra CPU overhead. What's more, in the case of crash occurrence, the ATEP scheme is still able to maintain rather steady throughput without sudden decrease.
Weibing Yang, Mingyu Chen 0001, Jianping Fan 0002
SNPD5
2010 Adaptive Power Control Based Spectrum Handover for Cognitive Radio Networks
abstract
This paper focuses on spectrum handover in cognitive radio networks where secondary users (SUs) opportunistically use licensed channels as long as the aggregate interference at the primary users (PUs) does not exceed a certain threshold. We incorporate power control into the proposed spectrum handover scheme to reduce the number of spectrum handovers and enhance the spectral efficiency. In our work, when a PU arrives, an SU first calculates the maximum transmission power that the SU does not interfere with the PU. If the SU can still reach its receiver, it lowers its power and continues to transmit; otherwise it switches to an idle band. Analysis results show that our proposed scheme can substantially reduce the spectrum handover ratio and improve the effective data rate by up to 30%.
Dianjie Lu, Xiaoxia Huang 0004, Jianping Fan 0002
WCNC4
2009 Probabilistic-constrained fuzzy logic for situation modeling
abstract
How to model situation user-friendly and precisely is a key issue for situation-aware applications. Fuzzy logic is an effective approach to model situation, but one obstacle is how to select the suitable operators between different fuzzy sets. One possibility is to combine the merit of both Fuzzy logic and Probability logic. The paper first introduces a set of constraints on conventional fuzzy logic and its operations, to setup a unified framework so as to combine the merits of the above two approaches. Such probabilistic-constrained fuzzy logic can be used in situation-aware applications. The paper then focuses on how to derive new fuzzy concepts from basic fuzzy partition, and how to compute the relationship between such derived and basic fuzzy concepts according to the probability constraints, which is different from the conventional ones.
Jinhua Xiong, Jianping Fan 0002
FUZZ-IEEE2
2009 Towards virtually cooking Chinese food
abstract
Chinese food is delicious but cooking Chinese food is a very complex process which involves in combing ingredients at the right time and temperature. In this paper, we present a multimedia technology for virtually cooking Chinese food. We focus on a popular Chinese food, shredded potato. We propose to use the theory from mechanics of materials to model shredded potato in its cooking process. The shredded potato is initially simplified with changed deformable beams by adjusting model parameters during the cooking process when shredded potato gets dasiasoftpsila deformations. We use superposition principle to cope with the multi-load problem when potato shreds pile up and contact with each other. We describe the modeling techniques and implementation issues in detail. We show the result from the proposed modeling technique by comparing it with the real image.
Jinghao Fei, Jie Yang 0001, Jianping Fan 0002
ICME3
2009 Confusion network based Video OCR post-processing approach
abstract
The paper originally presents a confusion network based framework for video OCR post-processing. The framework consists of four parts: selection of reference and hypotheses, construction of confusion network, decoding for final output, and a novel metric of quantitatively evaluating Video OCR post-processing approaches. By integrating both visual and textual information, we construct the character transition network to reduce the error rate for OCR outputs. The large-scale experimental results demonstrate that this approach can significantly improve the accuracy of Video OCR results with only little incremental time. Moreover, with comparison and the detailed analysis, we conclude that ldquoVoting+2-gramrdquo is the most applicable method for real application.
Anan Liu, Jinghao Fei, Jianping Fan 0002, Lin Pang, Yongdong Zhang 0001, Jintao Li 0001
ICME3
2009 Accurate Analytical Models for Message Passing on Multi-core Clusters
abstract
Memory hierarchy on multi-core clusters has two-fold characteristics: vertical memory hierarchy and horizontal memory hierarchy. Vertical memory hierarchy has been modeled by previous work (e.g. memory logP, lognP, log3P etc.) to analyze middlewarepsilas effects on point-to-point communication with different message sizes and message strides; Horizontal memory hierarchy has become more prominent due to distinct performance among three levels of communication in a multi-core cluster: intra-CMP, inter-CMP and inter-node, which should adequately be considered. Derived from lognP and log3P models, new analytical models mlognP and its reduction 2log{2,3}P are proposed to unitedly abstract memory hierarchy on multi-core clusters in vertical and horizontal levels. The results of performance evaluation show that it is indispensable to incorporate horizontal memory hierarchy into new models suitable for multi-core clusters, and 2log{2,3}P model can predict communication costs for message passing on multi-core clusters more accurately than log3P model.
Bibo Tu, Jianping Fan 0002, Jianfeng Zhan
PDP2
2008 Multi-core aware optimization for MPI collectives
abstract
MPI collective operations on multi-core clusters should be multi-core aware. In this paper, collective algorithms with hierarchical virtual topology focus on the performance difference among different communication levels on multi-core clusters, simply for intra-node and inter-node communication; Furthermore, to select befitting segment sizes for intra-node collective communication can cater to cache hierarchy in multi-core processors. Based on existing collective algorithms in MPICH2, above two techniques construct portable optimization methodology over MPICH2 for collective operations on multi-core clusters. Conforming to above optimization methodology, multi-core aware broadcast algorithm has been implemented and evaluated as a case study. The results of performance evaluation show that the multi-core aware optimization methodology over MPICH2 is efficient.
Bibo Tu, Ming Zou, Jianfeng Zhan, Jianping Fan 0002
CLUSTER5
2008 PGDC: Parallel Generation of Digital City
abstract
Digital cities are paid more and more attentions on for they are expected to make life, business, travel, city planning and so on more convenient and more effective. Existing digital cities are generated based on serial computing. However, in serial computing the speed of generating a digital city is so slow that digital cities can not update in sync with reality. Since parallel computing is more powerful and faster than serial computing, the parallel computing can be used to speed up and enhance the scale of the generation of digital city based on large number of parallelism in digital city. We studied Parallel Generation of Digital City (PGDC) to generate digital cities based on parallel computing.
Dingju Zhu, Jianping Fan 0002
CW2
2008 Application of Parallel Computing in Digital City
abstract
Digital cities are paid more and more attentions on for they are expected to make life, business, travel, city planning and so on more convenient and more effective. Existing digital cities are based on serial computing. However, the speed of the applications of digital city in serial computing is so slow that the applications can not run in real time or run in large scale. Since parallel computing is more powerful and faster than serial computing, parallel computing is used in our study to promote the scale and the speed of the applications of digital city. An experiment on recognizing objects in digital city based on parallel computing is given.
Dingju Zhu, Jianping Fan 0002
HPCC2
2008 GMIP: A Novel Optical Interconnect Gridded Memory Service Protocol
abstract
Dynamic self-organized computer architecture (DSAG) based on Grid-components departs computer components to grid components, and dynamically aggregates and organizes these components to realize architecture-on-demand. We take the important feature, that CPU centered design principle should possibly become memory centered design and optimization. We propose a novel computer architecture Gridded Memory Service (GMS), which is based on DSAG. We design and implement a serial, light weight and packet switching optical interconnect protocol, Gridded Memory Interconnect Protocol (GMIP), which is featured as high bandwidth and low latency protocol. It optimizes the link and physical layer to take advantage of very short reach optical interconnect technology. At last we study interconnect effects, and propose the main evaluation principles of bandwidth latency compensation.
Siyuan Liu 0001, Lei Li 0005, Jianping Fan 0002
ICPADS4
2008 Design Techniques for the Scalability of Cluster Management Software on Dawning Supercomputers
abstract
Cluster management software has faced more increased scalability challenge with ever enlarged cluster scale. Its good scalability rests with feasible design techniques focusing on hybrid software topologies with partitioning policy, non-blocking I/O multiplexing and message on demand. Design patterns are generic solutions to recurring software design problems, and above three important techniques are abstracted the design pattern of scalable cluster management software in this paper. According to this design pattern, some cluster management tools, such as job scheduling, MPI job launcher and so on, have been designed and applied on Dawning supercomputers. Some results of performance evaluation have shown that good scalability of cluster management software on Dawning supercomputers has benefited from this design pattern.
Bibo Tu, Ming Zou, Jianfeng Zhan, Jianping Fan 0002
ISPA4
2008 HMTT: a platform independent full-system memory trace monitoring system
abstract
Memory trace analysis is an important technology for architecture research, system software (i.e., OS, compiler) optimization, and application performance improvements. Many approaches have been used to track memory trace, such as simulation, binary instrumentation and hardware snooping. However, they usually have limitations of time, accuracy and capacity.In this paper we propose a platform independent memory trace monitoring system, which is able to track virtual memory reference trace of full systems (including OS, VMMs, libraries, and applications). The system adopts a DIMM-snooping mechanism that uses hardware boards plugged in DIMM slots to snoop. There are several advantages in this approach, such as fast, complete, undistorted, and portable. Three key techniques are proposed to address the system design challenges with this mechanism: (1) To keep up with memory speeds, the DDR protocol state machine is simplified, and large FIFOs are added between the state machine and the trace transmitting logic to handle burst memory accesses; (2) To reconstruct physical-tovirtual mapping and distinguish one process' address space from others, an OS kernel module, which collects page table information, and a synchronization mechanism, which synchronizes the page table information with the memory race, are developed; (3) To dump massive trace data, we employ a straightforward method to compress the trace and use Gigabit Ethernet and RAID to send and receive the compressed trace.We present our implementation of an initial monitoring system, named HMTT (Hyper Memory Trace Tracker). Using HMTT, we have observed that burst bandwidth utilization is much larger than average bandwidth utilization, by up to 5X in desktop applications. We have also confirmed that the stream memory accesses of many applications contribute even more than 40% of L2 Cache misses and OS virtual memory management may decrease stream accesses in view of memory controller (or L2 Cache), by up to 30.2%. Moreover, we have evaluated OS impact on memory performance in real systems. The evaluations and case studies show the feasibility and effectiveness of our proposed monitoring mechanism and techniques.
Yungang Bao, Mingyu Chen 0001, Yuan Ruan, Li Liu 0038, Jianping Fan 0002, Qingbo Yuan
SIGMETRICS5
2007 A Strategy-Proof Combinatorial Auction-Based Grid Resource Allocation System
Jianping Fan 0002, Ruihua Di
ICA3PP2
2007 DCGG: Digital City Group Grid
abstract
There is a serious problem in digital city construction which has not been intentioned and solved before. Different cities construct different digital cities using different data formats and different software, which causes that there are many separated digital cities which can not communicate with each other. We applied grid technology successfully in solving the problem, digital city group grid (DCGG) can support the communications and collaborations between different digital city application systems and can make all digital cities in DCGG can run together as a large digital city. Users can visit one digital city after another without logging out or switching and can search any information among all digital cities in DCGG.
Dingju Zhu, Jianping Fan 0002
PDCAT2
2006 Design Patterns of Scalable Cluster System Software
abstract
The design pattern of cluster system software has an important influence on scalability of massive cluster system. The paper presents design patterns of scalable cluster system software, including scalable software topologies and optimized communication modes. These design patterns have been widely applied in Dawning series of supercomputers and some results of performance evaluation show their good scalability
Bibo Tu, Ming Zou, Jianfeng Zhan, Lei Wang 0004, Jianping Fan 0002
PDCAT5
2006 Observability Statement Coverage Based on Dynamic Factored Use-Definition Chains for Functional Verification
Tao Lv 0001, Jianping Fan 0002, Xiaowei Li 0001, Ling-Yi Liu
J. Electron. Test.2
2005 Automatic Parsing of Sports Videos with Grammars
Kevin Lü 0001, Jintao Li 0001, Jianping Fan 0002
DEXA4
2005 A Reconfigurable Optical Interconnect System for DSAG
abstract
High performance computing research is facing challenges and innovation on architecture is urgent. DSAG architecture is proposed and delivers "Architecture on Demand" feature. In DSAG, components in different catalogs are parted, while the ones in same catalog are congregated. This architecture can be enabled by optical interconnect and reconfigurable computing technology. Using advanced optical devices and enhanced reconfigurable computing devices (FPGA), we build a prototype system for DSAG. Optical interconnect can reach 16Gbps bandwidth; DDRAM interface is selected as host communication interface to match the bandwidth of optical channel; Reconfigurable logic and embedded processors are employed for flexible reconfiguration. The system is featured by high bandwidth, owerful, flexible.
Lei Li 0005, Zheng Cao 0003, Mingyu Chen 0001, Jianping Fan 0002
PDCAT4
2004 HPL Performance Prevision to Intending System Improvement
Mingyu Chen 0001, Jianping Fan 0002
ISPA3
2003 An Efficient Observability Evaluation Algorithm Based on Factored Use-Def Chains
abstract
Coverage evaluation is indispensable for simulate modern designs. In this paper, we present an efficient algorithm to evaluate observability coverage, which is based on Factored Use-Def chains (FUD chains), a data-flow analysis technique in compilers. With the strategy of enhanced FUD chains and event-driven analysis, this method has three advantages. First, it is significantly more computationally efficient than prior efforts to assess observability information. Secondly, it could be easily integrated into hardware description language (HDL) compilers or simulators. Finally, it is universal, and can be combined with controllability metrics, such as statement coverage metric (SCM).
Tao Lv 0001, Jianping Fan 0002, Xiaowei Li 0001
Asian Test Symposium2