EDBT 2026 Demo / reviewers in the wild / expert
Jingyu Zhou
dblp:86/4786
· DBLP profile ↗
40ranked-venue papers
11as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-authorDatabases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Security and privacy · 4 · 2 first-authorArtificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Face, body and person analysis · 27% Vision and language · 23% Learning paradigms · 23% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 70% Information retrieval · 30% | |
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Storage systems · 64% Distributed systems · 23% Cloud and datacenter computing · 12% | |
| Software engineering, system software, and programming languages
2 papers |
Program analysis · 86% Operating systems · 14% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
multi-label classification |
1.0 | 1 | 2026 | Adapting Vision-Language Models from Iconic to Inclusive for Multi-label Recognition Without Labels · Int. J. Comput. Vis. 2026 |
Computer vision › Vision and language › vision-language model
vision-language model adaptation |
1.0 | 1 | 2026 | Adapting Vision-Language Models from Iconic to Inclusive for Multi-label Recognition Without Labels · Int. J. Comput. Vis. 2026 |
Data mining
anomaly detection |
1.0 | 1 | 2026 | ZUMA: Training-Free Zero-Shot Unified Multimodal Anomaly Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Data mining › anomaly detection
multimodal anomaly detection |
1.0 | 1 | 2026 | ZUMA: Training-Free Zero-Shot Unified Multimodal Anomaly Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-Id · IEEE Trans. Image Process. 2022 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.6 | 1 | 2022 | Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-Id · IEEE Trans. Image Process. 2022 |
Computer vision › Face, body and person analysis
person re-identification |
0.6 | 1 | 2022 | Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-Id · IEEE Trans. Image Process. 2022 |
Computer vision › Face, body and person analysis › person re-identification
unsupervised person re-identification |
0.6 | 1 | 2022 | Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-Id · IEEE Trans. Image Process. 2022 |
Distributed systems › distributed database
distributed transactions |
0.5 | 1 | 2021 | FoundationDB: A Distributed Unbundled Transactional Key Value Store · SIGMOD Conference 2021 |
Storage systems
key-value storage |
0.5 | 1 | 2021 | FoundationDB: A Distributed Unbundled Transactional Key Value Store · SIGMOD Conference 2021 |
Storage systems
storage reliability |
0.5 | 1 | 2021 | FoundationDB: A Distributed Unbundled Transactional Key Value Store · SIGMOD Conference 2021 |
Storage systems › key-value storage
transactional key-value store |
0.5 | 1 | 2021 | FoundationDB: A Distributed Unbundled Transactional Key Value Store · SIGMOD Conference 2021 |
Information retrieval › search engines
web crawling |
0.4 | 2 | 2016 | SmartCrawler: A Two-Stage Crawler for Efficiently Harvesting Deep-Web Interfaces · IEEE Trans. Serv. Comput. 2016 Extracting URLs from JavaScript via program analysis · ESEC/SIGSOFT FSE 2013 |
Information retrieval › search engines › web crawling
focused crawling |
0.2 | 1 | 2016 | SmartCrawler: A Two-Stage Crawler for Efficiently Harvesting Deep-Web Interfaces · IEEE Trans. Serv. Comput. 2016 |
Information retrieval › ranking › graph-based ranking
link-based ranking |
0.2 | 1 | 2016 | SmartCrawler: A Two-Stage Crawler for Efficiently Harvesting Deep-Web Interfaces · IEEE Trans. Serv. Comput. 2016 |
Program analysis › static analysis
program slicing |
0.2 | 1 | 2013 | Extracting URLs from JavaScript via program analysis · ESEC/SIGSOFT FSE 2013 |
Program analysis
static analysis |
0.2 | 1 | 2013 | Extracting URLs from JavaScript via program analysis · ESEC/SIGSOFT FSE 2013 |
Cloud and datacenter computing
cloud infrastructure |
0.1 | 1 | 2021 | FoundationDB: A Distributed Unbundled Transactional Key Value Store · SIGMOD Conference 2021 |
Wireless sensing and localization › RF sensing
RFID sensing |
0.1 | 1 | 2011 | TASA: Tag-Free Activity Sensing Using RFID Tag Arrays · IEEE Trans. Parallel Distributed Syst. 2011 |
Data mining › text mining
text classification |
0.1 | 1 | 2009 | A class-feature-centroid classifier for text categorization · WWW 2009 |
Storage systems
distributed storage |
0.1 | 2 | 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth Storage · SC 2004 A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 |
Operating systems › resource management › process management
CPU scheduling |
0.1 | 1 | 2006 | Request-Aware Scheduling for Busy Internet Services · INFOCOM 2006 |
Cloud and datacenter computing
load shedding |
0.1 | 1 | 2006 | Selective early request termination for busy internet services · WWW 2006 |
Cloud and datacenter computing
overload control |
0.1 | 1 | 2006 | Selective early request termination for busy internet services · WWW 2006 |
Distributed systems
fault tolerance |
0.1 | 1 | 2005 | Dependency isolation for thread-based multi-tier Internet services · INFOCOM 2005 |
Program analysis
dynamic analysis |
0.0 | 1 | 2013 | Extracting URLs from JavaScript via program analysis · ESEC/SIGSOFT FSE 2013 |
Distributed systems › peer-to-peer systems
consistent hashing |
0.0 | 1 | 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 |
Storage systems
data placement |
0.0 | 1 | 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 |
Storage systems
object storage |
0.0 | 1 | 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth Storage · SC 2004 |
Internet of things and sensor networks
RFID systems |
0.0 | 1 | 2011 | TASA: Tag-Free Activity Sensing Using RFID Tag Arrays · IEEE Trans. Parallel Distributed Syst. 2011 |
Methods — techniques the papers use, named apart from their topics
dynamic semantic interaction · 1.0cross-domain calibration · 1.0CLIP · 1.0self-supervised learning · 0.6contrastive learning · 0.6clustering · 0.6deterministic simulation · 0.5static analysis · 0.3statement coverage · 0.3program slicing · 0.3dynamic execution · 0.3site ranking · 0.2link tree data structure · 0.2reference tags · 0.1RSSI · 0.1feature centroid classifier · 0.1size-adaptive scheduling · 0.1dynamic feedback · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adapting Vision-Language Models from Iconic to Inclusive for Multi-label Recognition Without Labels
Jingyu Zhou, Yifan Zhao 0002, Jia Li 0003 |
Int. J. Comput. Vis. | 2 |
| 2026 | ZUMA: Training-Free Zero-Shot Unified Multimodal Anomaly DetectionabstractMultimodal anomaly detection (MAD) aims to exploit both texture and spatial attributes to identify deviations from normal patterns in complex scenarios. However, zero-shot (ZS) settings arising from privacy concerns or confidentiality constraints present significant challenges to existing MAD methods. To address this issue, we introduce ZUMA, a training-free, Zero-shot Unified Multimodal Anomaly detection framework that unleashes CLIP's cross-modal potential to perform ZS MAD. To mitigate the domain gap between CLIP's pretraining space and point clouds, we propose cross-domain calibration (CDC), which efficiently bridges the manifold misalignment through source-domain semantic transfer and establishes a hybrid semantic space, enabling a joint embedding of 2D and 3D representations. Subsequently, ZUMA performs dynamic semantic interaction (DSI) to enable structural decoupling of anomaly regions in the high-dimensional embedding space constructed by CDC, where natural languages serve as semantic anchors to help DSI establish discriminative hyperplanes within hybrid modality representations. Within this framework, ZUMA enables plug-and-play detection of 2D, 3D or multimodal anomalies, without training or fine-tuning even for cross-dataset or incomplete-modality scenarios. Additionally, to further investigate the potential of the training-free ZUMA within the training-based paradigm, we develop ZUMA-FT, a fine-tuned variant that achieves notable improvements with minimal parameter trade-off. Extensive experiments are conducted on two MAD benchmarks, MVTec 3D-AD and Eyecandies. Notably, the training-free ZUMA achieves state-of-the-art (SOTA) performance on both datasets, outperforming existing ZS MAD methods, including training-based approaches. Moreover, ZUMA-FT further extends the performance boundary of ZUMA with only 6.75 M learnable parameters. Yunfeng Ma, Min Liu 0008, Jingyu Zhou, Yuan Bian 0002, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | SCAP: Semantic Prototype Alignment for Robust Point Cloud RegistrationabstractPoint cloud rigid registration is a fundamental problem in robotics, 3D reconstruction, and augmented reality. However, existing methods predominantly rely on local geometric neighborhoods, which fail to capture higher-order semantic structures and thus degrade performance under noisy or complex geometry conditions. To address these limitations, we propose SCAP, a new point cloud registration paradigm that transforms feature interaction from geometry-driven to semantics–geometric co-driven. Specifically, a semantic prototype extractor is devised to abstract high-level semantic prototypes through graph embedding and clustering, thereby mitigating sensitivity to local feature noise. Since semantic abstraction alone cannot guarantee consistent correspondences across point clouds, SCAP performs a prototype alignment path learning to infer reliable semantic mappings through optimal transport. To enhance cross-layer feature integration and prevent redundant attention, an alignment-driven cross-layer transformer is proposed to incorporate the learned priors into the attention mechanism, thereby enabling feature aggregation with improved semantic coherence and local precision. Extensive experiments on ModelNet, ModelLoNet, 3DMatch, and 3DLoMatch demonstrate that our SCAP consistently surpasses state-of-the-art approaches, showing superior robustness and generalization in challenging scenarios with noise and partial overlap. The code will be available at https://github.com/Zhou-111jy/SCAP.git. Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Cross-View Dynamic Learning-Based Multi-Class Industrial Anomaly DetectionabstractIndustrial anomaly detection plays a crucial role in smart manufacturing. Traditional methods typically train separate models for each category, leading to substantial memory demands and computational cost. Moreover, relying solely on single-view images is prone to detection blind spots and poor sensitivity to subtle defects. To address these problems, this study proposes CVDL, a cross-view dynamic learning-based multi-class industrial anomaly detection method. Specifically, the CVDL leverages a proposed cross-view dynamic attention in conjunction with intra-view self-attention to dynamically modulate the model’s attention on multi-view information, thereby enhancing the detection performance of subtle defects. Furthermore, a category-guided prompt is developed to utilize object category information, which improves the model’s class-aware detection accuracy. To enhance the model’s robustness, we introduce a structured noise injection strategy and a region-wise mask into the CVDL, mitigating the “identity shortcut” that preserves anomalies during reconstruction. Extensive experiments on the authentic multi-view industrial datasets (Real-IAD) and well-known datasets (MVTec-AD and VisA) confirm the superior detection capability and robustness of the proposed CVDL, and the overall performance of CVDL is superior to all advanced approaches on Real-IAD, achieving SoTA performance of 90.1% image-level and 99.0% pixel-level AUROC. The code will be available at https://github.com/zfinn1/CVDL.git. Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Unified Multimodal Industrial Anomaly Detection via Few Normal SamplesabstractMultimodal industrial anomaly detection (MIAD) is the process of integrating multiple sensor data and utilizing visual intelligence to identify abnormal states in industrial production. In this article, we focus on two main practical but challenging issues in MIAD, i.e., a unified model for multiclass anomaly detection, and model training with only few normal samples. The current mainstream “one-for-one” paradigm requires training time that grows exponentially, and it relies on a sufficient number of samples (even just normal samples), which cannot adapt to practical industrial scenarios with rich abnormal classes. To this end, we offer aUnifiedMIAD model that trained using onlyFew (e.g., 1, 2, and 4) normal samples, termed UniMF. Specifically, we propose a fusion-guided prompt engineering process that generates paired antithetical instance-specific prompts with the assistance of multimodal fusion at both query and token levels. To enable cross-modal prompt learning under multimodal conditions, UniMF performs multi-proxy pairwise matching that involves alignment among multimodal feature patches, embeddings, and tokens of antithetical prompts. Experimental results show that UniMF stands state-of-the-art performance while remaining “one-for-all” paradigm, and even outperforms “one-for-one” methods under certain settings. Cross-dataset evaluation between MVTec 3D-AD and Eyecandies datasets also shows the transferability of UniMF. Yunfeng Ma, Jingyu Zhou, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Ind. Informatics | 3 |
| 2026 | Micro Surface Defect Inspection of Aero-Engine Blades via Dynamic Cross-Scale Semantic Aggregation
Kaijie Li, Jingyu Zhou, Xiangfei Meng, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Evaluation of Sustainable Economic and Environmental Development Evidence From OECD CountriesabstractThis study analyzes the economic and environmental performance of OECD countries over 2000–2019. A by-production approach is applied and the efficiency score is decomposed into its economic and environmental components. Unlike previous studies, we apply a refined model that allows for the correct modeling of by-production technology. The refined model can provide clear economic illustrations for balancing economic growth and environmental protection. The results indicate that environmental inefficiency is higher than the potential economic improvement. The environmental efficiency of OECD countries is improving, while economic performance is worsening over time. Therefore, instead of highly polluting energy, clean energy should be used to build a low-carbon economy. Worldwide, carbon-emitting countries and developed countries should shoulder their responsibilities to reduce carbon emissions and provide emission reduction funds for developing countries, while simultaneously sharing the green production technologies needed to reduce emissions. Jingyu Zhou, Xingyu Xu 0004, Ruochen Jiang, Jinyang Cai |
J. Glob. Inf. Manag. | 1 |
| 2022 | Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-IdabstractUnsupervised person re-identification (Re-Id) has attracted increasing attention due to its practical application in the read-world video surveillance system. The traditional unsupervised Re-Id are mostly based on the method alternating between clustering and fine-tuning with the classification or metric learning objectives on the grouped clusters. However, since person Re-Id is an open-set problem, the clustering based methods often leave out lots of outlier instances or group the instances into the wrong clusters, thus they can not make full use of the training samples as a whole. To solve these problems, we present the hybrid dynamic cluster contrast and probability distillation algorithm. It formulates the unsupervised Re-Id problem into an unified local-to-global dynamic contrastive learning and self-supervised probability distillation framework. Specifically, the proposed method can make the best of the self-supervised signals of all the clustered and un-clustered instances, from both the instances' self-contrastive level and the probability distillation respectives, in the memory-based non-parametric manner. Besides, the proposed hybrid local-to-global contrastive learning can take full advantage of the informative and valuable training examples for effective and robust training. Extensive experiment results show that the proposed method achieves superior performances to state-of-the-art methods, under both the purely unsupervised and unsupervised domain adaptation experiment settings. Our source code is released in https://github.com/zjy2050/HDCRL-ReID. De Cheng, Jingyu Zhou, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | FoundationDB: A Distributed Unbundled Transactional Key Value StoreabstractFoundationDB is an open source transactional key value store created more than ten years ago. It is one of the first systems to combine the flexibility and scalability of NoSQL architectures with the power of ACID transactions (a.k.a. NewSQL). FoundationDB adopts an unbundled architecture that decouples an in-memory transaction management system, a distributed storage system, and a built-in distributed configuration system. Each sub-system can be independently provisioned and configured to achieve the desired scalability, high-availability and fault tolerance properties. FoundationDB uniquely integrates a deterministic simulation framework, used to test every new feature of the system under a myriad of possible faults. This rigorous testing makes FoundationDB extremely stable and allows developers to introduce and release new features in a rapid cadence. FoundationDB offers a minimal and carefully chosen feature set, which has enabled a range of disparate systems (from semi-relational databases, document and object stores, to graph databases and more) to be built as layers on top. FoundationDB is the underpinning of cloud infrastructure at Apple, Snowflake and other companies, due to its consistency, robustness and availability for storing user data, system metadata and configuration, and other critical information. Jingyu Zhou, Meng Xu 0023, Alexander Shraer, Bala Namasivayam, Evan Tschannen, Steve Atherton, Andrew J. Beamon, Rusty Sears, John Leach, Dave Rosenthal, Will Wilson, Ben Collins, David Scherer, Alec Grieser, Young Liu, Alvin Moore, Bhaskar Muppana, Xiaoge Su, Vishesh Yadav |
SIGMOD Conference | 1 |
| 2017 | A Novel Noise-assisted Prognostic Method for Linear Analog Circuits
Liyue Yan, Houjun Wang, Zhen Liu 0003, Jingyu Zhou, Bing Long |
J. Electron. Test. | 4 |
| 2017 | Reverse Furthest Neighbors Query in Road Networks
Xiao-Jun Xu, Jinsong Bao, Bin Yao 0002, Jingyu Zhou, Feilong Tang 0001, Minyi Guo, Jianqiu Xu |
J. Comput. Sci. Technol. | 4 |
| 2016 | SmartCrawler: A Two-Stage Crawler for Efficiently Harvesting Deep-Web InterfacesabstractAs deep web grows at a very fast pace, there has been increased interest in techniques that help efficiently locate deep-web interfaces. However, due to the large volume of web resources and the dynamic nature of deep web, achieving wide coverage and high efficiency is a challenging issue. We propose a two-stage framework, namely SmartCrawler, for efficient harvesting deep web interfaces. In the first stage, SmartCrawler performs site-based searching for center pages with the help of search engines, avoiding visiting a large number of pages. To achieve more accurate results for a focused crawl, SmartCrawler ranks websites to prioritize highly relevant ones for a given topic. In the second stage, SmartCrawler achieves fast in-site searching by excavating most relevant links with an adaptive link-ranking. To eliminate bias on visiting some highly relevant links in hidden web directories, we design a link tree data structure to achieve wider coverage for a website. Our experimental results on a set of representative domains show the agility and accuracy of our proposed crawler framework, which efficiently retrieves deep-web interfaces from large-scale sites and achieves higher harvest rates than other crawlers. Feng Zhao 0003, Jingyu Zhou, Chang Nie, Heqing Huang 0001, Hai Jin 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2015 | Cowic: A Column-Wise Independent Compression for Log Stream AnalysisabstractNowadays massive log streams are generated from many Internet and cloud services. Storing log streams consumes a large amount of disk space and incurs high cost. Traditional compression methods can be applied to reduce storage cost, but are inefficient for log analysis, because fetching relevant log entries from compressed data often requires retrieval and decompression of large blocks of data. We propose a column-wise compression approach for well-formatted log streams, where each log entry can be independently compressed or decompressed for analysis. Specifically, we separate a log entry into several columns and compress each column with different models. We have implemented our approach as a library and integrated it into two applications, a log search system and a log joining system. Experimental results show that our compression scheme outperforms traditional compression methods for decompression times and has a competitive compression ratio. For log search, our approach achieves better query times than using traditional compression algorithms for both in-core and out-of-core cases. For joining log streams, our approach achieves the same join quality with only 30% memory of uncompressed streams. Jingyu Zhou, Bin Yao 0002, Minyi Guo, Jie Li 0002 |
CCGRID | 2 |
| 2015 | Fast Proof Generation for Verifying Cloud SearchabstractAs cloud computing has become prominent, the need for searching cloud data has grown increasingly urgent. However, cloud search may be incorrect due to errors of cloud providers and attacks from other malicious tenants. Previous work on verifiable computing returns results with probabilistically checkable proofs, which targets at different applications other than search and requires a large computation overhead. We propose a hybrid approach for generating proofs of cloud search results. Specifically, we model search indices as sets and search operations as set intersections, and build proofs based on RSA accumulators and aggregated membership and no membership witnesses. Because generating witnesses for large sets is computationally expensive, we employ interval-based witnesses for fast proof generation. To reduce proof size, our hybrid method uses Bloom filters when set difference is large. Evaluation on real datasets shows that our hybrid approach generates proofs in an average of 0.197s, up to 83.2% faster than previous work with a smaller proof size. Experiments also show our approach allows incremental updates with constant cost. Jingyu Zhou, Jiannong Cao 0001, Bin Yao 0002, Minyi Guo |
IPDPS | 1 |
| 2015 | LS-AMS: An Adaptive Indexing Structure for Realtime Search on MicroblogsabstractIndexing microblogs for realtime search is challenging, because new microblogs are created at tremendous speed, and user query requests keep constantly changing. To guarantee user obtain complete query results, micro-blogging site maintains huge indices which leads to index fragmentation or extra merging overhead during realtime search. This paper proposes an efficient LogStructured index structure with Adaptive Merging Strategy (LS-AMS) for realtime search on microblogs. LS-AMS structure consists of an inverted index buffer and a sequence of dynamically adjustable index packages with exponentially increasing sizes. These index packages manage their inverted indices using adaptive merging strategy, which can reduce the merging overhead to improve query performance and can adjust the index structure based on environmental factors, such as the arrival rate of query requests and new microblogs. Experimental results show that LS-AMS can greatly improve query performance without increasing the update cost and improve the self-adaptability in dynamic environment. Feng Zhao 0003, Jun Liu 0002, Jingyu Zhou, Hai Jin 0001, Laurence T. Yang |
IEEE Trans. Big Data | 3 |
| 2014 | LSShare: an efficient multiple query optimization system in the cloud
Xing Ge, Bin Yao 0002, Minyi Guo, Changliang Xu, Jingyu Zhou, Chentao Wu, Guangtao Xue |
Distributed Parallel Databases | 5 |
| 2014 | Preference-based mining of top-K influential nodes in social networks
Jingyu Zhou |
Future Gener. Comput. Syst. | 1 |
| 2013 | ShmStreaming: A Shared Memory Approach for Improving Hadoop Streaming PerformanceabstractThe Map-Reduce programming model is now drawing both academic and industrial attentions for processing large data. Hadoop, one of the most popular implementations of the model, has been widely adopted. To support application programs written in languages other than Java, Hadoop introduces a streaming mechanism that allows it to communicate with external programs through pipes. Because of the added overhead associated with pipes and context switches, the performance of Hadoop streaming is significantly worse than native Hadoop jobs. We propose ShmStreaming, a mechanism that takes advantages of shared memory to realize Hadoop streaming for better performance. Specifically, ShmStreaming uses shared memory to implement a lockless FIFO queue that connects Hadoop and external programs. To further reduce the number of context switches, the FIFO queue adopts a batching technique to allow multiple key-value pairs to be processed together. For typical benchmarks of word count, grep and inverted index, experimental results show 20-30% performance improvement comparing to the native Hadoop streaming implementation. Longbin Lai, Jingyu Zhou, Long Zheng 0001, Huakang Li, Yanchao Lu, Feilong Tang 0001, Minyi Guo |
AINA | 2 |
| 2013 | Versioned File Backup and Synchronization for Storage CloudsabstractCloud storage has become widely used for data backup and archiving. The current archiving systems typically support a specific cloud service and this vendor lock-in problem can cause a challenge in data migration and even data loss when a cloud service provider ceases to exist. We propose a system called Rosy Cloud for automatic backup and synchronization of client data on different clouds. Rosy Cloud uses a thin interface of storage cloud through HTTP requests, keeps all encrypted file versions on the cloud, and allows different devices to asynchronously synchronize with the cloud. Conflicts due to concurrent writes are detected and resolved using a DAG model. Rosy Cloud also supports secure file sharing among different users. We have implemented the Rosy Cloud prototype with three popular storage clouds. Experiments show that Rosy Cloud is effective for both backup and synchronization with low monetary cost. Jingyu Zhou, Tao Yang 0009 |
CCGRID | 2 |
| 2013 | Extracting URLs from JavaScript via program analysisabstractWith the extensive use of client-side JavaScript in web applications, web contents are becoming more dynamic than ever before. This poses significant challenges for search engines, because more web URLs are now embedded or hidden inside JavaScript code and most web crawlers are script-agnostic, significantly reducing the coverage of search engines. We present a hybrid approach that combines static analysis with dynamic execution, overcoming the weakness of a purely static or dynamic approach that either lacks accuracy or suffers from huge execution cost. We also propose to integrate program analysis techniques such as statement coverage and program slicing to improve the performance of URL mining. Qi Wang 0017, Jingyu Zhou, Yuting Chen 0001, Jianjun Zhao 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2013 | Fast dimension reduction for document classification based on imprecise spectrum analysis
Hu Guan, Jingyu Zhou, Bin Xiao 0001, Minyi Guo, Tao Yang 0009 |
Inf. Sci. | 2 |
| 2013 | Detecting Hot Road Mobility of Vehicular Ad Hoc Networks
Daqiang Zhang 0001, Hongyu Huang 0001, Jingyu Zhou, Feng Xia 0001, Zhe Chen 0011 |
Mob. Networks Appl. | 3 |
| 2011 | A Scalable Multiprocessor Architecture for Pervasive Computing
Long Zheng 0001, Yanchao Lu, Jingyu Zhou, Minyi Guo, Hai Jin 0001, Song Guo 0001, Jiehan Zhou, Jukka Riekki |
GPC | 3 |
| 2011 | CAB: Cache Aware Bi-tier Task-Stealing in Multi-socket Multi-core ArchitectureabstractModern multi-core computers often adopt a multi-socket multi-core architecture with shared caches in each socket. However, traditional task-stealing schedulers tend to pollute the shared cache and incur more cache misses due to their random stealing. To relieve this problem, this paper proposes a Cache Aware Bi-tier (CAB) task-stealing scheduler, which improves the performance of memory-bound applications by reducing memory footprint and cache misses of tasks running inside the same CPU socket. CAB uses an automatic partitioning method to divide an execution Directed Acyclic Graph (DAG) into the inter-socket tier and the intra-socket tier. Tasks generated in the inter-socket tier are scheduled across sockets, while tasks generated in the intra-socket tier are scheduled within the same socket. Experimental results show that CAB can improve the performance of memory-bound applications up to 68.7% compared with the traditional task-stealing. Quan Chen 0002, Zhiyi Huang 0001, Minyi Guo, Jingyu Zhou |
ICPP | 4 |
| 2011 | Preference-Based Top-K Influential Nodes Mining in Social NetworksabstractFinding top-K influential nodes in social networks has many important applications. Previous work only considered that one node in the network can influence other nodes with a uniform probability, which doesn't take user preferences into account and greatly affects the accuracy of results. We propose a two-stage mining algorithm (GAUP) for mining most influential nodes on a specific topic. In the first stage, GAUP uses a collaborative filtering technique to determine user preferences on a topic. Then in the second stage, GAUP adopts a greedy algorithm to find top-K nodes in the network. Our evaluation shows that our GAUP algorithm can successfully mine top nodes for a given topic. Jingyu Zhou |
TrustCom | 2 |
| 2011 | TASA: Tag-Free Activity Sensing Using RFID Tag ArraysabstractRadio Frequency IDentification (RFID) has attracted considerable attention in recent years for its low cost, general availability, and location sensing functionality. Most existing schemes require the tracked persons to be labeled with RFID tags. This requirement may not be satisfied for some activity sensing applications due to privacy and security concerns and uncertainty of objects to be monitored, e.g., group behavior monitoring in warehouses with privacy limitations, and abnormal customers in banks. In this paper, we propose TASA-Tag-free Activity Sensing using RFID tag Arrays for location sensing and frequent route detection. TASA relaxes the monitored objects from attaching RFID tags, online recovers and checks frequent trajectories by capturing the Received Signal Strength Indicator (RSSI) series for passive RFID tag arrays where objects traverse. In order to improve the accuracy for estimated trajectories and accelerate location sensing, TASA introduces reference tags with known positions. With the readings from reference tags, TASA can locate objects more accurately. Extensive experiment shows that TASA is an effective approach for certain activity sensing applications. Daqiang Zhang 0001, Jingyu Zhou, Minyi Guo, Jiannong Cao 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | Fast dimension reduction for document classification based on imprecise spectrum analysisabstractThis paper proposes an algorithm called Imprecise Spectrum Analysis (ISA) to carry out fast dimension reduction for document classification. ISA is designed based on the one-sided Jacobi method for Singular Value Decomposition (SVD). To speedup dimension reduction, it simplifies the orthogonalization process of Jacobi computation and introduces a new mapping formula for transforming original document-term vectors. To improve classification accuracy using ISA, a feature selection method is further developed to make inter-class feature vectors more orthogonal in building the initial weighted term-document matrix. Our experimental results show that ISA is extremely fast in handling large term-document matrices and delivers better or competitive classification accuracy compared to SVD-based LSI. Hu Guan, Bin Xiao 0001, Jingyu Zhou, Minyi Guo, Tao Yang 0009 |
CIKM | 3 |
| 2010 | Programming support and adaptive checkpointing for high-throughput data services with log-based recoveryabstractMany applications in large-scale data mining and offline processing are organized as network services, running continuously or for a long period of time. To sustain high-throughput, these services often keep their data in memory, thus susceptible to failures. On the other hand, the availability requirement for these services is not as stringent as online services exposed to millions of users. But those data-intensive offline or mining applications do require data persistence to survive failures. This paper presents programming and runtime support called SLACH for building multi-threaded high-throughput persistent services. To keep in-memory objects persistent, SLACH employs application-assisted logging and checkpointing for log-based recovery while maximizing throughput and concurrency. SLACH adaptively adjusts checkpointing frequency based on log growth and throughput demand to balance between runtime overhead and recovery speed. This paper describes the design and API of SLACH, adaptive checkpoint control, and our experiences and experiments in using SLACH at Ask.com. Jingyu Zhou, Caijie Zhang, Hong Tang 0004, Jiesheng Wu, Tao Yang 0009 |
DSN | 1 |
| 2010 | Context reasoning using extended evidence theory in pervasive computing environments
Daqiang Zhang 0001, Minyi Guo, Jingyu Zhou, Dazhou Kang, Jiannong Cao 0001 |
Future Gener. Comput. Syst. | 3 |
| 2009 | An Efficient Collaborative Filtering Approach Using Smoothing and FusingabstractCollaborative filtering (CF) has achieved widespread success in recommender systems such as Amazon and Yahoo! music. However, CF usually suffers from two fundamental problems - data sparsity and limited scalability. Among the two broad classes of CF approaches, namely, memory-based and model-based, the former usually falls short of the system scalability demands, because these approaches predict user preferences over the entire item-user matrix. The latter often achieves unsatisfactory accuracy, because they cannot capture precisely the diversity in user rating styles. In this paper, we propose an efficient collaborative filtering approach using smoothing and fusing (CFSF) strategies. CFSF formulates the CF problem as a local prediction problem by mapping it from the entire large-scale item-user matrix to a locally reduced item-user matrix. Given an active item and a user, CFSF dynamically constructs a local item-user matrix as the basis of prediction. To alleviate data sparsity, CFSF presents a fusion strategy for the local item-user matrix that fuses ratings of the same user makes on similar items, and ratings of like-minded users make on the same and similar items. To eliminate diversity in user rating styles, CFSF uses a smoothing strategy that clusters users over the entire item-user matrix and then smoothes ratings within each user cluster. Empirical study shows that CFSF outperforms the state-of-the-art CF approaches in terms of both accuracy and scalability. Daqiang Zhang 0001, Jiannong Cao 0001, Jingyu Zhou, Minyi Guo, Vaskar Raychoudhury |
ICPP | 3 |
| 2009 | A class-feature-centroid classifier for text categorizationabstractAutomated text categorization is an important technique for many web applications, such as document indexing, document filtering, and cataloging web resources. Many different approaches have been proposed for the automated text categorization problem. Among them, centroid-based approaches have the advantages of short training time and testing time due to its computational efficiency. As a result, centroid-based classifiers have been widely used in many web applications. However, the accuracy of centroid-based classifiers is inferior to SVM, mainly because centroids found during construction are far from perfect locations. Hu Guan, Jingyu Zhou, Minyi Guo |
WWW | 2 |
| 2008 | iShadow: Yet Another Pervasive Computing EnvironmentabstractPrevious architectures of pervasive computing are customized for specific types of applications. In this paper, we propose a new architecture named iShadow, which facilitates the design and implementation of generic applications in pervasive computing environment. iShadow gracefully integrates physical spaces and human attention, and provides fundamental and flexible support to construct pervasive applications rapidly. Significant differences of iShadow from previous works are lightweight user-shadow model, scalable resource discovery and potent context inference mechanism. Our prototypes demonstrate that the iShadow architecture is robust, feasible and effective for pervasive applications. Daqiang Zhang 0001, Hu Guan, Jingyu Zhou, Feilong Tang 0001, Minyi Guo |
ISPA | 3 |
| 2006 | Request-Aware Scheduling for Busy Internet ServicesabstractAbstract — Internet traffic is bursty and network servers are often overloaded with surprising events or abnormal client request patterns. This paper studies scheduling algorithms for interactive network services that use multiple threads to handle incoming requests continuously and concurrently. Our investigation with applications from Ask Jeeves search shows that during overloaded situations, requests that require excessive computing resource can dramatically affect the overall system throughput and response time. The most effective method is to manage resource usage at a request level instead of a thread or process level. We propose a new size-adaptive request-aware scheduling algorithm called SRQ with dynamic feedbacks to control queue properties and have implemented SRQ in the Linux kernel level. Our experimental results with several application service benchmarks indicate that the proposed scheduler can significantly outperform the standard Linux scheduler. I. Jingyu Zhou, Caijie Zhang, Tao Yang 0009, Lingkun Chu |
INFOCOM | 1 |
| 2006 | Selective early request termination for busy internet servicesabstractInternet traffic is bursty and network servers are often overloaded with surprising events or abnormal client request patterns. This paper studies a load shedding mechanism called selective early request termination (SERT) for network services that use threads to handle multiple incoming requests continuously and concurrently. Our investigation with applications from Ask.com shows that during overloaded situations, a relatively small percentage of long requests that require excessive computing resource can dramatically affect other short requests and reduce the overall system throughput. By actively detecting and aborting overdue long requests, services can perform significantly better to achieve QoS objectives compared to a purely admission based approach. We have proposed a termination scheme that monitors running time of requests, accounts for their resource usage, adaptively adjusts the selection threshold, and performs a safe termination for a class of requests. This paper presents the design and implementation of this scheme and describes experimental results to validate the proposed approach. Jingyu Zhou, Tao Yang 0009 |
WWW | 1 |
| 2005 | Dependency isolation for thread-based multi-tier Internet servicesabstractMulti-tier Internet service clusters often contain complex calling dependencies among service components spreading across cluster nodes. Without proper handling, partial failure or overload at one component can cause cascading performance degradation in the entire system. While dependency management may not present significant challenges for even-driven services (particularly in the context of staged event-driven architecture), there is a lack of system support for thread-based online services to achieve dependency isolation automatically. To this end, we propose dependency capsule, a new mechanism that supports automatic recognition of dependency states and per-dependency management for thread-based services. Our design employs a number of dependency capsules at each service node: one for each remote service component. Dependency capsules monitor and manage threads that block on supporting services and isolate their performance impact on the capsule host and the rest of the system. In addition to the failure and overload isolation, each capsule can also maintain dependency-specific feedback information to adjust control strategies for better availability and performance. In our implementation, dependency capsules are transparent to application-level services and clustering middleware, which is achieved by intercepting dependency-induced system calls. Additionally, we employ two-level thread management so that only light-weight user-level threads block in dependency capsules. Using four applications on two different clustering middleware platforms, we demonstrate the effectiveness of dependency capsules in improving service availability and throughput during component failures and overload. Lingkun Chu, Hong Tang 0004, Tao Yang 0009, Jingyu Zhou |
INFOCOM | 5 |
| 2004 | Detecting Attacks That Exploit Application-Logic Errors Through Application-Level AuditingabstractHost security is achieved by securing both the operating system kernel and the privileged applications that run on top of it. Application-level bugs are more frequent than kernel-level bugs, and, therefore, applications are often the means to compromise the security of a system. Detecting these attacks can be difficult, especially in the case of attacks that exploit application-logic errors. These attacks seldom exhibit characterizing patterns as in the case of buffer overflows and format string attacks. In addition, the data used by intrusion detection systems is either too low-level, as in the case of system calls, or incomplete, as in the case of syslog entries. This paper presents a technique to enforce nonbypassable, application-level auditing that does not require the recompilation of legacy systems. The technique is implemented as a kernel-level component, a privileged daemon, and an offline language tool. The technique uses binary rewriting to instrument applications so that meaningful and complete audit information can be extracted. This information is then matched against application-specific signatures to detect attacks that exploit application-logic errors. The technique has been successfully applied to detect attacks against widely-deployed applications, including the Apache Web server and the OpenSSH server. Jingyu Zhou, Giovanni Vigna |
ACSAC | 1 |
| 2004 | Parallel Simulation of Fluid Slip in a MicrochannelabstractSummary form only given. We investigate the parallel simulation of fluid slip along microchannel walls using the multicomponent lattice Boltzmann method (LBM) with domain decomposition. Because of the high complexity for microscale simulation, even a parallel computation of fluid slip can take days or weeks. Any slowness in the participating nodes in a cluster can drag the entire computation substantially, due to frequent node synchronization involved in each computational phase of the algorithm. We augment the parallel LBM algorithm with filtered dynamic remapping for lattice points. This filtered scheme uses lazy remapping and over-redistribution strategies to balance the computational speed of participating nodes and to minimize the performance impact of slow nodes on synchronized phases. Our experimental results indicate that the proposed technique can greatly speed up fluid slip simulation on a nondedicated cluster over a long period of execution time. Jingyu Zhou, Luoding Zhu, Linda R. Petzold, Tao Yang 0009 |
IPDPS | 1 |
| 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive ApplicationsabstractCluster-based storage systems are popular for data-intensive applications and it is desirable yet challenging to provide incremental expansion and high availability while achieving scalability and strong consistency. This paper presents the design and implementation of a self-organizing storage cluster called Sorrento, which targets data-intensive workload with highly parallel requests and low write-sharing patterns. Sorrento automatically adapts to storage node joins and departures, and the system can be configured and maintained incrementally without interrupting its normal operation. Data location information is distributed across storage nodes using consistent hashing and the location protocol differentiates small and large data objects for access efficiency. It adopts versioning to achieve single-file serializability and replication consistency. In this paper, we present experimental results to demonstrate features and performance of Sorrento using microbenchmarks, application benchmarks, and application trace replay. Hong Tang 0004, Aziz Gulbeden, Jingyu Zhou, William Strathearn, Tao Yang 0009, Lingkun Chu |
SC | 3 |
| 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth StorageabstractFundamental advances in high-level storage architectures and low-level storage-device interfaces greatly improve the performance and scalability of storage systems. Specifically, the decoupling of storage control (i.e., file system policy) from datapath operations (i.e., read, write) allows client applications to leverage the readily available bandwidth of storage devices while continuing to rely on the rich semantics of today’s file systems. Further, the evolution of storage interfaces from block-based devices with no protection to object-based devices with per-command access control enables storage to become secure, first-class IP-based network citizens. This paper examines how the Panasas ActiveScale Storage Cluster leverages distributed storage and object-based devices to achieve linear scalability of storage bandwidth. Specifically, we focus on implementation issues with our Object-based Storage Device, aggregation algorithms across the collection of OSDs, and the close coupling of networking and storage to achieve scalability. Hong Tang 0004, Aziz Gulbeden, Jingyu Zhou, William Strathearn, Tao Yang 0009, Lingkun Chu |
SC | 3 |
| 2002 | Composable Tools For Network Discovery and Security AnalysisabstractSecurity analysis should take advantage of a reliable knowledge base that contains semantically-rich information about a protected network. This knowledge is provided by network mapping tools. These tools rely on models to represent the entities of interest, and they leverage off network discovery techniques to populate the model structure with the data that is pertinent to a specific target network. Unfortunately, existing tools rely on incomplete data models. Networks are complex systems and most approaches oversimplify their target models in an effort to limit the problem space. In addition, the techniques used to populate the models are limited in scope and are difficult to extend. This paper presents NetMap, a security tool for network modeling, discovery, and analysis. NetMap relies on a comprehensive network model that is not limited to a specific network level; it integrates network information throughout the layers. The model contains information about topology, infrastructure, and deployed services. In addition, the relationships among different entities in different layers of the model are made explicit. The modeled information is managed by using a suite of composable network tools that can determine various aspects of network configurations through scanning techniques and heuristics. Tools in the suite are responsible for a single, well-defined task. Giovanni Vigna, Fredrik Valeur, Jingyu Zhou, Richard A. Kemmerer |
ACSAC | 3 |