VLDB 2026 Research / reviewers in the wild / expert
Jian Ni
dblp:75/2470
· DBLP profile ↗
43ranked-venue papers
26as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 21 · 13 first-authorArtificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Theory of computation · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
16 papers |
Wireless networking · 56% Network measurement and analytics · 16% Network optimization and economics · 12% | |
| Artificial intelligence
6 papers |
Information extraction and text analysis · 49% Transfer learning and domain adaptation · 16% Knowledge representation and reasoning · 16% | |
| Network and information security
2 papers |
Cryptographic primitives and cryptanalysis · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Knowledge graphs · 59% Web and social media mining · 41% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Integrated circuit design · 65% Performance modeling and evaluation · 35% | |
| Theoretical computer science
5 papers |
Algorithms and data structures · 71% Graph algorithms and graph theory · 23% Information theory · 7% |
Topics — the 30 heaviest of 73, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cryptographic primitives and cryptanalysis › public-key cryptography
modular multiplication |
0.8 | 2 | 2020 | High Performance Modular Multiplication for SIDH · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Optimized Modular Multiplication for Supersingular Isogeny Diffie-Hellman · IEEE Trans. Computers 2019 |
Cryptographic primitives and cryptanalysis
post-quantum cryptography |
0.8 | 2 | 2020 | High Performance Modular Multiplication for SIDH · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Optimized Modular Multiplication for Supersingular Isogeny Diffie-Hellman · IEEE Trans. Computers 2019 |
Cryptographic primitives and cryptanalysis › post-quantum cryptography › isogeny-based cryptography
supersingular isogeny diffie-hellman |
0.8 | 2 | 2020 | High Performance Modular Multiplication for SIDH · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Optimized Modular Multiplication for Supersingular Isogeny Diffie-Hellman · IEEE Trans. Computers 2019 |
Wireless networking
medium access control |
0.8 | 6 | 2013 | Throughput-Optimal CSMA With Imperfect Carrier Sensing · IEEE/ACM Trans. Netw. 2013 Q-CSMA: Queue-Length-Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless Networks · IEEE/ACM Trans. Netw. 2012 Fast Mixing of Parallel Glauber Dynamics and Low-Delay CSMA Scheduling · IEEE Trans. Inf. Theory 2012 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.6 | 1 | 2022 | Knowledge-Based News Event Analysis and Forecasting Toolkit · IJCAI 2022 |
Natural language and speech › Information extraction and text analysis › event analysis
event prediction |
0.6 | 1 | 2022 | Knowledge-Based News Event Analysis and Forecasting Toolkit · IJCAI 2022 |
Knowledge graphs › temporal knowledge graph
event knowledge graph |
0.6 | 1 | 2022 | Knowledge-Based News Event Analysis and Forecasting Toolkit · IJCAI 2022 |
Web and social media mining › news analysis
news event analysis |
0.6 | 1 | 2022 | Knowledge-Based News Event Analysis and Forecasting Toolkit · IJCAI 2022 |
Natural language and speech › Information extraction and text analysis › named entity recognition
cross-lingual named entity recognition |
0.5 | 2 | 2017 | Weakly Supervised Cross-Lingual Named Entity Recognition via Effective Annotation and Representation Projection · ACL (1) 2017 Improving Multilingual Named Entity Recognition with Wikipedia Entity Type Mapping · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.5 | 2 | 2017 | Weakly Supervised Cross-Lingual Named Entity Recognition via Effective Annotation and Representation Projection · ACL (1) 2017 Improving Multilingual Named Entity Recognition with Wikipedia Entity Type Mapping · EMNLP 2016 |
Cryptographic primitives and cryptanalysis › post-quantum cryptography
isogeny-based cryptography |
0.4 | 1 | 2020 | High Performance Modular Multiplication for SIDH · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Network optimization and economics
throughput-optimal scheduling |
0.4 | 3 | 2013 | Throughput-Optimal CSMA With Imperfect Carrier Sensing · IEEE/ACM Trans. Netw. 2013 Q-CSMA: Queue-Length-Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless Networks · IEEE/ACM Trans. Netw. 2012 Q-CSMA: Queue-Length Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless Networks · INFOCOM 2010 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
generalized zero-shot learning |
0.4 | 1 | 2019 | Dual Adversarial Semantics-Consistent Network for Generalized Zero-Shot Learning · NeurIPS 2019 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Dual Adversarial Semantics-Consistent Network for Generalized Zero-Shot Learning · NeurIPS 2019 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.4 | 1 | 2019 | Neural Cross-Lingual Relation Extraction Based on Bilingual Word Embedding Mapping · EMNLP/IJCNLP (1) 2019 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.4 | 1 | 2019 | Dual Adversarial Semantics-Consistent Network for Generalized Zero-Shot Learning · NeurIPS 2019 |
Network measurement and analytics › network tomography
topology inference |
0.3 | 3 | 2011 | Network Tomography Based on Additive Metrics · IEEE Trans. Inf. Theory 2011 Efficient and dynamic routing topology inference from end-to-end measurements · IEEE/ACM Trans. Netw. 2010 Network Routing Topology Inference from End-to-End Measurements · INFOCOM 2008 |
Natural language and speech › Machine translation
annotation projection |
0.3 | 1 | 2017 | Weakly Supervised Cross-Lingual Named Entity Recognition via Effective Annotation and Representation Projection · ACL (1) 2017 |
Natural language and speech › Information extraction and text analysis › named entity recognition
weakly supervised named entity recognition |
0.3 | 1 | 2017 | Weakly Supervised Cross-Lingual Named Entity Recognition via Effective Annotation and Representation Projection · ACL (1) 2017 |
Wireless networking › medium access control › channel access scheduling
CSMA scheduling |
0.3 | 2 | 2012 | Fast Mixing of Parallel Glauber Dynamics and Low-Delay CSMA Scheduling · IEEE Trans. Inf. Theory 2012 Fast mixing of parallel Glauber dynamics and low-delay CSMA scheduling · INFOCOM 2011 |
Knowledge graphs
knowledge graph construction |
0.2 | 1 | 2016 | Improving Multilingual Named Entity Recognition with Wikipedia Entity Type Mapping · EMNLP 2016 |
Integrated circuit design
digital circuit design |
0.2 | 2 | 2020 | High Performance Modular Multiplication for SIDH · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 Optimized Modular Multiplication for Supersingular Isogeny Diffie-Hellman · IEEE Trans. Computers 2019 |
Wireless networking
channel assignment |
0.2 | 2 | 2011 | Coloring spatial point processes with applications to peer discovery in large wireless networks · IEEE/ACM Trans. Netw. 2011 Coloring spatial point processes with applications to peer discovery in large wireless networks · SIGMETRICS 2010 |
Network measurement and analytics › network tomography › topology inference
routing topology inference |
0.2 | 2 | 2010 | Efficient and dynamic routing topology inference from end-to-end measurements · IEEE/ACM Trans. Netw. 2010 Network Routing Topology Inference from End-to-End Measurements · INFOCOM 2008 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
causal knowledge extraction |
0.2 | 1 | 2022 | Knowledge-Based News Event Analysis and Forecasting Toolkit · IJCAI 2022 |
Wireless networking › medium access control
carrier sense multiple access |
0.2 | 1 | 2013 | Throughput-Optimal CSMA With Imperfect Carrier Sensing · IEEE/ACM Trans. Netw. 2013 |
Wireless networking › wireless network protocols
peer discovery |
0.2 | 2 | 2011 | Coloring spatial point processes with applications to peer discovery in large wireless networks · IEEE/ACM Trans. Netw. 2011 Coloring spatial point processes with applications to peer discovery in large wireless networks · SIGMETRICS 2010 |
Wireless networking › medium access control › collision avoidance
CSMA/CA |
0.1 | 1 | 2012 | Q-CSMA: Queue-Length-Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless Networks · IEEE/ACM Trans. Netw. 2012 |
Network performance modeling › markov chain model
mixing time |
0.1 | 1 | 2012 | Fast Mixing of Parallel Glauber Dynamics and Low-Delay CSMA Scheduling · IEEE Trans. Inf. Theory 2012 |
Wireless networking › scheduling › queueing discipline
queue-length-based scheduling |
0.1 | 1 | 2012 | Q-CSMA: Queue-Length-Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless Networks · IEEE/ACM Trans. Netw. 2012 |
Methods — techniques the papers use, named apart from their topics
neuro-symbolic techniques · 1.1knowledge graph reasoning · 1.1modular multiplication algorithm · 0.9interleaved hardware architecture · 0.9hardware-software co-design · 0.8glauber dynamics · 0.6weakly supervised learning · 0.5visual-semantic embedding · 0.4semantics-consistent adversarial learning · 0.4neural network · 0.4bilingual word embedding mapping · 0.4greedy coloring · 0.4word embedding projection · 0.3co-decoding · 0.3annotation projection · 0.3wikipedia mining · 0.2perturbation theory of markov chains · 0.2mixing time analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment RetrievalabstractAccurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LVMR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1,200 seconds in duration, and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, and object-level, covering common tasks like action recognition, object localization, causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker to facilitate future research in this area. Huaying Yuan, Jian Ni, Zheng Liu 0011, Yueze Wang, Junjie Zhou 0001, Zhengyang Liang, Bo Zhao 0015, Zhao Cao, Ji-Rong Wen, Zhicheng Dou |
NeurIPS | 2 |
| 2025 | DRCL: rethinking jigsaw puzzles for unsupervised medical image segmentation
Jian Ni, Wenjian Tao |
Vis. Comput. | 1 |
| 2024 | Distilling Event Sequence Knowledge From Large Language Models
Somin Wadhwa, Oktie Hassanzadeh, Debarun Bhattacharjya, Ken Barker 0002, Jian Ni |
ISWC (1) | 5 |
| 2024 | Identification and Evolutionary Analysis of User Collusion Behavior in Blockchain Online Social MediaabstractBlockchain technology has given rise to a series of new blockchain online social media (BOSMs), of which Steemit is representative. Such communities are based on a token reward system and attempt to engross users in the knowledge activities of the community through knowledge payment. Studies have found that the reward system of such communities has been abused (e.g., collusion for profit), but few studies have performed an in-depth analysis for this phenomenon. Consequently, real data for Steemit are used as a case study herein to examine the collusion of users in BOSMs. Two user collusion behaviors (group-voting and vote-buying) are defined and measured. On this basis, an identification and evolutionary survival analysis of the two collusion behaviors are conducted for colluding users and colluding groups, and the behavior patterns of user collusion under the token system are deconstructed. The results of this study improve stakeholders’ understanding of user participation behavior in new online communities, and serve as a reference for decision-making in community governance and token design. Hongting Tang, Jian Ni, Yanlin Zhang |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Semantics-Disentangled Contrastive Embedding for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) is a challenging class of vision and knowledge transfer problems in which both seen and unseen classes appear during testing. Most existing GZSL methods achieve knowledge transfer based on the original features of samples that inevitably contain information irrelevant to recognition, resulting in negative influence for the performance. In this paper, we propose a novel contrastive disentanglement learning framework for the GZSL task (SDCE-GZSL), where the original and generated visual features are factorized into semantic-consistent and semantic-unrelated representations via a novel mutual information (MI)-based constraint. In addition, we propose a contrastive learning framework that leverages class-level and instance-level supervision to further facilitate disentanglement. Extensive experiments show that our approach achieves significant improvements over the state-of-the-art approaches. Jian Ni |
ICASSP | 1 |
| 2023 | Structure-Preserving and Redundancy-Free Features Refinement for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) aims to recognize images from seen and unseen classes. Most models achieve competitive performance but still suffer from two problems: (1) Topological structure neglection; (2) Redundant information interference. In this paper, we propose a Structure-preserving and Redundancy-free Features Refinement model (referred to as SP-RFFR) to address these problems correspondingly in two modules: (1) Structure-preserving, to explicitly incorporate the topological structure into the learning of the latent space and the generator; (2) Redundancy-free features refinement, to remove the redundant information from the visual features and learn class- and semantically-relevant representations. To the best of our knowledge, this is the first work that incorporates topological structure preserving and redundancy-free features refinement into a unified framework for GZSL. Extensive experiments show that SP-RFFR outperforms the state-of-the-art methods on four benchmarks. Jian Ni |
ICASSP | 1 |
| 2022 | Knowledge-Based News Event Analysis and Forecasting ToolkitabstractWe present a toolkit for knowledge-based news event analysis and forecasting. The toolkit is powered by a Knowledge Graph (KG) of events curated from structured and unstructured sources of event-related knowledge. The toolkit provides functions for 1) mapping ongoing news headlines to concepts in the KG, 2) retrieval, reasoning, and visualization for causal analysis and forecasting, and 3) extraction of causal knowledge from text documents to augment the KG with additional domain knowledge. Each function has a number of implementations using a wide range of state-of-the-art neuro-symbolic techniques. We show how the toolkit enables building a human-in-the-loop explainable solution for event analysis and forecasting. Oktie Hassanzadeh, Parul Awasthy, Ken Barker 0002, Onkar Bhardwaj, Debarun Bhattacharjya, Mark Feblowitz, Lee Martie, Jian Ni, Kavitha Srinivas, Lucy Yip |
IJCAI | 8 |
| 2022 | Copper price movement prediction using recurrent neural networks and ensemble averaging
Jian Ni, Zhi Li 0053 |
Soft Comput. | 1 |
| 2021 | Multi-class Cardiovascular Disease Detection and Classification from 12-Lead ECG Signals Using an Inception Residual NetworkabstractA number of deep neural network (DNN)-based models have been applied to help classify and detect severe cardiovascular diseases using 12-lead electrocardiogram (ECG) signals. These models, however, suffer from poor performance in detecting one or two specific cardiac abnormalities, like ST-segment abnormalities, with their accuracy lingering at only around 60%, which in turn limits their applicability in clinical practice. In this paper, we show that the convolution layers of these DNN models can cause the diminishment of the key features of ST-segment abnormalities, making it, compared to cardiac arrhythmia and normal sinus rhythm (NSR), hard to be classified out from the ECG. Correspondingly, we, in this paper, propose a novel DNN-based model that is able to achieve high accuracy of detecting 9 classes of rhythms. In specific, the moving averages of the ECG signals are used as the second input to the proposed model so that it can help fully mine the features relevant to the ST abnormalities. In addition, as opposed to the convolution layers, our model utilizes the inception-residual layers to preserve features from the shallow layers for reuse in the deep layers. On top of these network architecture improvements, we further introduce a customized differential layer so that all the relevant features can be preserved and/or amplified for classification purposes. Trained with CPSC 2018 dataset, our proposed model is able to accurately classify the 9 rhythm classes, with an F1 score as high as 89.7%, up by merely 4.2% from the best result reported in the literature. As far as the ST segment abnormalities are concerned, the proposed model achieves an 89.6% F1 score for STD and an 80.8% F1 score for STE, which is 7.8% and 13.1% respectively higher than that of the best result. The proposed method is thus poised to become a viable solution for cardiovascular health monitoring with the increasing availability of portable and home-based ECG devices. Jian Ni, Yingtao Jiang, Shengjie Zhai, Amei Amei, Dieu-My. T. Tran, Lijie Zhai, Yu Kuang |
COMPSAC | 1 |
| 2020 | High Performance Modular Multiplication for SIDHabstractThe latest research indicates that quantum computers will be realized in the near future. In theory, the computation speed of a quantum computer is much faster than current computers, which will pose a serious threat to current cryptosystems. Post-quantum cryptography (PQC) is a class of cryptography based on underlying mathematical problems that are considered infeasible to crack even with access to a quantum computer. The supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol is a new post-quantum cryptosystem, which offers advantages in reduced secret key length and attack resistance. SIDH is the basis of the supersingular isogeny key encapsulation (SIKE) protocol, which is in the second round of the U.S. National Institute of Standards and Technology (NIST) PQC standardization process. In this article, we propose a new modular multiplication algorithm and a new interleaved hardware architecture for SIDH. Performance results for the proposed modular multiplier using four parameter sets for the prime, p that correspond to the SIKE Round 2 parameter sets show significant advantages in speed. Weiqiang Liu 0001, Ziying Ni, Jian Ni, Ciara Rafferty, Máire O'Neill |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Neural Cross-Lingual Relation Extraction Based on Bilingual Word Embedding MappingabstractJian Ni, Radu Florian. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jian Ni, Radu Florian |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Dual Adversarial Semantics-Consistent Network for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) is a challenging class of vision and knowledge transfer problems in which both seen and unseen classes appear during testing. Existing GZSL approaches either suffer from semantic loss and discard discriminative information at the embedding stage, or cannot guarantee the visual-semantic interactions. To address these limitations, we propose a Dual Adversarial Semantics-Consistent Network (referred to as DASCN), which learns both primal and dual Generative Adversarial Networks (GANs) in a unified framework for GZSL. In DASCN, the primal GAN learns to synthesize inter-class discriminative and semantics-preserving visual features from both the semantic representations of seen/unseen classes and the ones reconstructed by the dual GAN. The dual GAN enforces the synthetic visual features to represent prior semantic knowledge well via semantics-consistent adversarial learning. To the best of our knowledge, this is the first work that employs a novel dual-GAN mechanism for GZSL. Extensive experiments show that our approach achieves significant improvements over the state-of-the-art approaches. Jian Ni, Shanghang Zhang, Haiyong Xie 0001 |
NeurIPS | 1 |
| 2019 | Optimized Modular Multiplication for Supersingular Isogeny Diffie-HellmanabstractRecent progress in quantum physics shows that quantum computers may be a reality in the not too distant future. Post-quantum cryptography (PQC) refers to cryptographic schemes that are based on hard problems which are believed to be resistant to attacks from quantum computers. The supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol shows promising security properties among various post-quantum cryptosystems that have been proposed. In this paper, we propose two efficient modular multiplication algorithms with special primes that can be used in SIDH key exchange protocol. Hardware architectures for the two proposed algorithms are also proposed. The hardware implementations are provided and compared with the original modular multiplication algorithm. The results show that the proposed finite field multiplier is over 6.79 times faster than the original multiplier in hardware. Moreover, the SIDH hardware/software codesign implementation using the proposed FFM2 hardware is over 31 percent faster than the best SIDH software implementation. Weiqiang Liu 0001, Jian Ni, Zhe Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 2 |
| 2018 | Design and Optimization of Modular Multiplication for SIDHabstractRecent progress on quantum physics shows that quantum computers may be a reality in the not too distant future. Based on new mathematical hard problems, post-quantum cryptography (PQC) has been studied to make sure the attacks from quantum computers can be resistant. The latest supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol shows promising security properties among various post-quantum cryptosystems. In this paper, we propose an improved modular multiplication algorithm with special primes that can be used in SIDH key exchange protocol. Both software and hardware implementations are provided and compared with original modular multiplication algorithm. The results show that the software results of improved algorithm can be 24% faster than the original software implementation, while the hardware implementation based on the proposed hardware architecture can be 6 times faster than previous hardware implementation. Jian Ni, Weiqiang Liu 0001, Zhe Liu 0001, Máire O'Neill |
ISCAS | 2 |
| 2018 | Whether Android Applications Broadcast Your Private information: A Naive Bayesian-based Analysis Approach (S)abstractWith the rapid development of android smart terminals, android applications are exhibiting explosive growth.However, there remains a challenging issue facing android system, a malicious application may broadcast user's private information.In this paper, we propose a Naive Bayesian-based approach for analyzing private information leakage under the android broadcast mechanism, which calls BRbysA.Firstly, broadcast actions registered in manifest.xmlare picked up statically by keyword matching technique.Secondly, with the Xposed framework, the broadcast actions specified at run time are discovered by hooking broadcast callback onReceive() function.Combining the above two ways, we can capture all real-time broadcast actions in an android application.Thirdly, we adopt Naive Bayesian learning algorithm, all broadcast actions which involved in users' privacy leakage are analyzed and classified.Finally, we evaluate the proposed approach by using the dataset from Drebin and Google Play. Li Lin 0008, Jian Ni, Xinya Mao, Jian-biao Zhang |
SEKE | 2 |
| 2017 | Weakly Supervised Cross-Lingual Named Entity Recognition via Effective Annotation and Representation ProjectionabstractThe state-of-the-art named entity recognition (NER) systems are supervised machine learning models that require large amounts of manually annotated data to achieve high accuracy.However, annotating NER data by human is expensive and time-consuming, and can be quite difficult for a new language.In this paper, we present two weakly supervised approaches for cross-lingual NER with no human annotation in a target language.The first approach is to create automatically labeled NER data for a target language via annotation projection on comparable corpora, where we develop a heuristic scheme that effectively selects goodquality projection-labeled data from noisy data.The second approach is to project distributed representations of words (word embeddings) from a target language to a source language, so that the sourcelanguage NER system can be applied to the target language without re-training.We also design two co-decoding schemes that effectively combine the outputs of the two projection-based approaches.We evaluate the performance of the proposed approaches on both in-house and open NER data for several target languages.The results show that the combined systems outperform three other weakly supervised approaches on the CoNLL data. Jian Ni, Georgiana Dinu, Radu Florian |
ACL (1) | 1 |
| 2016 | Improving Multilingual Named Entity Recognition with Wikipedia Entity Type MappingabstractThe state-of-the-art named entity recognition (NER) systems are statistical machine learning models that have strong generalization capability (i.e., can recognize unseen entities that do not appear in training data) based on lexical and contextual information.However, such a model could still make mistakes if its features favor a wrong entity type.In this paper, we utilize Wikipedia as an open knowledge base to improve multilingual NER systems.Central to our approach is the construction of high-accuracy, highcoverage multilingual Wikipedia entity type mappings.These mappings are built from weakly annotated data and can be extended to new languages with no human annotation or language-dependent knowledge involved.Based on these mappings, we develop several approaches to improve an NER system.We evaluate the performance of the approaches via experiments on NER systems trained for 6 languages.Experimental results show that the proposed approaches are effective in improving the accuracy of such systems on unseen entities, especially when a system is applied to a new domain or it is trained with little training data (up to 18.3 F 1 score improvement). Jian Ni, Radu Florian |
EMNLP | 1 |
| 2013 | A statistical machine learning approach for ticket mining in IT service delivery
Ea-Ee Jan, Jian Ni, Niyu Ge, Naga Ayachitula, Xiaolan Zhang 0001 |
IM | 2 |
| 2013 | Throughput-Optimal CSMA With Imperfect Carrier SensingabstractRecently, it has been shown that a simple, distributed backlog-based carrier-sense multiple access (CSMA) algorithm is throughput-optimal. However, throughput optimality is established under the perfect or ideal carrier-sensing assumption, i.e., each link can precisely sense the presence of other active links in its neighborhood. In this paper, we investigate the achievable throughput of the CSMA algorithm under imperfect carrier sensing. Through the analysis on both false positive and negative carrier sensing failures, we show that CSMA can achieve an arbitrary fraction of the capacity region if certain access probabilities are set appropriately. To establish this result, we use the perturbation theory of Markov chains. Tae Hyun Kim 0001, Jian Ni, R. Srikant 0001, Nitin H. Vaidya |
IEEE/ACM Trans. Netw. | 2 |
| 2012 | Fast Mixing of Parallel Glauber Dynamics and Low-Delay CSMA SchedulingabstractGlauber dynamics is a powerful tool to generate randomized, approximate solutions to combinatorially difficult problems. It has been recently used to design distributed carrier-sense multiple-access (CSMA) scheduling algorithms for multihop wireless networks. In this paper, we derive bounds on the mixing time of a generalization of Glauber dynamics where multiple links update their states in parallel and the fugacity of each link can be different. The results are used to prove that the average queue length (and hence, the delay) under the parallel-Glauber-dynamics-based CSMA grows polynomially in the number of links for wireless networks with bounded-degree interference graphs when the arrival rate lies in a fraction of the capacity region. Other versions of adaptive CSMA can be analyzed similarly. We also show that in specific network topologies, the low-delay capacity region can be further improved. Libin Jiang, Mathieu Leconte, Jian Ni, R. Srikant 0001, Jean C. Walrand |
IEEE Trans. Inf. Theory | 3 |
| 2012 | Q-CSMA: Queue-Length-Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless NetworksabstractRecently, it has been shown that carrier-sense multiple access (CSMA)-type random access algorithms can achieve the maximum possible throughput in ad hoc wireless networks. However, these algorithms assume an idealized continuous-time CSMA protocol where collisions can never occur. In addition, simulation results indicate that the delay performance of these algorithms can be quite bad. On the other hand, although some simple heuristics (such as greedy maximal scheduling) can yield much better delay performance for a large set of arrival rates, in general they may only achieve a fraction of the capacity region. In this paper, we propose a discrete-time version of the CSMA algorithm. Central to our results is a discrete-time distributed randomized algorithm that is based on a generalization of the so-called Glauber dynamics from statistical physics, where multiple links are allowed to update their states in a single timeslot. The algorithm generates collision-free transmission schedules while explicitly taking collisions into account during the control phase of the protocol, thus relaxing the perfect CSMA assumption. More importantly, the algorithm allows us to incorporate heuristics that lead to very good delay performance while retaining the throughput-optimality property. Jian Ni, Bo Tan 0002, R. Srikant 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | Fast mixing of parallel Glauber dynamics and low-delay CSMA schedulingabstractGlauber dynamics is a powerful tool to generate randomized, approximate solutions to combinatorially difficult problems. It has been recently used to design distributed CSMA scheduling algorithms for multi-hop wireless networks. In this paper, we derive bounds on the mixing time of a generalization of Glauber dynamics where multiple links update their states in parallel and the fugacity of each link can be different. The results are used to prove that the average queue length (and hence, the delay) under the parallel-Glauber-dynamics-based CSMA grows polynomially in the number of links for wireless networks with bounded-degree interference graphs when the arrival rate lies in a fraction of the capacity region. Other versions of adaptive CSMA can be analyzed similarly. Libin Jiang, Mathieu Leconte, Jian Ni, R. Srikant 0001, Jean C. Walrand |
INFOCOM | 3 |
| 2011 | On the achievable throughput of CSMA under imperfect carrier sensingabstractRecently, it has been shown that a simple, distributed CSMA algorithm can achieve throughput-optimality. However, the optimality is established under the ideal carrier sensing assumption, i.e., each link can precisely sense the presence of other active links in its neighborhood. This paper, in contrast, investigates the achievable throughput of the CSMA algorithm under imperfect carrier sensing. The main result is that CSMA can achieve an arbitrary fraction of the capacity region if certain access probabilities are set appropriately. To establish this result, we use the perturbation theory of Markov chains. Tae Hyun Kim 0001, Jian Ni, R. Srikant 0001, Nitin H. Vaidya |
INFOCOM | 2 |
| 2011 | Network Tomography Based on Additive MetricsabstractNetwork tomography studies the inference of network structure and dynamics based on indirect measurements when direct measurements are unavailable or difficult to collect. In this paper, we design and analyze routing tree topology and link performance inference algorithms for communication networks using tools from phylogenetic inference in evolutionary biology. We develop polynomial-time distance-based inference algorithms and derive sufficient conditions for the correctness of the algorithms. We show that the algorithms are consistent and robust. In particular, the algorithms achieve the optimall∞-radius 1/2 for binary trees and 1/4 for general trees when a threshold neighbor selection criterion is used. Jian Ni, Sekhar Tatikonda |
IEEE Trans. Inf. Theory | 1 |
| 2011 | Improved bounds on the throughput efficiency of greedy maximal scheduling in wireless networksabstractIn this paper, we derive new bounds on the throughput efficiency of Greedy Maximal Scheduling (GMS) for wireless networks of arbitrary topology under the generalk-hop interference model. These results improve the known bounds for networks with up to 26 nodes under the 2-hop interference model. We also prove that GMS is throughput-optimal in small networks. In particular, we show that GMS achieves 100% throughput in networks with up to eight nodes under the 2-hop interference model. Furthermore, we provide a simple proof to show that GMS can be implemented using only local neighborhood information in networks of any size. Mathieu Leconte, Jian Ni, R. Srikant 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2011 | Coloring spatial point processes with applications to peer discovery in large wireless networksabstractIn this paper, we study distributed channel assignment in wireless networks with applications to peer discovery in ad hoc wireless networks. We model channel assignment as a coloring problem for spatial point processes in whichnnodes are located in a unit cube uniformly at random and each node is assigned one ofKcolors, where each color represents a channel. The objective is to maximize the spatial separation between nodes of the same color. In general, it is hard to derive the optimal coloring algorithm, and we therefore consider a natural online greedy coloring algorithm first proposed by Ko and Rubenstein in 2005. We prove two key results: 1) with just logn/log logncolors, the distance separation achieved by the greedy coloring algorithm asymptotically matches the optimal distance separation that can be achieved by an algorithm which is allowed to optimally place the nodes but is allowed to use only one color; and 2) when K=Ω(log n), the greedy coloring algorithm asymptotically achieves the best distance separation that can be achieved by an algorithm which is allowed to both optimally color and place nodes. The greedy coloring algorithm is also shown to dramatically outperform a simple random coloring algorithm. Moreover, the results continue to hold under node mobility. Jian Ni, R. Srikant 0001, Xinzhou Wu |
IEEE/ACM Trans. Netw. | 1 |
| 2010 | Q-CSMA: Queue-Length Based CSMA/CA Algorithms for Achieving Maximum Throughput and Low Delay in Wireless NetworksabstractRecently, it has been shown that CSMA-type random access algorithms can achieve the maximum possible throughput in ad hoc wireless networks. However, these algorithms assume an idealized continuous-time CSMA protocol where collisions can never occur. In addition, simulation results indicate that the delay performance of these algorithms can be quite bad. On the other hand, although some simple heuristics (such as distributed approximations of greedy maximal scheduling) can yield much better delay performance for a large set of arrival rates, they may only achieve a fraction of the capacity region in general. In this paper, we propose a discrete-time version of the CSMA algorithm. Central to our results is a discrete-time distributed randomized algorithm which is based on a generalization of the so-called Glauber dynamics from statistical physics, where multiple links are allowed to update their states in a single time slot. The algorithm generates collision-free transmission schedules while explicitly taking collisions into account during the control phase of the protocol, thus relaxing the perfect CSMA assumption. More importantly, the algorithm allows us to incorporate delay-reduction mechanisms which lead to very good delay performance while retaining the throughput-optimality property. Jian Ni, Bo Tan 0002, R. Srikant 0001 |
INFOCOM | 1 |
| 2010 | Coloring spatial point processes with applications to peer discovery in large wireless networksabstractIn this paper, we study distributed channel assignment in wireless networks with applications to peer discovery in ad hoc wireless networks. We model channel assignment as a coloring problem for spatial point processes in which n nodes are located in a unit cube uniformly at random and each node is assigned one of K colors, where each color represents a channel. The objective is to maximize the spatial separation between nodes of the same color. In general, it is hard to derive the optimal coloring algorithm and therefore, we consider a natural greedy coloring algorithm, first proposed in [5]. We prove two key results: (i) with just a small number of colors when K is roughly of the order of log(n) loglog(n), the distance separation achieved by the greedy coloring algorithm asymptotically matches the optimal distance separation that can be achieved by an algorithm which is allowed to select the locations of the nodes but is allowed to use only one color, and (ii) when K = Omega(log(n)), the greedy coloring algorithm asymptotically achieves the best distance separation that can be achieved by an algorithm which is allowed to both optimally color and place nodes. The greedy coloring algorithm is also shown to dramatically outperform a simple random coloring algorithm. Moreover, the results continue to hold under node mobilities. Jian Ni, R. Srikant 0001, Xinzhou Wu |
SIGMETRICS | 1 |
| 2010 | Efficient and dynamic routing topology inference from end-to-end measurements
Jian Ni, Haiyong Xie 0001, Sekhar Tatikonda, Yang Richard Yang |
IEEE/ACM Trans. Netw. | 1 |
| 2009 | Improved bounds on the throughput efficiency of greedy maximal scheduling in wireless networksabstractDue to its low complexity, Greedy Maximal Scheduling (GMS), also known as Longest Queue First (LQF), has been studied extensively for wireless networks. However, GMS can result in degraded throughput performance in general wireless networks. In this paper, we prove that GMS achieves 100% throughput in all networks with eight nodes or less, under the two-hop interference model. Further, we obtain performance bounds that improve upon previous results for larger networks up to a certain size. We also provide a simple proof to show that GMS can be implemented using only local neighborhood information in networks of any size. Mathieu Leconte, Jian Ni, R. Srikant 0001 |
MobiHoc | 2 |
| 2008 | Packet doppler: network monitoring using packet shift detectionabstractDue to recent large-scale deployments of delay and loss-sensitive applications, there are increasingly stringent demands on the monitoring of service level agreement metrics. Although many end-to-end monitoring methods have been proposed, they are mainly based on active probing and thus inject measurement traffic into the network. In this paper, we propose a new scheme for monitoring service level agreement metrics, in particular, delay distribution. Our scheme is passive and therefore will not cause perturbation to real traffic. Using realistic delay and traffic demands, we show that our scheme achieves high accuracy and can detect burst events that will be missed by probing based methods. Tongqing Qiu, Jian Ni, Hao Wang 0010, Nan Hua, Yang Richard Yang, Jun (Jim) Xu |
CoNEXT | 2 |
| 2008 | Designing File Replication Schemes for Peer-to-Peer File Sharing SystemsabstractPeer-to-peer (P2P) file sharing systems are becoming increasingly popular due to their flexibility and scalability. We propose a new model to design file replication schemes for P2P file sharing systems. The model introduces expected costs for serving user requests for the files which are computed from node up/down statistics. Based on the model we introduce and develop several methods to determine the sets of nodes to store copies of the files in order to optimize certain performance metrics (e.g., maximize the system hit rate, minimize the total expected cost). We verify the effectiveness of the file replication schemes via simulation. We also outline a framework to implement the file replication schemes for P2P file sharing systems in a distributed and adaptive manner. The framework scales to a large number of nodes and files and can handle user request pattern change via file migration. Jian Ni, S. J. Harrington, Naveen Sharma |
ICC | 1 |
| 2008 | Network Routing Topology Inference from End-to-End MeasurementsabstractInference of the routing topology and link performance from a node to a set of other nodes is an important component of network monitoring and application design. In this paper we propose a general framework for designing topology inference algorithms based on additive metrics. Our framework allows the integration of both end-to-end packet probing measurements and traceroute type measurements. Based on this framework we design several computationally efficient topology inference algorithms. In particular, we propose a novel sequential topology inference algorithm to address the probing scalability problem and handle dynamic node joining and leaving. We provide sufficient conditions for the correctness of our algorithms and derive lower bounds on the probability of correct topology inference. We conduct Internet experiments to evaluate and demonstrate the effectiveness of our algorithms. Jian Ni, Haiyong Xie 0001, Sekhar Tatikonda, Yang Richard Yang |
INFOCOM | 1 |
| 2007 | Performance Evaluation of Loss Networks via Factor Graphs and the Sum-Product AlgorithmabstractLoss networks provide a powerful tool for the analysis and design of many communication and networking systems. It is well known that a large number of loss networks have product-form steady-state probabilities. However, for most networks of practical interest, evaluating the system performance is a difficult task due to the presence of a normalization constant. In this paper, we present a new framework based on probabilistic graphical models to tackle this task. Specifically, we propose to use factor graphs to model the stationary distribution of a network. Based on the factor graph model, we can easily derive recursive formulas for symmetric networks. Most importantly, for networks with arbitrary topology, we can apply efficient message-passing algorithms like the sum-product algorithm to compute the exact or approximate marginal distributions of all state variables and the related performance measures such as call blocking probabilities. Through extensive numerical experiments, we show that the sum-product algorithm returns very accurate blocking probabilities and greatly outperforms the reduced load approximation for both single-service and multiservice loss networks with a variety of topologies. Jian Ni, Sekhar Tatikonda |
INFOCOM | 1 |
| 2007 | Analyzing Product-Form Stochastic Networks Via Factor Graphs and the Sum-Product AlgorithmabstractA large number of stochastic networks including loss networks and certain queueing networks have product-form steady-state probabilities. However, for most practical networks, evaluating the system performance is a difficult task due to the presence of a normalization constant. We propose a new framework based on probabilistic graphical models to tackle this task. Specifically, we use factor graphs to model the stationary distribution of a network. For networks with arbitrary topology, we can apply efficient message-passing algorithms like the sum-product algorithm to compute the exact or approximate marginal distributions of all state variables and related performance measures such as blocking probabilities. Through extensive numerical experiments, we show that the sum-product algorithm returns very accurate blocking probabilities and greatly outperforms the reduced load approximation for loss networks with a variety of topologies. The factor graph model also provides a promising approach for analyzing product-form queueing networks. Jian Ni, Sekhar Tatikonda |
IEEE Trans. Commun. | 1 |
| 2007 | Optimal and Structured Call Admission Control Policies for Resource-Sharing SystemsabstractMany communication and networking systems can be modeled as resource-sharing systems with multiple classes of calls. Call admission control (CAC) is an essential component of such systems. Markov decision process (MDP) tools can be applied to analyze and compute the optimal CAC policy that optimizes certain performance metrics of the system. But for most practical systems, it is prohibitively difficult to compute the optimal CAC policy using any MDP algorithm because of the "curse of dimensionality". We are, therefore, motivated to consider two families of structured CAC policies: reservation and threshold policies. These policies are easy to implement and have good performance in practice. However, since the number of structured policies grows exponentially with the number of call classes and the capacity of the system, finding the optimal structured policy is a complex unsolved problem. In this paper, we develop fast and efficient search algorithms to determine the parameters of the structured policies. We prove the convergence of the algorithms. Through extensive numerical experiments, we show that the search algorithms converge quickly and work for systems with large capacity and many call classes. In addition, the returned structured policies have optimal or near-optimal performance, and outperform those structured policies with parameters chosen based on simple heuristics Jian Ni, Danny H. K. Tsang, Sekhar Tatikonda, Brahim Bensaou |
IEEE Trans. Commun. | 1 |
| 2006 | A Large-Scale Distributed Traffic Matrix Estimation AlgorithmabstractAs today's communication networks (e.g., the Internet) grow in size and diversity, accurate, large-scale, and distributed traffic matrix estimation techniques will become increasingly important for many network control and management tasks. In this paper, we study the gravity model with entropy penalization approach for estimating traffic matrices based on link traffic measurements, which is known to have remarkable accuracy for real networks [7], [12]. We propose a dual approach to convert the constrained primal optimization problem under the gravity model into an unconstrained dual optimization problem. For most practical networks in which the number of links is much smaller than the number of origin-destination pairs, the dual problem has a much smaller dimension and hence scales for large networks. In addition, the solution algorithm for the dual problem can be implemented in a distributed manner. Jian Ni, Sekhar Tatikonda, Edmund M. Yeh |
GLOBECOM | 1 |
| 2006 | A Markov Random Field Approach to Multicast-Based Network Inference ProblemsabstractIn this paper, we provide a new unified approach to analyze and solve multicast-based network inference problems. We show that the outcome variables induced by the transmission of a multicast packet form a Markov random field on the multicast tree. We present an algorithm that recovers the multicast tree topology based on the values of an additive tree metric on pairs of the terminal nodes. We prove the correctness of the algorithm. We also give several examples of an additive tree metric for which the values on pairs of the terminal nodes can be estimated from traffic measurements taken at the receivers. In addition, we propose an algorithm to recover the link performance parameters from the joint distribution of the outcome variables at the terminal nodes Jian Ni, Sekhar Tatikonda |
ISIT | 1 |
| 2005 | Revenue optimization via call admission control and pricing for mobile cellular systemsabstractService differentiation is now becoming a more and more desirable feature in mobile cellular systems making the case for pricing differentiation among the users. One of the main objectives of a mobile service provider is to maximize its revenue. In general there are two approaches to optimize the revenue rate of a base station. One is through call admission control (CAC), the other is through pricing. For most real systems it is prohibitively difficult to compute the optimal CAC policy or pricing scheme because of the 'curse of dimensionality'. Therefore we focus on CAC policies that have simple structures and static pricing schemes in which the charging rates are independent of the system state. However, finding the optimal structured CAC policies or static pricing scheme is still combinatorial in nature. In this paper we show that our recently proposed iterative coordinate search algorithm provides an efficient and effective method to find the optimal or near-optimal structured CAC policies and static pricing schemes for mobile cellular systems. Jian Ni, Sekhar Tatikonda |
ICC | 1 |
| 2005 | Threshold and reservation based call admission control policies for multiservice resource-sharing systemsabstractMany communications and networking systems can be modelled as resource-sharing systems with multiple classes of calls. Call admission control (CAC) is an essential component of such systems. For most practical systems it is prohibitively difficult to compute the optimal CAC policy that optimizes certain performance metrics because of the 'curse of dimensionality'. In this paper we study two families of structured CAC policies: threshold and reservation policies. These policies are easy to implement and have good performance in practice. However, since the number of structured policies grows exponentially with the number of call classes and the capacity of the system, finding the optimal structured policies is a complex unsolved problem. In this paper efficient search algorithms are proposed to find the coordinate optimal structured policies among all structured policies. Through extensive numerical experiments we show that the search algorithms converge quickly and work for systems with large capacity and many call classes. In addition, the returned structured policies have optimal or near-optimal performance, and outperform those structured policies with parameters chosen based on simple heuristics. Jian Ni, Danny H. K. Tsang, Sekhar Tatikonda, Brahim Bensaou |
INFOCOM | 1 |
| 2005 | Calculating blocking probabilities for loss networks based on probabilistic graphical modelsabstractLoss networks are a class of resource-sharing models which provide a powerful tool to the analysis and design of many communications and networking systems. For most loss networks of practical interest calculating the exact blocking probabilities is a difficult task. In this paper we present a new framework based on probabilistic graphical models to tackle this task. Specifically, we propose to use factor graphs to model the stationary distribution of product-form loss networks. We also propose to use the sum-product algorithm to compute the marginal distributions and the blocking probabilities of all call classes. Through extensive numerical experiments we show that the sum-product algorithm returns very accurate blocking probabilities and greatly outperforms the reduced load approximation for both single-service and multiservice loss networks with a variety of topologies. In addition, the sum-product algorithm converges very fast and can be implemented in a distributed way Jian Ni, Sekhar Tatikonda |
ISIT | 1 |
| 2003 | Hierarchical content routing in large-scale multimedia content delivery networkabstractContent delivery network (CDN) is an intermediate layer of infrastructure that helps to efficiently deliver the ever increasing multimedia content from content providers to a large community of geographically distributed clients. Content routing is an essential component of CDN architecture. In this paper we propose a hierarchical content routing architecture for large-scale CDN, in which CDN servers perform inter-cluster content routing based on two-level hierarchical overlay network. We analyze the routing overhead and the corresponding CDN performance of different intra-cluster content routing schemes. In particular, we propose a semi-hashing based scheme for intra-cluster content routing and a content-query based scheme for inter-cluster content routing. Through qualitative analysis and simulations we show that the semi-hashing based scheme is scalable (small routing overhead), efficient (high content sharing efficiency), and flexible (adjustable parameters). Jian Ni, Danny H. K. Tsang, S.-H. Ivan Yeung, Xiaojun Hei |
ICC | 1 |
| 1995 | Traffic Smoothing and Bandwidth Allocation for VBR MPEG-2 Video Connections in ATM NetworksabstractWe studied the traffic smoothing and bandwidth allocation issues for VBR MPEG-2 video connections in an ATM network. First of all, the statistical characteristics of VBR MPEG-2 video sequences were examined by real observation. The obtained results showed that it would be very difficult to efficiently allocate an appropriate amount of bandwidth if such traffic were straightforwardly feed into the network. We therefore introduced a so called macro-frame (M-frame) smoothing scheme. It was found that the smoothed traffic not only possesses relatively simple characteristics but also makes the bandwidth allocation more efficient. In consideration of the fact that there is not yet a proper bandwidth allocation algorithm for VBR MPEG-2 connections, we examined and compared the simple peak rate allocation and Gaussian approximation. It was found that (1) peak rate allocation is far more inefficient than the actual need and (2) Gaussian approximation is surprisingly accurate and efficient in estimating the bandwidth, especially in the case of smoothed traffic. Zhongyong Guo, Jian Ni |
ICCCN | 2 |