VLDB 2026 Research / reviewers in the wild / expert
Shihao Ji 0001
dblp:35/4137 · also Jonathan Shihao Ji
· DBLP profile ↗
45ranked-venue papers
9as first author
23since 2021 · last 2025
0000-0002-3573-5379ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reliable and Efficient Container Orchestration of LLMs via MCPabstractThis paper presents a structured decoding approach to support reliable and efficient container orchestration using large language models (LLMs) in conjunction with the Model Context Protocol (MCP), a standard interface for LLMs to interact with Docker and Kubernetes. We address key challenges in using LLMs for container orchestration: high token overhead from outputs and the risk of generating invalid or unsafe commands. Empirical results demonstrate up to a 76.2% latency reduction. Han Xu 0014, Shihao Ji 0001 |
CIKM | 3 |
| 2025 | HyperGCL: Multi-Modal Graph Contrastive Learning via Learnable Hypergraph ViewsabstractRecent advancements in Graph Contrastive Learning (GCL) have demonstrated remarkable effectiveness in improving graph representations. However, relying on predefined augmentations (e.g., node dropping, edge perturbation, attribute masking) may result in the loss of task-relevant information and a lack of adaptability to diverse input data. Furthermore, the selection of negative samples remains rarely explored. In this paper, we introduce HyperGCL, a novel multimodal GCL framework from a hypergraph perspective. HyperGCLconstructs three distinct hypergraph views by jointly utilizing the input graph’s structure and attributes, enabling a comprehensive integration of multiple modalities in contrastive learning. A learnable adaptive topology augmentation technique enhances these views by preserving important relations and filtering out noise. View-specific encoders capture essential characteristics from each view, while a network-aware contrastive loss leverages the underlying topology to define positive and negative samples effectively. Extensive experiments on benchmark datasets demonstrate that HyperGCLachieves state-of-the-art node classification performance. Khaled Mohammed Saifuddin, Shihao Ji 0001, Esra Akbas |
IJCNN | 2 |
| 2025 | Uni-LoRA: One Vector is All You NeedabstractLow-Rank Adaptation (LoRA) has become the de facto parameter-efficient fine-tuning (PEFT) method for large language models (LLMs) by constraining weight updates to low-rank matrices. Recent works such as Tied-LoRA, VeRA, and VB-LoRA push efficiency further by introducing additional constraints to reduce the trainable parameter space. In this paper, we show that the parameter space reduction strategies employed by these LoRA variants can be formulated within a unified framework, Uni-LoRA, where the LoRA parameter space, flattened as a high-dimensional vector space R^D, can be reconstructed through a projection from a subspace R^d, with d << D. We demonstrate that the fundamental difference among various LoRA methods lies in the choice of the projection matrix, P ∈ R^{D×d}.
Most existing LoRA variants rely on layer-wise or structure-specific projections that limit cross-layer parameter sharing, thereby compromising parameter efficiency. In light of this, we introduce an efficient and theoretically grounded projection matrix that is isometric, enabling global parameter sharing and reducing computation overhead. Furthermore, under the unified view of Uni-LoRA, this design requires only a single trainable vector to reconstruct LoRA parameters for the entire LLM -- making Uni-LoRA both a unified framework and a “one-vector-only” solution. Extensive experiments on GLUE, mathematical reasoning, and instruction tuning benchmarks demonstrate that Uni-LoRA achieves state-of-the-art parameter efficiency while outperforming or matching prior approaches in predictive performance. Kaiyang Li 0001, Shaobo Han, Qing Su 0001, Wei Li 0059, Zhipeng Cai 0001, Shihao Ji 0001 |
NeurIPS | 6 |
| 2024 | Towards Energy-Efficient Llama2 Architecture on Embedded FPGAsabstractLarge language models (LLMs) have shown immense potential for applications in information retrieval and knowledge management, but their computational and memory demands pose challenges for resource-constrained devices. In response, this work introduces an FPGA-based accelerator designed to improve LLM inference performance on embedded devices. We leverage quantization techniques, asynchronous computation, and a fully-pipelined accelerator to enhance efficiency. Our empirical evaluations, conducted using the TinyLlama 1.1B model on a Xilinx ZCU102 platform, demonstrate a 14.3-15.8x speedup and a 6.1x energy efficiency improvement over running exclusively on the ZCU102 processing system (PS). Han Xu 0014, Shihao Ji 0001 |
CIKM | 3 |
| 2024 | Unsqueeze [CLS] Bottleneck to Learn Rich Representations
Qing Su 0001, Shihao Ji 0001 |
ECCV (42) | 2 |
| 2024 | VB-LoRA: Extreme Parameter Efficient Fine-Tuning with Vector BanksabstractAs the adoption of large language models increases and the need for per-user or per-task model customization grows, the parameter-efficient fine-tuning (PEFT) methods, such as low-rank adaptation (LoRA) and its variants, incur substantial storage and transmission costs. To further reduce stored parameters, we introduce a "divide-and-share" paradigm that breaks the barriers of low-rank decomposition across matrix dimensions, modules, and layers by sharing parameters globally via a vector bank. As an instantiation of the paradigm to LoRA, our proposed VB-LoRA composites all the low-rank matrices of LoRA from a shared vector bank with a differentiable top-$k$ admixture module. VB-LoRA achieves extreme parameter efficiency while maintaining comparable or better performance compared to state-of-the-art PEFT methods. Extensive experiments demonstrate the effectiveness of VB-LoRA on natural language understanding, natural language generation, instruction tuning, and mathematical reasoning tasks. When fine-tuning the Llama2-13B model, VB-LoRA only uses 0.4% of LoRA's stored parameters, yet achieves superior results. Our source code is available at https://github.com/leo-yangli/VB-LoRA. This method has been merged into the Hugging Face PEFT package. Yang Li 0146, Shaobo Han, Shihao Ji 0001 |
NeurIPS | 3 |
| 2024 | UAV3D: A Large-scale 3D Perception Benchmark for Unmanned Aerial VehiclesabstractUnmanned Aerial Vehicles (UAVs), equipped with cameras, are employed in numerous applications, including aerial photography, surveillance, and agriculture. In these applications, robust object detection and tracking are essential for the effective deployment of UAVs. However, existing benchmarks for UAV applications are mainly designed for traditional 2D perception tasks, restricting thedevelopment of real-world applications that require a 3D understanding of the environment. Furthermore, despite recent advancements in single-UAV perception, limited views of a single UAV platform significantly constrain its perception capabilities over long distances or in occluded areas. To address these challenges, we introduce UAV3D – a benchmark designed to advance research in both 3D andcollaborative 3D perception tasks with UAVs. UAV3D comprises 1,000 scenes, each of which has 20 frames with fully annotated 3D bounding boxes on vehicles. We provide the benchmark for four 3D perception tasks: single-UAV 3D object detection, single-UAV object tracking, collaborative-UAV 3D object detection, and collaborative-UAV object tracking. Our dataset and code are available athttps://huiyegit.github.io/UAV3D_Benchmark/. Rajshekhar Sunderraman, Shihao Ji 0001 |
NeurIPS | 3 |
| 2024 | MatchXML: An Efficient Text-Label Matching Framework for Extreme Multi-Label Text ClassificationabstractThe eXtreme Multi-label text Classification (XMC) refers to training a classifier that assigns a text sample with relevant labels from an extremely large-scale label set (e.g., millions of labels). We propose MatchXML, an efficient textlabel matching framework for XMC. We observe that the label embeddings generated from the sparse Term Frequency-Inverse Document Frequency (TF–IDF) features have several limitations. We thus propose label2vec to effectively train the semantic dense label embeddings by the Skip-gram model. The dense label embeddings are then used to build a Hierarchical Label Tree by clustering. In fine-tuning the pre-trained encoder Transformer, we formulate the multi-label text classification as a text-label matching problem in a bipartite graph. We then extract the dense text representations from the fine-tuned Transformer. Besides the fine-tuned dense text embeddings, we also extract the static dense sentence embeddings from a pre-trained Sentence Transformer. Finally, a linear ranker is trained by utilizing the sparse TF–IDF features, the fine-tuned dense text representations, and static dense sentence features. Experimental results demonstrate that MatchXML achieves the state-of-the-art accuracies on five out of six datasets. As for the training speed, MatchXML outperforms the competing methods on all the six datasets. Our source code is publicly available athttps://github.com/huiyegit/MatchXML. Rajshekhar Sunderraman, Shihao Ji 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Towards Bridging the Performance Gaps of Joint Energy-Based ModelsabstractCan we train a hybrid discriminative-generative model with a single network? This question has recently been answered in the affirmative, introducing the field of Joint Energy-based Model (JEM) [17], [48], which achieves high classification accuracy and image generation quality simultaneously. Despite recent advances, there remain two performance gaps: the accuracy gap to the standard softmax classifier, and the generation quality gap to state-of-the-art generative models. In this paper, we introduce a variety of training techniques to bridge the accuracy gap and the generation quality gap of JEM. 1) We incorporate a recently proposed sharpness-aware minimization (SAM) framework to train JEM, which promotes the energy landscape smoothness and the generalization of JEM. 2) We exclude data augmentation from the maximum likelihood estimate pipeline of JEM, and mitigate the negative impact of data augmentation to image generation quality. Extensive experiments on multiple datasets demonstrate our SADA-JEM achieves state-of-the-art performances and outperforms JEM in image classification, image generation, calibration, out-of-distribution detection and adversarial robustness by a notable margin. Our code is available at https://github.com/sndnyang/SADAJEM. Xiulong Yang, Qing Su 0001, Shihao Ji 0001 |
CVPR | 3 |
| 2023 | FLSL: Feature-level Self-supervised LearningabstractCurrent self-supervised learning (SSL) methods (e.g., SimCLR, DINO, VICReg, MOCOv3) target primarily on representations at instance level and do not generalize well to dense prediction tasks, such as object detection and segmentation. Towards aligning SSL with dense predictions, this paper demonstrates for the first time the underlying mean-shift clustering process of Vision Transformers (ViT), which aligns well with natural image semantics (e.g., a world of objects and stuffs). By employing transformer for joint embedding and clustering, we propose a bi-level feature clustering SSL method, coined Feature-Level Self-supervised Learning (FLSL). We present the formal definition of the FLSL problem and construct the objectives from the mean-shift and k-means perspectives. We show that FLSL promotes remarkable semantic cluster representations and learns an embedding scheme amenable to intra-view and inter-view feature clustering. Experiments show that FLSL yields significant improvements in dense prediction tasks, achieving 44.9 (+2.8)% AP and 46.5% AP in object detection, as well as 40.8 (+2.3)% AP and 42.1% AP in instance segmentation on MS-COCO, using Mask R-CNN with ViT-S/16 and ViT-S/8 as backbone, respectively. FLSL consistently outperforms existing SSL methods across additional benchmarks, including UAV object detection on UAVDT, and video instance segmentation on DAVIS 2017. We conclude by presenting visualization and various ablation studies to better understand the success of FLSL. The source code is available at https://github.com/ISL-CV/FLSL. Qing Su 0001, Anton Netchaev, Shihao Ji 0001 |
NeurIPS | 4 |
| 2023 | M-EBM: Towards Understanding the Manifolds of Energy-Based Models
Xiulong Yang, Shihao Ji 0001 |
PAKDD (1) | 2 |
| 2023 | Sparse Graph Attention NetworksabstractGraph Neural Networks (GNNs) have proved to be an effective representation learning framework for graph-structured data, and have achieved state-of-the-art performance on many practical predictive tasks. Among the variants of GNNs, Graph Attention Networks (GATs) improve the performance of many graph learning tasks through a dense attention mechanism. However, real-world graphs are often very large and noisy, and GATs are prone to overfitting if not regularized properly. In this paper, we propose Sparse Graph Attention Networks (SGATs) that learn sparse attention coefficients under an L0-norm regularization, and the learned sparse attentions are then used for all GNN layers, resulting in an edge-sparsified graph. By doing so, we can identify noisy/task-irrelevant edges, and thus perform feature aggregation on most informative neighbors. Extensive experiments on synthetic and real-world (assortative and disassortative) graph learning benchmarks demonstrate the superior performance of SGATs. Furthermore, the removed edges can be interpreted intuitively and quantitatively. To the best of our knowledge, this is the first graph learning algorithm that shows significant redundancies in graphs and edge-sparsified graphs can achieve similar (on assortative graphs) or sometimes higher (on disassortative graphs) predictive performances than original graphs. Our code is available at https://github.com/Yangyeeee/SGAT. Shihao Ji 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | APSNet: Attention Based Point Cloud Sampling
Xiulong Yang, Shihao Ji 0001 |
BMVC | 3 |
| 2022 | Semantic Structure Based Query Graph Prediction for Question Answering over Knowledge GraphabstractBuilding query graphs from natural language questions is an important step in complex question answering over knowledge graph (Complex KGQA). In general, a question can be correctly answered if its query graph is built correctly and the right answer is then retrieved by issuing the query graph against the KG. Therefore, this paper focuses on query graph generation from natural language questions. Existing approaches for query graph generation ignore the semantic structure of a question, resulting in a large number of noisy query graph candidates that undermine prediction accuracies. In this paper, we define six semantic structures from common questions in KGQA and develop a novel Structure-BERT to predict the semantic structure of a question. By doing so, we can first filter out noisy candidate query graphs by the predicted semantic structures, and then rank the remaining candidates with a BERT-based ranking model. Extensive experiments on two popular benchmarks MetaQA and WebQuestionsSP (WSP) demonstrate the effectiveness of our method as compared to state-of-the-arts. Shihao Ji 0001 |
COLING | 2 |
| 2022 | ChiTransformer: Towards Reliable Stereo from CuesabstractCurrent stereo matching techniques are challenged by restricted searching space, occluded regions and sheer size. While single image depth estimation is spared from these challenges and can achieve satisfactory results with the extracted monocular cues, the lack of stereoscopic relationship renders the monocular prediction less reliable on its own especially in highly dynamic or cluttered environments. To address these issues in both scenarios, we present an optic-chiasm-inspired self-supervised binocular depth estimation method, wherein vision transformer (ViT) with a gated positional cross-attention (GPCA) layer is designed to enable feature-sensitive pattern retrieval between views, while retaining the extensive context information aggregated through self-attentions. Monocular cues from a single view are thereafter conditionally rectified by a blending layer with the retrieved pattern pairs. This crossover design is biologically analogous to the optic-chasma structure in human visual system and hence the name, Chi-Transformer. Our experiments show that this architecture yields substantial improvements over state-of-the-art self-supervised stereo approaches by 11%, and can be used on both rectilinear and non-rectilinear (e.g., fisheye) images.11https://github.com/ISL-CV/ChiTransformer.git Qing Su 0001, Shihao Ji 0001 |
CVPR | 2 |
| 2021 | Generative Dynamic Patch Attack
Xiang Li 0080, Shihao Ji 0001 |
BMVC | 2 |
| 2021 | Improving Text-to-Image Synthesis Using Contrastive Learning
Xiulong Yang, Martin Takác 0001, Rajshekhar Sunderraman, Shihao Ji 0001 |
BMVC | 5 |
| 2021 | JEM++: Improved Techniques for Training JEMabstractJoint Energy-based Model (JEM) [12] is a recently proposed hybrid model that retains strong discriminative power of modern CNN classifiers, while generating samples rivaling the quality of GAN-based approaches. In this paper, we propose a variety of new training procedures and architecture features to improve JEM’s accuracy, training stability, and speed altogether. 1) We propose a proximal SGLD to generate samples in the proximity of samples from previous step, which improves the stability. 2) We further treat the approximate maximum likelihood learning of EBM as a multi-step differential game, and extend the YOPO framework [47] to cut out redundant calculations during backpropagation, which accelerates the training substantially. 3) Rather than initializing SGLD chain from random noise, we introduce a new informative initialization that samples from a distribution estimated from training data. 4) This informative initialization allows us to enable batch normalization in JEM, which further releases the power of modern CNN architectures for hybrid modeling.1 Xiulong Yang, Shihao Ji 0001 |
ICCV | 2 |
| 2021 | A Unified Density-Driven Framework For Effective Data Denoising And Robust AbstentionabstractThe success of Deep Neural Networks (DNNs) highly depends on data quality. Moreover, predictive uncertainty reduces reliability of DNNs for real-world applications. In this paper, we aim to address these two issues by proposing a unified filtering framework leveraging underlying data density, that effectively denoises training data as well as avoids predicting confusing samples. Our proposed framework differentiates noise from clean data samples without modifying existing DNN architectures or loss functions. Extensive experiments on multiple benchmark datasets and recent COVIDx dataset demonstrate the effectiveness of our framework over state-of-the-art (SOTA) methods in denoising training data and abstaining uncertain test data. Krishanu Sarker, Xiulong Yang, Yang Li 0146, Saeid Belkasim, Shihao Ji 0001 |
ICIP | 5 |
| 2021 | Neural Plasticity NetworksabstractNeural plasticity is an important functionality of human brain, in which number of neurons and synapses can shrink or expand in response to stimuli throughout the span of life. We model this dynamic learning process as an$L_{0}$- norm regularized binary optimization problem, in which each unit of a neural network (e.g., weight, neuron or channel, etc.) is attached with a stochastic binary gate, whose parameters determine the level of activity of a unit in the network. At the beginning, only a small portion of binary gates (therefore the corresponding neurons) are activated, while the remaining neurons are in a hibernation mode. As the learning proceeds, some neurons might be activated or deactivated if doing so can be justified by the cost-benefit tradeoff measured by the$L_{0}$-norm regularized objective. As the training gets mature, the probability of transition between activation and deactivation will diminish until a final hardening stage. We demonstrate that all of these learning dynamics can be modulated by a single parameter$k$seamlessly. Our neural plasticity network (NPN) can prune or expand a network depending on the initial capacity of network provided by the user; it also unifies dropout (when$k=0$), traditional training of DNNs (when$k=\infty$) and interpolates between these two. To the best of our knowledge, this is the first learning framework that unifies network sparsification and network expansion in an end-to-end training pipeline. Extensive experiments on synthetic dataset and multiple image classification benchmarks demonstrate the superior performance of NPN. We show that both network sparsification and network expansion can yield compact models of similar architectures, while retaining competitive accuracies of the original networks.11Our source code is available at https://github.com/leo-yangli/npns. Yang Li 0146, Shihao Ji 0001 |
IJCNN | 2 |
| 2021 | Dep-L0: Improving L0-Based Network Sparsification via Dependency Modeling
Yang Li 0146, Shihao Ji 0001 |
ECML/PKDD (3) | 2 |
| 2021 | Generative Max-Mahalanobis Classifiers for Image Classification, Generation and More
Xiulong Yang, Xiang Li 0080, Shihao Ji 0001 |
ECML/PKDD (2) | 5 |
| 2021 | Adversarial Privacy-Preserving Graph Embedding Against Inference AttackabstractRecently, the surge in popularity of the Internet of Things (IoT), mobile devices, social media, etc., has opened up a large source for graph data. Graph embedding has been proved extremely useful to learn low-dimensional feature representations from graph-structured data. These feature representations can be used for a variety of prediction tasks from node classification to link prediction. However, the existing graph embedding methods do not consider users' privacy to prevent inference attacks. That is, adversaries can infer users' sensitive information by analyzing node representations learned from graph embedding algorithms. In this article, we propose adversarial privacy graph embedding (APGE), a graph adversarial training framework that integrates the disentangling and purging mechanisms to remove users' private information from learned node representations. The proposed method preserves the structural information and utility attributes of a graph while concealing users' private attributes from inference attacks. Extensive experiments on real-world graph data sets demonstrate the superior performance of APGE compared to the state-of-the-arts. Our source code can be found at https://github.com/KaiyangLi1992/Privacy-Preserving-Social-Network-Embedding. Kaiyang Li 0001, Guangchun Luo, Wei Li 0059, Shihao Ji 0001, Zhipeng Cai 0001 |
IEEE Internet Things J. | 5 |
| 2020 | Learning with Multiplicative PerturbationsabstractAdversarial Training (AT) and Virtual Adversarial Training (VAT) are the regularization techniques that train Deep Neural Networks (DNNs) with adversarial examples generated by adding small but worst-case perturbations to input examples. In this paper, we propose xAT and xVAT, new adversarial training algorithms that generate multiplicative perturbations to input examples for robust training of DNNs. Such perturbations are much more perceptible and interpretable than their additive counterparts exploited by AT and VAT. Furthermore, the multiplicative perturbations can be generated transductively or inductively, while the standard AT and VAT only support a transductive implementation. We conduct a series of experiments that analyze the behavior of the multiplicative perturbations and demonstrate that xAT and xVAT match or outperform state-of-the-art classification accuracies across multiple established benchmarks while being about 30% faster than their additive counterparts. Our source code can be found at https://github.com/sndnyang/xvat. Xiulong Yang, Shihao Ji 0001 |
ICPR | 2 |
| 2019 | Extreme Stochastic Variational Inference: Distributed Inference for Large Scale Mixture ModelsabstractMixture of exponential family models are among the most fundamental and widely used statistical models. Stochastic variational inference (SVI), the state-of-the-art algorithm for parameter estimation in such models is inherently serial. Moreover, it requires the parameters to fit in the memory of a single processor; this poses serious limitations on scalability when the number of parameters is in billions. In this paper, we present extreme stochastic variational inference (ESVI), a distributed, asynchronous and lock-free algorithm to perform variational inference for mixture models on massive real world datasets. ESVI overcomes the limitations of SVI by requiring that each processor only access a subset of the data and a subset of the parameters, thus providing data and model parallelism simultaneously. Our empirical study demonstrates that ESVI not only outperforms VI and SVI in wallclock-time, but also achieves a better quality solution. To further speed up computation and save memory when fitting large number of topics, we propose a variant ESVI-TOPK which maintains only the top-k important topics. Empirically, we found that using top 25% topics suffices to achieve the same accuracy as storing all the topics. Jiong Zhang 0001, Parameswaran Raman, Shihao Ji 0001, Hsiang-Fu Yu, S. V. N. Vishwanathan, Inderjit S. Dhillon |
AISTATS | 3 |
| 2019 | Toward Filament Segmentation Using Deep Neural NetworksabstractWe use a well-known deep neural network framework, called Mask R-CNN, for identification of solar filaments in full-disk H-$\alpha$ images from Big Bear Solar Observatory (BBSO). The image data, collected from BBSO's archive, are integrated with the spatiotemporal metadata of filaments retrieved from the Heliophysics Events Knowledgebase (HEK) system. This integrated data is then treated as the ground-truth in the training process of the model. The available spatial metadata are the output of a currently running filament-detection module developed and maintained by the Feature Finding Team; an international consortium selected by NASA. Despite the known challenges in the identification and characterization of filaments by the existing module, which in turn are inherited into any other module that intends to learn from such outputs, Mask R-CNN shows promising results. Trained and validated on two years worth of BBSO data, this model is then tested on the three following years. Our case-by-case and overall analyses show that Mask R-CNN can clearly compete with the existing module and in some cases even perform better. Several cases of false positives and false negatives, that are correctly segmented by this model are also shown. The overall advantages of using the proposed model are two-fold: First, deep neural networks' performance generally improves as more annotated data, or better annotations are provided. Second, such a model can be scaled up to detect other solar events, as well as a single multi-purpose module. The results presented in this study introduce a proof of concept in benefits of employing deep neural networks for detection of solar events, and in particular, filaments. Azim Ahmadzadeh, Sushant S. Mahajan, Dustin Kempton, Rafal A. Angryk, Shihao Ji 0001 |
IEEE BigData | 5 |
| 2019 | L0-ARM: Network Sparsification via Stochastic Binary Optimization
Yang Li 0146, Shihao Ji 0001 |
ECML/PKDD (2) | 2 |
| 2019 | Parallelizing Word2Vec in Shared and Distributed MemoryabstractWord2vec is a widely used algorithm for extracting low-dimensional vector representations of words. State-of-the-art algorithms including those by Mikolov et al. [1] , [2] have been parallelized for multi-core CPU architectures, but are based on vector-vector operations with “Hogwild” updates that are memory-bandwidth intensive and do not efficiently use computational resources. In this paper, we propose “HogBatch” by improving reuse of various data structures in the algorithm through the use of minibatching and negative sample sharing, hence allowing us to express the problem using matrix multiply operations. We also explore different techniques to distribute word2vec computation across nodes in a computer cluster, and demonstrate good strong scalability up to 32 nodes. The new algorithm is particularly suitable for modern multi-core/many-core architectures, especially Intel's latest Knights Landing processors, and allows us to scale up the computation near linearly across cores and nodes, and process hundreds of millions of words per second, which is the fastest word2vec implementation to the best of our knowledge. We released the source code for reproducible research and general usage. Shihao Ji 0001, Nadathur Satish, Sheng Li 0007, Pradeep Dubey |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Towards Robust Human Activity Recognition from RGB Video Stream with Limited Labeled DataabstractHuman activity recognition based on video streams has received numerous attentions in recent years. Due to lack of depth information, RGB video based activity recognition performs poorly compared to RGB-D video based solutions. On the other hand, acquiring depth information, inertia etc. is costly and requires special equipment, whereas RGB video streams are available in ordinary cameras. Hence, our goal is to investigate whether similar or even higher accuracy can be achieved with RGB-only modality. In this regard, we propose a novel framework that couples skeleton data extracted from RGB video and deep Bidirectional Long Short Term Memory (BLSTM) model for activity recognition. A big challenge of training such a deep network is the limited training data, and exploring RGB-only stream significantly exaggerates the difficulty. We therefore propose a set of algorithmic techniques to train this model effectively, e.g., data augmentation, dynamic frame dropout and gradient injection. The experiments demonstrate that our RGB-only solution surpasses the state-of-the-art approaches that all exploit RGB-D video streams by a notable margin. This makes our solution widely deployable with ordinary cameras. Krishanu Sarker, Mohamed Masoud, Saeid Belkasim, Shihao Ji 0001 |
ICMLA | 4 |
| 2016 | WordRank: Learning Word Embeddings via Robust RankingabstractEmbedding words in a vector space has gained a lot of attention in recent years.While stateof-the-art methods provide efficient computation of word similarities via a low-dimensional matrix embedding, their motivation is often left unclear.In this paper, we argue that word embedding can be naturally viewed as a ranking problem due to the ranking nature of the evaluation metrics.Then, based on this insight, we propose a novel framework Wor-dRank that efficiently estimates word representations via robust ranking, in which the attention mechanism and robustness to noise are readily achieved via the DCG-like ranking losses.The performance of WordRank is measured in word similarity and word analogy benchmarks, and the results are compared to the state-of-the-art word embedding techniques.Our algorithm is very competitive to the state-of-the-arts on large corpora, while outperforms them by a significant margin when the training set is limited (i.e., sparse and noisy).With 17 million tokens, WordRank performs almost as well as existing methods using 7.2 billion tokens on a popular word similarity benchmark.Our multi-node distributed implementation of WordRank is publicly available for general usage. Shihao Ji 0001, Hyokun Yun, Pinar Yanardag Delul, Shin Matsushima, S. V. N. Vishwanathan |
EMNLP | 1 |
| 2011 | Intent-based diversification of web search results: metrics and algorithms
Olivier Chapelle, Shihao Ji 0001, Ciya Liao, Emre Velipasaoglu, Larry Lai, Su-Lin Wu |
Inf. Retr. | 2 |
| 2010 | User behavior driven ranking without editorial judgmentsabstractWe explore the potential of using users click-through logs where no editorial judgment is available to improve the ranking function of a vertical search engine. We base our analysis on the Cumulate Relevance Model, a user behavior model recently proposed as a way to extract relevance signal from click-through logs. We propose a novel way of directly learning the ranking function, effectively by-passing the need to have explicit editorial relevance label for each query-document pair. This approach potentially adjusts more closely the ranking function to a variety of user behaviors both at the individual and at the aggregate levels. We investigate two ways of using behavioral model; First, we consider the parametric approach where we learn the estimates of document relevance and use them as targets for the machine learned ranking schemes. In the second, functional approach, we learn a function that maximizes the behavioral model likelihood, effectively by-passing the need to estimate a substitute for document labels. Experiments using user session data collected from a commercial vertical search engine demonstrate the potential of our approach. While in terms of DCG, the editorial model out-perform the behavioral one, online experiments show that the behavioral model is on par --if not superior-- to the editorial model. To our knowledge, this is the first report in the Literature of a competitive behavioral model in a commercial setting Taesup Moon, Georges Dupret, Shihao Ji 0001, Ciya Liao, Zhaohui Zheng 0001 |
CIKM | 3 |
| 2009 | Incorporating robustness into web ranking evaluationabstractIn many Web search engines, a ranking function is selected for deployment mainly by comparing the relevance measurements over candidates. Due to the dynamical nature of the Web, the ranking features and the query and URL distribution on which the ranking functions are built, may change dramatically over time. The actual relevance of the function may degrade, and thus the previous function selection conclusions become invalid. In this work we suggest to select Web ranking functions according to both their relevance and robustness to the changes that may lead to relevance degradation over time. We argue that the ranking robustness can be effectively measured by taking into account the ranking score distribution across search results. We then propose two alternatives to the NDCG metric that both incorporate ranking robustness into ranking function evaluation and selection. A machine learning approach is developed to learn the parameters that control the metric sensitivity to score turbulence, from human-judged preference data. Shihao Ji 0001, Zhaohui Zheng 0001, Yi Chang 0001, Anlei Dong |
CIKM | 3 |
| 2009 | Empirical Exploitation of Click Data for Task Specific Ranking
Anlei Dong, Yi Chang 0001, Shihao Ji 0001, Ciya Liao, Zhaohui Zheng 0001 |
EMNLP | 3 |
| 2009 | Global ranking by exploiting user clicksabstractIt is now widely recognized that user interactions with search results can provide substantial relevance information on the documents displayed in the search results. In this paper, we focus on extracting relevance information from one source of user interactions, i.e., user click data, which records the sequence of documents being clicked and not clicked in the result set during a user search session. We formulate the problem as a global ranking problem, emphasizing the importance of the sequential nature of user clicks, with the goal to predict the relevance labels of all the documents in a search session. This is distinct from conventional learning to rank methods that usually design a ranking model defined on a single document; in contrast, in our model the relational information among the documents as manifested by an aggregation of user clicks is exploited to rank all the documents jointly. In particular, we adapt several sequential supervised learning algorithms, including the conditional random field (CRF), the sliding window method and the recurrent sliding window method, to the global ranking problem. Experiments on the click data collected from a commercial search engine demonstrate that our methods can outperform the baseline models for search results re-ranking. Shihao Ji 0001, Ke Zhou 0002, Ciya Liao, Zhaohui Zheng 0001, Gui-Rong Xue, Olivier Chapelle, Gordon Sun, Hongyuan Zha |
SIGIR | 1 |
| 2009 | Comparing both relevance and robustness in selection of web ranking functionsabstractIn commercial search engines, a ranking function is selected for deployment mainly by comparing the relevance measurements over candidates. In this paper we suggest to select Web ranking functions according to both their relevance and robustness to the changes that may lead to relevance degradation over time. We argue that the ranking robustness can be effectively measured by taking into account the ranking score distribution across Web pages. We then improve NDCG with two new metrics and show their superiority in terms of stability to ranking score turbulence and stability in function selection. Shihao Ji 0001, Zhaohui Zheng 0001 |
SIGIR | 3 |
| 2009 | Semisupervised Learning of Hidden Markov Models via a Homotopy Method
Shihao Ji 0001, Layne T. Watson, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Multitask Classification by Learning the Task RelevanceabstractWe consider the problem of multitask learning (MTL), in which we simultaneously learn classifiers for multiple data sets (tasks), with sharing of intertask data as appropriate. We introduce a set of relevance parameters that control the degree to which data from other tasks are used in estimating the current task's classifier parameters. The set of relevance parameters are learned by maximizing their posterior probability, yielding an expectation-maximization (EM) algorithm. We illustrate the effectiveness of our approach through experimental results on a practical data set. Jun Fang 0001, Shihao Ji 0001, Ya Xue, Lawrence Carin |
IEEE Signal Process. Lett. | 2 |
| 2007 | Point-Based Policy Iteration
Shihao Ji 0001, Ronald Parr, Hui Li 0068, Xuejun Liao, Lawrence Carin |
AAAI | 1 |
| 2007 | Bayesian compressive sensing and projection optimizationabstractThis paper introduces a new problem for which machine-learning tools may make an impact. The problem considered is termed "compressive sensing", in which a real signal of dimension N is measured accurately based on K << N real measurements. This is achieved under the assumption that the underlying signal has a sparse representation in some basis (e.g., wavelets). In this paper we demonstrate how techniques developed in machine learning, specifically sparse Bayesian regression and active learning, may be leveraged to this new problem. We also point out future research directions in compressive sensing of interest to the machine-learning community. Shihao Ji 0001, Lawrence Carin |
ICML | 1 |
| 2007 | Cost-sensitive feature acquisition and classification
Shihao Ji 0001, Lawrence Carin |
Pattern Recognit. | 1 |
| 2007 | Adaptive Multimodality Sensing of LandminesabstractThe problem of adaptive multimodality sensing of landmines is considered based on electromagnetic induction (EMI) and ground-penetrating radar (GPR) sensors. Two formulations are considered based on a partially observable Markov decision process (POMDP) framework. In the first formulation, it is assumed that sufficient training data are available, and a POMDP model is designed based on physics-based features, with model selection performed via a variational Bayes analysis of several possible models. In the second approach, the training data are assumed absent or insufficient, and a lifelong-learning approach is considered, in which exploration and exploitation are integrated. We provide a detailed description of both formulations, with example results presented using measured EMI and GPR data, for buried mines and clutter Lihan He, Shihao Ji 0001, Waymond R. Scott, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2006 | Homotopy-Based Semi-Supervised Hidden Markov Tree for Texture AnalysisabstractA semi-supervised hidden Markov tree (HMT) model is developed for texture analysis, incorporating both labeled and unlabeled data for training; the optimal balance between labeled and unlabeled data is estimated via the homotopy method. In traditional EM-based semi-supervised modeling, this balance is dictated by the relative size of labeled and unlabeled data, often leading to poor performance. Semi-supervised modeling may be viewed as a source allocation problem between labeled and unlabeled data, controlled by a parameter λ ∈ [0, 1], where λ = 0 and 1 correspond to the purely supervised HMT model and purely unsupervised HMT-based clustering, respectively. We consider the homotopy method to track a path of fixed points starting from λ = 0, with the optimal source allocation identified as a critical transition point where the solution is unsupported by the initial labeled data. Experimental results on real textures demonstrate the superiority of this method compared to the EM-based semi-supervised HMT training. Nilanjan Dasgupta, Shihao Ji 0001, Lawrence Carin |
ICASSP (2) | 2 |
| 2006 | Variational Bayes for Continuous Hidden Markov Models and Its Application to Active LearningabstractIn this paper, we present a varitional Bayes (VB) framework for learning continuous hidden Markov models (CHMMs), and we examine the VB framework within active learning. Unlike a maximum likelihood or maximum a posteriori training procedure, which yield a point estimate of the CHMM parameters, VB-based training yields an estimate of the full posterior of the model parameters. This is particularly important for small training sets since it gives a measure of confidence in the accuracy of the learned model. This is utilized within the context of active learning, for which we acquire labels for those feature vectors for which knowledge of the associated label would be most informative for reducing model-parameter uncertainty. Three active learning algorithms are considered in this paper: 1) query by committee (QBC), with the goal of selecting data for labeling that minimize the classification variance, 2) a maximum expected information gain method that seeks to label data with the goal of reducing the entropy of the model parameters, and 3) an error-reduction-based procedure that attempts to minimize classification error over the test data. The experimental results are presented for synthetic and measured data. We demonstrate that all of these active learning methods can significantly reduce the amount of required labeling, compared to random selection of samples for labeling. Shihao Ji 0001, Balaji Krishnapuram, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Adaptive multi-aspect target classification and detection with hidden Markov modelsabstractWe consider target classification and detection based on backscattered observations measured from a sequence of target-sensor orientations. The multi-aspect scattered waves from a given target are modeled with a hidden Markov model (HMM). The targets are assumed concealed and the absolute target-sensor orientation is assumed unknown; therefore, it is possible to control only the angular displacements (change in orientation) between consecutive measurements. The performance of the HMM classifiers/detectors is influenced by the choice of the angular displacements, the optimization of which motivates the developed adaptive search strategies, based on entropy-driven optimality criteria. The search proceeds in a sequential fashion. Based on the previous observations and their associated angular displacements, one determines the optimal next displacement to perform an associated observation. The search strategies are detailed and example results presented on adaptive classification and detection of underwater targets. Shihao Ji 0001, Xuejun Liao, Lawrence Carin |
ICASSP (2) | 1 |