VLDB 2026 Research / reviewers in the wild / expert
Wei-Ying Ma
dblp:m/WYMa · also Weiying Ma
· DBLP profile ↗
281ranked-venue papers
13as first author
40since 2021 · last 2026
0000-0002-7384-0735ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 113 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 105 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 100 · 3 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 1 first-authorHuman-computer interaction and ubiquitous computing · 6Computer networks · 4Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningabstractVirtual screening (VS) is an essential task in drug discovery, focusing on the identification of small-molecule ligands that bind to specific protein pockets. Existing deep learning methods, from early regression models to recent contrastive learning approaches, primarily rely on structural data while overlooking protein sequences, which are more accessible and can enhance generalizability. However, directly integrating protein sequences poses challenges due to the redundancy and noise in large-scale protein-ligand datasets. To address these limitations, we propose S²Drug, a two-stage framework that explicitly incorporates protein Sequence information and 3D Structure context in protein-ligand contrastive representation learning. In the first stage, we perform protein sequence pretraining on ChemBL using an ESM2-based backbone, combined with a tailored data sampling strategy to reduce redundancy and noise on both protein and ligand sides. In the second stage, we fine-tune on PDBBind by fusing sequence and structure information through a residue-level gating module, while introducing an auxiliary binding site prediction task. This auxiliary task guides the model to accurately localize binding residues within the protein sequence and capture their 3D spatial arrangement, thereby refining protein-ligand matching. Across multiple benchmarks, S²Drug consistently improves virtual screening performance and achieves strong results on binding site prediction, demonstrating the value of bridging sequence and structure in contrastive learning. Bowei He, Yankai Chen 0001, Yanyan Lan, Chen Ma 0001, Philip S. Yu, Ya-Qin Zhang, Wei-Ying Ma |
AAAI | 8 |
| 2026 | Learning Protein-Ligand Binding in Hyperbolic SpaceabstractProtein-ligand binding prediction is central to virtual screening and affinity ranking, two fundamental tasks in drug discovery. While recent retrieval-based methods embed ligands and protein pockets into Euclidean space for similarity-based search, the geometry of Euclidean embeddings often fails to capture the hierarchical structure and fine-grained affinity variations intrinsic to molecular interactions. In this work, we propose HypSeek, a hyperbolic representation learning framework that embeds ligands, protein pockets, and sequences into Lorentz-model hyperbolic space. By leveraging the exponential geometry and negative curvature of hyperbolic space, HypSeek enables expressive, affinity-sensitive embeddings that can effectively model both global activity and subtle functional differences–particularly in challenging cases such as activity cliffs, where structurally similar ligands exhibit large affinity gaps. Our model unifies virtual screening and affinity ranking in a single framework, introducing a protein-guided three-tower architecture to enhance representational structure. HypSeek improves early enrichment in virtual screening on DUD-E from 42.63 to 51.44 (+20.7%) and affinity ranking correlation on JACS from 0.5774 to 0.7239 (+25.4%), demonstrating the benefits of hyperbolic geometry across both tasks and highlighting its potential as a powerful inductive bias for protein-ligand modeling. Wenyu Zhu, Ya-Qin Zhang, Wei-Ying Ma, Yanyan Lan |
AAAI | 6 |
| 2026 | R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement LearningabstractYiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lihao Wang, Lei Bai, Lijun Wu, Wei-Ying Ma, Hao Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. YiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lei Bai 0001, Lijun Wu 0003, Wei-Ying Ma, Hao Zhou 0012 |
ACL (1) | 9 |
| 2025 | RetroDiff: Retrosynthesis as Multi-stage Distribution InterpolationabstractRetrosynthesis poses a key challenge in biopharmaceuticals, aiding chemists in finding appropriate reactant molecules for given product molecules. With reactants and products represented as 2D graphs, retrosynthesis constitutes a conditional graph-to-graph (G2G) generative task. Inspired by advancements in discrete diffusion models for graph generation, we aim to design a diffusion-based method to address this problem. However, integrating a diffusion-based G2G framework while retaining essential chemical reaction template information presents a notable challenge. Our key innovation involves a multi-stage diffusion process. We decompose the retrosynthesis procedure to first sample external groups from the dummy distribution given products, then generate external bonds to connect products and generated groups. Interestingly, this generation process mirrors the reverse of the widely adapted semi-template retrosynthesis workflow, i.e., from reaction center identification to synthon completion. Based on these designs, we introduce Retrosynthesis Diffusion (RetroDiff), a novel diffusion-based method for the retrosynthesis task. Experimental results demonstrate that RetroDiff surpasses all semi-template methods in accuracy, and outperforms template-based and template-free methods in large-scale scenarios and molecular validity, respectively. Yiming Wang 0011, Yuxuan Song 0002, Minkai Xu, Rui Wang 0015, Hao Zhou 0012, Wei-Ying Ma |
AISTATS | 7 |
| 2025 | UniGEM: A Unified Approach to Generation and Property Prediction for MoleculesabstractMolecular generation and molecular property prediction are both crucial for drug discovery, but they are often developed independently. Inspired by recent studies, which demonstrate that diffusion model, a prominent generative approach, can learn meaningful data representations that enhance predictive tasks, we explore the potential for developing a unified generative model in the molecular domain that effectively addresses both molecular generation and property prediction tasks. However, the integration of these tasks is challenging due to inherent inconsistencies, making simple multi-task learning ineffective. To address this, we propose UniGEM, the first unified model to successfully integrate molecular generation and property prediction, delivering superior performance in both tasks. Our key innovation lies in a novel two-phase generative process, where predictive tasks are activated in the later stages, after the molecular scaffold is formed. We further enhance task balance through innovative training strategies. Rigorous theoretical analysis and comprehensive experiments demonstrate our significant improvements in both tasks. The principles behind UniGEM hold promise for broader applications, including natural language processing and computer vision. Shikun Feng, Yuyan Ni, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
ICLR | 5 |
| 2025 | Reframing Structure-Based Drug Design Model Evaluation via Metrics Correlated to Practical NeedsabstractRecent advances in structure-based drug design (SBDD) have produced surprising results, with models often generating molecules that achieve better Vina docking scores than actual ligands. However, these results are frequently overly optimistic due to the limitations of docking score accuracy and the challenges of wet-lab validation. While generated molecules may demonstrate high QED (drug-likeness) and SA (synthetic accessibility) scores, they often lack true drug-like properties or synthesizability. To address these limitations, we propose a model-level evaluation framework that emphasizes practical metrics aligned with real-world applications. Inspired by recent findings on the utility of generated molecules in ligand-based virtual screening, our framework evaluates SBDD models by their ability to produce molecules that effectively retrieve active compounds from chemical libraries via similarity-based searches. This approach provides a direct indication of therapeutic potential, bridging the gap between theoretical performance and real-world utility. Our experiments reveal that while SBDD models may excel in theoretical metrics like Vina scores, they often fall short in these practical metrics. By introducing this new evaluation strategy, we aim to enhance the relevance and impact of SBDD models for pharmaceutical research and development. Haichuan Tan, Yanwen Huang, Minsi Ren, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
ICLR | 6 |
| 2025 | Steering Protein Family Design through Profile Bayesian FlowabstractProtein family design emerges as a promising alternative by combining the advantages of de novo protein design and mutation-based directed evolution.In this paper, we propose ProfileBFN, the Profile Bayesian Flow Networks, for specifically generative modeling of protein families. ProfileBFN extends the discrete Bayesian Flow Network from an MSA profile perspective, which can be trained on single protein sequences by regarding it as a degenerate profile, thereby achieving efficient protein family design by avoiding large-scale MSA data construction and training. Empirical results show that ProfileBFN has a profound understanding of proteins. When generating diverse and novel family proteins, it can accurately capture the structural characteristics of the family. The enzyme produced by this method is more likely than the previous approach to have the corresponding function, offering better odds of generating diverse proteins with the desired functionality. Jingjing Gong, Siyu Long, Yuxuan Song 0002, Wenhao Huang 0001, Ziyao Cao, Hao Zhou 0012, Wei-Ying Ma |
ICLR | 10 |
| 2025 | Redefining the task of Bioactivity PredictionabstractSmall molecules are vital to modern medicine, and accurately predicting their bioactivity against protein targets is crucial for therapeutic discovery and development. However, current machine learning models often rely on spurious features, leading to biased outcomes. Notably, a simple pocket-only baseline can achieve results comparable to, and sometimes better than, more complex models that incorporate both the protein pockets and the small molecules. Our analysis reveals that this phenomenon arises from insufficient training data and an improper evaluation process, which is typically conducted at the pocket level rather than the small molecule level. To address these issues, we redefine the bioactivity prediction task by introducing the SIU dataset-a million-scale Structural small molecule-protein Interaction dataset for Unbiased bioactivity prediction task, which is 50 times larger than the widely used PDBbind. The bioactivity labels in SIU are derived from wet experiments and organized by label types, ensuring greater accuracy and comparability. The complexes in SIU are constructed using a majority vote from three commonly used docking software programs, enhancing their reliability. Additionally, the structure of SIU allows for multiple small molecules to be associated with each protein pocket, enabling the redefinition of evaluation metrics like Pearson and Spearman correlations across different small molecules targeting the same protein pocket. Experimental results demonstrate that this new task provides a more challenging and meaningful benchmark for training and evaluating bioactivity prediction models, ultimately offering a more robust assessment of model performance. Yanwen Huang, Yinjun Jia, Hongbo Ma, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
ICLR | 5 |
| 2025 | A Periodic Bayesian Flow for Material GenerationabstractGenerative modeling of crystal data distribution is an important yet challenging task due to the unique periodic physical symmetry of crystals. Diffusion-based methods have shown early promise in modeling crystal distribution. More recently, Bayesian Flow Networks were introduced to aggregate noisy latent variables, resulting in a variance-reduced parameter space that has been shown to be advantageous for modeling Euclidean data distributions with structural constraints (Song, et al.,2023). Inspired by this, we seek to unlock its potential for modeling variables located in non-Euclidean manifolds e.g. those within crystal structures, by overcoming challenging theoretical issues. We introduce CrysBFN, a novel crystal generation method by proposing a periodic Bayesian flow, which essentially differs from the original Gaussian-based BFN by exhibiting non-monotonic entropy dynamics. To successfully realize the concept of periodic Bayesian flow, CrysBFN integrates a new entropy conditioning mechanism and empirically demonstrates its significance compared to time-conditioning. Extensive experiments over both crystal ab initio generation and crystal structure prediction tasks demonstrate the superiority of CrysBFN, which consistently achieves new state-of-the-art on all benchmarks. Surprisingly, we found that CrysBFN enjoys a significant improvement in sampling efficiency, e.g., 200x speedup (10 v.s. 2000 steps network forwards) compared with previous Diffusion-based methods on MP-20 dataset. Yuxuan Song 0002, Jingjing Gong, Ziyao Cao, Yawen Ouyang, Hao Zhou 0012, Wei-Ying Ma |
ICLR | 8 |
| 2025 | Piloting Structure-Based Drug Design via Modality-Specific Optimal ScheduleabstractStructure-Based Drug Design (SBDD) is crucial for identifying bioactive molecules. Recent deep generative models are faced with challenges in geometric structure modeling. A major bottleneck lies in the twisted probability path of multi-modalities—continuous 3D positions and discrete 2D topologies—which jointly determine molecular geometries. By establishing the fact that noise schedules decide the Variational Lower Bound (VLB) for the twisted probability path, we propose VLB-Optimal Scheduling (VOS) strategy in this under-explored area, which optimizes VLB as a path integral for SBDD. Our model effectively enhances molecular geometries and interaction modeling, achieving state-of-the-art PoseBusters passing rate of 95.9\% on CrossDock, more than 10\% improvement upon strong baselines, while maintaining high affinities and robust intramolecular validity evaluated on held-out test set. Code is available at https://github.com/AlgoMole/MolCRAFT. Keyue Qiu, Yuxuan Song 0002, Zhehuan Fan, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma |
ICML | 8 |
| 2025 | Empower Structure-Based Molecule Optimization with Gradient Guided Bayesian Flow NetworksabstractStructure-based molecule optimization (SBMO) aims to optimize molecules with both continuous coordinates and discrete types against protein targets.
A promising direction is to exert gradient guidance on generative models given its remarkable success in images, but it is challenging to guide discrete data and risks inconsistencies between modalities.
To this end, we leverage a continuous and differentiable space derived through Bayesian inference, presenting Molecule Joint Optimization (MolJO), the gradient-based SBMO framework that facilitates joint guidance signals across different modalities while preserving SE(3)-equivariance.
We introduce a novel backward correction strategy that optimizes within a sliding window of the past histories, allowing for a seamless trade-off between explore-and-exploit during optimization.
MolJO achieves state-of-the-art performance on CrossDocked2020 benchmark (Success Rate 51.3\%, Vina Dock -9.05 and SA 0.78), more than 4x improvement in Success Rate compared to the gradient-based counterpart, and 2x ``Me-Better'' Ratio as much as 3D baselines.
Furthermore, we extend MolJO to a wide range of optimization settings, including multi-objective optimization and challenging tasks in drug design such as R-group optimization and scaffold hopping, further underscoring its versatility.
Code is available at https://github.com/AlgoMole/MolCRAFT. Keyue Qiu, Yuxuan Song 0002, Hongbo Ma, Ziyao Cao, Yushuai Wu, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma |
ICML | 10 |
| 2025 | Smooth Interpolation for Improved Discrete Graph Generative ModelsabstractThough typically represented by the discrete node and edge attributes, the graph topological information can be sufficiently captured by the graph spectrum in a continuous space. It is believed that incorporating the continuity of graph topological information into the generative process design could establish a superior paradigm for graph generative modeling. Motivated by such prior and recent advancements in the generative paradigm, we propose Graph Bayesian Flow Networks (GraphBFN) in this paper, a principled generative framework that designs an alternative generative process emphasizing the dynamics of topological information. Unlike recent discrete-diffusion-based methods, GraphBFNemploys the continuous counts derived from sampling infinite times from a categorical distribution as latent to facilitate a smooth decomposition of topological information, demonstrating enhanced effectiveness. To effectively realize the concept, we further develop an advanced sampling strategy and new time-scheduling techniques to overcome practical barriers and boost performance. Through extensive experimental validation on both generic graph and molecular graph generation tasks, GraphBFN could consistently achieve superior or competitive performance with significantly higher training and sampling efficiency. Yuxuan Song 0002, Juntong Shi, Jingjing Gong, Minkai Xu, Stefano Ermon, Hao Zhou 0012, Wei-Ying Ma |
ICML | 7 |
| 2025 | Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent SpaceabstractMedicinal chemists often optimize drugs considering their 3D structures and designing structurally distinct molecules that retain key features, such as shapes, pharmacophores, or chemical properties. Previous deep learning approaches address this through supervised tasks like molecule inpainting or property-guided optimization. In this work, we propose a flexible zero-shot molecule manipulation method by navigating in a shared latent space of 3D molecules. We introduce a Variational AutoEncoder (VAE) for 3D molecules, named MolFLAE, which learns a fixed-dimensional, E(3)-equivariant latent space independent of atom counts. MolFLAE encodes 3D molecules using an E(3)-equivariant neural network into fixed number of latent nodes, distinguished by learned embeddings. The latent space is regularized, and molecular structures are reconstructed via a Bayesian Flow Network (BFN) conditioned on the encoder’s latent output. MolFLAE achieves competitive performance on standard unconditional 3D molecule generation benchmarks. Moreover, the latent space of MolFLAE enables zero-shot molecule manipulation, including atom number editing, structure reconstruction, and coordinated latent interpolation for both structure and properties. We further demonstrate our approach on a drug optimization task for the human glucocorticoid receptor, generating molecules with improved hydrophilicity while preserving key interactions, under computational evaluations. These results highlight the flexibility, robustness, and real-world utility of our method, opening new avenues for molecule editing and optimization. Yinjun Jia, Zitong Tian, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 4 |
| 2025 | FIGRDock: Fast Interaction-Guided Regression for Flexible DockingabstractFlexible docking, which predicts the binding conformations of both proteins and small molecules by modeling their structural flexibility, plays a vital role in structure-based drug design. Although recent generative approaches, particularly diffusion-based models, have shown promising results, they require iterative sampling to generate candidate structures and depend on separate scoring functions for pose selection. This leads to an inefficient pipeline that is difficult to scale in real-world drug discovery workflows. To overcome these challenges, we introduce FIGRDock, a fast and accurate flexible docking framework that understands complicated interactions between molecules and proteins with a regression-based approach. FIGRDock leverages initial docking poses from conventional tools to distill interaction-aware distance patterns, which serve as explicit structural conditions to directly guide the prediction of the final protein-ligand complex via a regression model. This one-shot inference paradigm enables rapid and precise pose prediction without reliance on multi-step sampling or external scoring stages. Experimental results show that FIGRDock achieves up to 100× faster inference than diffusion-based docking methods, while consistently surpassing them in accuracy across standard benchmarks. These results suggest that FIGRDock has the potential to offer a scalable and efficient solution for flexible docking, advancing the pace of structure-based drug discovery. Shikun Feng, Bicheng Lin, Yuanhuan Mo, Yuyan Ni, Wenyu Zhu, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 7 |
| 2025 | CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMsabstractStructure-guided molecular generation is pivotal in early-stage drug discovery, enabling the design of compounds tailored to specific protein targets. However, despite recent advances in 3D generative modeling, particularly in improving docking scores, these methods often produce rare and intrinsically irrational molecular structures that deviate from drug-like chemical space. To quantify this issue, we propose a novel metric, the Molecule Reasonable Ratio (MRR), which measures structural rationality and reveals a critical gap between existing models and real-world approved drugs. To address this, we introduce the Collaborative Intelligence Drug Design (CIDD) framework, the first approach to unify the 3D interaction modeling capabilities of generative models with the general knowledge and reasoning power of large language models (LLMs). By leveraging LLM-based Chain-of-Thought reasoning, CIDD generates molecules that not only bind effectively to protein pockets but also exhibit strong structural drug-likeness, rationality, and synthetic accessibility. On the CrossDocked2020 benchmark, CIDD consistently improves drug-likeness metrics, including QED, SA, and MRR, across different base generative models, while maintaining competitive binding affinity. Notably, it raises the combined success rate (balancing drug-likeness and binding) from 15.72% to 34.59%, more than doubling previous results. These findings demonstrate the value of integrating knowledge reasoning with geometric generation to advance AI-driven drug design. Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Bowei He, Haichuan Tan, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
NeurIPS | 7 |
| 2025 | MOF-BFN: Metal-Organic Frameworks Structure Prediction via Bayesian Flow NetworksabstractMetal-Organic Frameworks (MOFs) have attracted considerable attention due to their unique properties including high surface area and tunable porosity, and promising applications in catalysis, gas storage, and drug delivery. Structure prediction for MOFs is a challenging task, as these frameworks are intrinsically periodic and hierarchically organized, where the entire structure is assembled from building blocks like metal nodes and organic linkers. To address this, we introduce MOF-BFN, a novel generative model for MOF structure prediction based on Bayesian Flow Networks (BFNs). Given the local geometry of building blocks, MOF-BFN jointly predicts the lattice parameters, as well as the positions and orientations of all building blocks within the unit cell. In particular, the positions are modelled in the fractional coordinate system to naturally incorporate the periodicity. Meanwhile, the orientations are modeled as unit quaternions sampled from learned Bingham distributions via the proposed Bingham BFN, enabling effective orientation generation on the 4D unit hypersphere. Experimental results demonstrate that MOF-BFN achieves state-of-the-art performance across multiple tasks, including structure prediction, geometric property evaluation, and de novo generation, offering a promising tool for designing complex MOF materials. Wenbing Huang 0001, Yuxuan Song 0002, Yawen Ouyang, Yu Rong 0001, Tingyang Xu, Hao Zhou 0012, Wei-Ying Ma, Yang Liu 0005 |
NeurIPS | 10 |
| 2025 | Retro-R1: LLM-based Agentic RetrosynthesisabstractRetrosynthetic planning is a fundamental task in chemical discovery. Due to the vast combinatorial search space, identifying viable synthetic routes remains a significant challenge--even for expert chemists. Recent advances in Large Language Models (LLMs), particularly equipped with reinforcement learning, have demonstrated strong human-like reasoning and planning abilities, especially in mathematics and code problem solving. This raises a natural question: Can the reasoning capabilities of LLMs be harnessed to develop an AI chemist capable of learning effective policies for multi-step retrosynthesis? In this study, we introduce Retro-R1, a novel LLM-based retrosynthesis agent trained via reinforcement learning to design molecular synthesis pathways. Unlike prior approaches, which typically rely on single-turn, question-answering formats, Retro-R1 interacts dynamically with plug-in single-step retrosynthesis tools and learns from environmental feedback. Experimental results show that Retro-R1 achieves a 55.79\% pass@1 success rate, surpassing the previous state of the art by 8.95\%. Notably, Retro-R1 demonstrates strong generalization to out-of-domain test cases, where existing methods tend to fail despite their high in-domain performance. Our work marks a significant step toward equipping LLMs with advanced, chemist-like reasoning abilities, highlighting the promise of reinforcement learning for enabling data-efficient, generalizable, and sophisticated scientific problem-solving in LLM-based agents. Jiangtao Feng, Hongli Yu, Yuxuan Song 0002, Shufei Zhang, Lei Bai 0001, Wei-Ying Ma, Hao Zhou 0012 |
NeurIPS | 8 |
| 2025 | Straight-Line Diffusion Model for Efficient 3D Molecular GenerationabstractDiffusion-based models have shown great promise in molecular generation but often require a large number of sampling steps to generate valid samples. In this paper, we introduce a novel Straight-Line Diffusion Model (SLDM) to tackle this problem, by formulating the diffusion process to follow a linear trajectory. The proposed process aligns well with the noise sensitivity characteristic of molecular structures and uniformly distributes reconstruction effort across the generative process, thus enhancing learning efficiency and efficacy. Consequently, SLDM achieves state-of-the-art performance on 3D molecule generation benchmarks, delivering a 100-fold improvement in sampling efficiency. Yuyan Ni, Shikun Feng, Haohan Chi, Huan-ang Gao, Wei-Ying Ma, Zhiming Ma, Yanyan Lan |
NeurIPS | 6 |
| 2025 | ShortListing Model: A Streamlined Simplex Diffusion for Discrete Variable GenerationabstractGenerative modeling of discrete variables is challenging yet crucial for applications in natural language processing and biological sequence design. We introduce the Shortlisting Model (SLM), a novel simplex-based diffusion model inspired by progressive candidate pruning. SLM operates on simplex centroids, reducing generation complexity and enhancing scalability. Additionally, SLM incorporates a flexible implementation of classifier-free guidance, enhancing unconditional generation performance. Extensive experiments on DNA promoter and enhancer design, protein design, character-level and large-vocabulary language modeling demonstrate the competitive performance and strong potential of SLM. Our code can be found at https://github.com/GenSI-THUAIR/SLM. Yuxuan Song 0002, Jingjing Gong, Qiying Yu, Zheng Zhang 0001, Mingxuan Wang, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 10 |
| 2025 | Rationalized All-Atom Protein Design with Unified Multi-Modal Bayesian FlowabstractDesigning functional proteins is a critical yet challenging problem due to the intricate interplay between backbone structures, sequences, and side-chains. Current approaches often decompose protein design into separate tasks, which can lead to accumulated errors, while recent efforts increasingly focus on all-atom protein design. However, we observe that existing all-atom generation approaches suffering from an information shortcut issue, where models inadvertently infer sequences from side-chain information, compromising their ability to accurately learn sequence distributions. To address this, we introduce a novel rationalized information flow strategy to eliminate the information shortcut. Furthermore, motivated by the advantages of Bayesian flows over differential equation–based methods, we propose the first Bayesian flow formulation for protein backbone orientations by recasting orientation modeling as an equivalent hyperspherical generation problem with antipodal symmetry. To validate, our method delivers consistently exceptional performance in both peptide and antibody design tasks. Yuxuan Song 0002, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 6 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 32 |
| 2025 | Accelerating 3D Molecule Generative Models with Trajectory DiagnosisabstractGeometric molecule generative models have found expanding applications across various scientific domains, but their generation inefficiency has become a critical bottleneck. Through a systematic investigation of the generative trajectory, we discover a unique challenge for molecule geometric graph generation: generative models require determining the permutation order of atoms in the molecule before refining its atomic feature values. Based on this insight, we decompose the generation process into permutation phase and adjustment phase, and propose a geometric-informed prior and consistency parameter objective to accelerate each phase. Extensive experiments demonstrate that our approach achieves competitive performance with approximately 10 sampling steps, 7.5 × faster than previous state-of-the-art models and approximately 100 × faster than diffusion-based models, offering a significant step towards scalable molecular generation. Yuxuan Song 0002, Jingjing Gong, Dongzhan Zhou, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 8 |
| 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and AggregationabstractVirtual screening (VS) is a critical component of modern drug discovery, yet most existing methods—whether physics-based or deep learning-based—are developed around *holo* protein structures with known ligand-bound pockets. Consequently, their performance degrades significantly on *apo* or predicted structures such as those from AlphaFold2, which are more representative of real-world early-stage drug discovery, where pocket information is often missing. In this paper, we introduce an alignment-and-aggregation framework to enable accurate virtual screening under structural uncertainty. Our method comprises two core components: (1) a tri-modal contrastive learning module that aligns representations of the ligand, the *holo* pocket, and cavities detected from structures, thereby enhancing robustness to pocket localization error; and (2) a cross-attention based adapter for dynamically aggregating candidate binding sites, enabling the model to learn from activity data even without precise pocket annotations. We evaluated our method on a newly curated benchmark of *apo* structures, where it significantly outperforms state-of-the-art methods in blind apo setting, improving the early enrichment factor (EF1\%) from 11.75 to 37.19. Notably, it also maintains strong performance on *holo* structures. These results demonstrate the promise of our approach in advancing first-in-class drug discovery, particularly in scenarios lacking experimentally resolved protein-ligand complexes. Our implementation is publicly available at [https://github.com/Wiley-Z/AANet](https://github.com/Wiley-Z/AANet). Wenyu Zhu, Yinjun Jia, Haichuan Tan, Ya-Qin Zhang, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 7 |
| 2025 | Artificial intelligence without restriction surpassing human intelligence with probability one: Theoretical insight into secrets of the brain with AI twins of the brain
Guang-Bin Huang, M. Brandon Westover, Eng-King Tan, Dongshun Cui, Wei-Ying Ma, Tiantong Wang, Haikun Wei, Qiyuan Tian, Kwok-Yan Lam, Tien Yin Wong |
Neurocomputing | 6 |
| 2024 | Protein-ligand binding representation learning from fine-grained interactionsabstractThe binding between proteins and ligands plays a crucial role in the realm of drug discovery. Previous deep learning approaches have shown promising results over traditional computationally intensive methods, but resulting in poor generalization due to limited supervised data. In this paper, we propose to learn protein-ligand binding representation in a self-supervised learning manner. Different from existing pre-training approaches which treat proteins and ligands individually, we emphasize to discern the intricate binding patterns from fine-grained interactions. Specifically, this self-supervised learning problem is formulated as a prediction of the conclusive binding complex structure given a pocket and ligand with a Transformer based interaction module, which naturally emulates the binding process. To ensure the representation of rich binding information, we introduce two pre-training tasks, i.e. atomic pairwise distance map prediction and mask ligand reconstruction, which comprehensively model the fine-grained interactions from both structure and feature space. Extensive experiments have demonstrated the superiority of our method across various binding tasks, including protein-ligand affinity prediction, virtual screening and protein-ligand docking. Shikun Feng, Yinjun Jia, Wei-Ying Ma, Yanyan Lan |
ICLR | 4 |
| 2024 | Self-supervised Pocket Pretraining via Protein Fragment-Surroundings AlignmentabstractPocket representations play a vital role in various biomedical applications, such as druggability estimation, ligand affinity prediction, and de novo drug design. While existing geometric features and pretrained representations have demonstrated promising results, they usually treat pockets independent of ligands, neglecting the fundamental interactions between them. However, the limited pocket-ligand complex structures available in the PDB database (less than 100 thousand non-redundant pairs) hampers large-scale pretraining endeavors for interaction modeling. To address this constraint, we propose a novel pocket pretraining approach that leverages knowledge from high-resolution atomic protein structures, assisted by highly effective pretrained small molecule representations. By segmenting protein structures into drug-like fragments and their corresponding pockets, we obtain a reasonable simulation of ligand-receptor interactions, resulting in the generation of over 5 million complexes. Subsequently, the pocket encoder is trained in a contrastive manner to align with the representation of pseudo-ligand furnished by some pretrained small molecule encoders. Our method, named ProFSA, achieves state-of-the-art performance across various tasks, including pocket druggability prediction, pocket matching, and ligand binding affinity prediction. Notably, ProFSA surpasses other pretraining methods by a substantial margin. Moreover, our work opens up a new avenue for mitigating the scarcity of protein-ligand complex data through the utilization of high-quality and diverse protein structure databases. Yinjun Jia, Yuanle Mo, Yuyan Ni, Wei-Ying Ma, Zhiming Ma, Yanyan Lan |
ICLR | 5 |
| 2024 | Sliced Denoising: A Physics-Informed Molecular Pre-Training MethodabstractWhile molecular pre-training has shown great potential in enhancing drug discovery, the lack of a solid physical interpretation in current methods raises concerns about whether the learned representation truly captures the underlying explanatory factors in observed data, ultimately resulting in limited generalization and robustness. Although denoising methods offer a physical interpretation, their accuracy is often compromised by ad-hoc noise design, leading to inaccurate learned force fields. To address this limitation, this paper proposes a new method for molecular pre-training, called sliced denoising (SliDe), which is based on the classical mechanical intramolecular potential theory. SliDe utilizes a novel noise strategy that perturbs bond lengths, angles, and torsion angles to achieve better sampling over conformations. Additionally, it introduces a random slicing approach that circumvents the computationally expensive calculation of the Jacobian matrix, which is otherwise essential for estimating the force field. By aligning with physical principles, SliDe shows a 42\% improvement in the accuracy of estimated force fields compared to current state-of-the-art denoising methods, and thus outperforms traditional baselines on various molecular property prediction tasks. Yuyan Ni, Shikun Feng, Wei-Ying Ma, Zhiming Ma, Yanyan Lan |
ICLR | 3 |
| 2024 | Unified Generative Modeling of 3D Molecules with Bayesian Flow NetworksabstractAdvanced generative model (\textit{e.g.}, diffusion model) derived from simplified continuity assumptions of data distribution, though showing promising progress, has been difficult to apply directly to geometry generation applications due to the \textit{multi-modality} and \textit{noise-sensitive} nature of molecule geometry.
This work introduces Geometric Bayesian Flow Networks (GeoBFN), which naturally fits molecule geometry by modeling diverse modalities in the differentiable parameter space of distributions. GeoBFN maintains the SE-(3) invariant density modeling property by incorporating equivariant inter-dependency modeling on parameters of distributions and unifying the probabilistic modeling of different modalities.
Through optimized training and sampling techniques, we demonstrate that GeoBFN achieves state-of-the-art performance on multiple 3D molecule generation benchmarks in terms of generation quality (90.87\% molecule stability in QM9 and 85.6\% atom stability in GEOM-DRUG\footnote{The scores are reported at 1k sampling steps for fair comparison, and our scores could be further improved if sampling sufficiently longer steps.}). GeoBFN can also conduct sampling with any number of steps to reach an optimal trade-off between efficiency and quality (\textit{e.g.}, 20$\times$ speedup without sacrificing performance). Yuxuan Song 0002, Jingjing Gong, Hao Zhou 0012, Mingyue Zheng, Wei-Ying Ma |
ICLR | 6 |
| 2024 | UniCorn: A Unified Contrastive Learning Approach for Multi-view Molecular Representation LearningabstractRecently, a noticeable trend has emerged in developing pre-trained foundation models in the domains of CV and NLP. However, for molecular pre-training, there lacks a universal model capable of effectively applying to various categories of molecular tasks, since existing prevalent pre-training methods exhibit effectiveness for specific types of downstream tasks. Furthermore, the lack of profound understanding of existing pre-training methods, including 2D graph masking, 2D-3D contrastive learning, and 3D denoising, hampers the advancement of molecular foundation models. In this work, we provide a unified comprehension of existing pre-training methods through the lens of contrastive learning. Thus their distinctions lie in clustering different views of molecules, which is shown beneficial to specific downstream tasks. To achieve a complete and general-purpose molecular representation, we propose a novel pre-training framework, named UniCorn, that inherits the merits of the three methods, depicting molecular views in three different levels. SOTA performance across quantum, physicochemical, and biological tasks, along with comprehensive ablation study, validate the universality and effectiveness of UniCorn. Shikun Feng, Yuyan Ni, Yanwen Huang, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
ICML | 6 |
| 2024 | Rethinking Specificity in SBDD: Leveraging Delta Score and Energy-Guided DiffusionabstractIn the field of Structure-based Drug Design (SBDD), deep learning-based generative models have achieved outstanding performance in terms of docking score. However, further study shows that the existing molecular generative methods and docking scores both have lacked consideration in terms of specificity, which means that generated molecules bind to almost every protein pocket with high affinity. To address this, we introduce the Delta Score, a new metric for evaluating the specificity of molecular binding. To further incorporate this insight for generation, we develop an innovative energy-guided approach using contrastive learning, with active compounds as decoys, to direct generative models toward creating molecules with high specificity. Our empirical results show that this method not only enhances the delta score but also maintains or improves traditional docking scores, successfully bridging the gap between SBDD and real-world needs. Minsi Ren, Yuyan Ni, Yanwen Huang, Bo Qiang, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
ICML | 7 |
| 2024 | MolCRAFT: Structure-Based Drug Design in Continuous Parameter SpaceabstractGenerative models for structure-based drug design (SBDD) have shown promising results in recent years. Existing works mainly focus on how to generate molecules with higher binding affinity, ignoring the feasibility prerequisites for generated 3D poses and resulting in false positives. We conduct thorough studies on key factors of ill-conformational problems when applying autoregressive methods and diffusion to SBDD, including mode collapse and hybrid continuous-discrete space. In this paper, we introduce MolCRAFT, the first SBDD model that operates in the continuous parameter space, together with a novel noise reduced sampling strategy. Empirical results show that our model consistently achieves superior performance in binding affinity with more stable 3D structure, demonstrating our ability to accurately model interatomic interactions. To our best knowledge, MolCRAFT is the first to achieve reference-level Vina Scores (-6.59 kcal/mol) with comparable molecular size, outperforming other strong baselines by a wide margin (-0.84 kcal/mol). Code is available at https://github.com/AlgoMole/MolCRAFT. Yanru Qu, Keyue Qiu, Yuxuan Song 0002, Jingjing Gong, Jiawei Han 0001, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma |
ICML | 8 |
| 2024 | Mol-AE: Auto-Encoder Based Molecular Representation Learning With 3D Cloze Test Objectiveabstract3D molecular representation learning has gained tremendous interest and achieved promising performance in various downstream tasks. A series of recent approaches follow a prevalent framework: an encoder-only model coupled with a coordinate denoising objective. However, through a series of analytical experiments, we prove that the encoder-only model with coordinate denoising objective exhibits inconsistency between pre-training and downstream objectives, as well as issues with disrupted atomic identifiers. To address these two issues, we propose Mol-AE for molecular representation learning, an auto-encoder model using positional encoding as atomic identifiers. We also propose a new training objective named 3D Cloze Test to make the model learn better atom spatial relationships from real molecular substructures. Empirical results demonstrate that Mol-AE achieves a large margin performance gain compared to the current state-of-the-art 3D molecular modeling approach. Kangjie Zheng, Siyu Long, Zaiqing Nie, Ming Zhang 0004, Xinyu Dai, Wei-Ying Ma, Hao Zhou 0012 |
ICML | 7 |
| 2024 | ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular ModelingabstractProtein language models have demonstrated significant potential in the field of protein engineering. However, current protein language models primarily operate at the residue scale, which limits their ability to provide information at the atom level. This limitation prevents us from fully exploiting the capabilities of protein language models for applications involving both proteins and small molecules. In this paper, we propose ESM-AA (ESM All-Atom), a novel approach that enables atom-scale and residue-scale unified molecular modeling. ESM-AA achieves this by pre-training on multi-scale code-switch protein sequences and utilizing a multi-scale position encoding to capture relationships among residues and atoms. Experimental results indicate that ESM-AA surpasses previous methods in protein-molecule tasks, demonstrating the full utilization of protein language models. Further investigations reveal that through unified molecular modeling, ESM-AA not only gains molecular knowledge but also retains its understanding of proteins. Kangjie Zheng, Siyu Long, Tianyu Lu, Xinyu Dai, Ming Zhang 0004, Zaiqing Nie, Wei-Ying Ma, Hao Zhou 0012 |
ICML | 8 |
| 2023 | Fractional Denoising for 3D Molecular Pre-trainingabstractCoordinate denoising is a promising 3D molecular pre-training method, which has achieved remarkable performance in various downstream drug discovery tasks. Theoretically, the objective is equivalent to learning the force field, which is revealed helpful for downstream tasks. Nevertheless, there are two challenges for coordinate denoising to learn an effective force field, i.e. low coverage samples and isotropic force field. The underlying reason is that molecular distributions assumed by existing denoising methods fail to capture the anisotropic characteristic of molecules. To tackle these challenges, we propose a novel hybrid noise strategy, including noises on both dihedral angel and coordinate. However, denoising such hybrid noise in a traditional way is no more equivalent to learning the force field. Through theoretical deductions, we find that the problem is caused by the dependency of the input conformation for covariance. To this end, we propose to decouple the two types of noise and design a novel fractional denoising method (Frad), which only denoises the latter coordinate part. In this way, Frad enjoys both the merits of sampling more low-energy structures and the force field equivalence. Extensive experiments show the effectiveness of Frad in molecule representation, with a new state-of-the-art on 9 out of 12 tasks of QM9 and on 7 out of 8 targets of MD17. Shikun Feng, Yuyan Ni, Yanyan Lan, Zhiming Ma, Wei-Ying Ma |
ICML | 5 |
| 2023 | Coarse-to-Fine: a Hierarchical Diffusion Model for Molecule Generation in 3DabstractGenerating desirable molecular structures in 3D is a fundamental problem for drug discovery. Despite the considerable progress we have achieved, existing methods usually generate molecules in atom resolution and ignore intrinsic local structures such as rings, which leads to poor quality in generated structures, especially when generating large molecules. Fragment-based molecule generation is a promising strategy, however, it is nontrivial to be adapted for 3D non-autoregressive generations because of the combinational optimization problems. In this paper, we utilize a coarse-to-fine strategy to tackle this problem, in which a Hierarchical Diffusion-based model (i.e. HierDiff) is proposed to preserve the validity of local segments without relying on autoregressive modeling. Specifically, HierDiff first generates coarse-grained molecule geometries via an equivariant diffusion process, where each coarse-grained node reflects a fragment in a molecule. Then the coarse-grained nodes are decoded into fine-grained fragments by a message-passing process and a newly designed iterative refined sampling module. Lastly, the fine-grained fragments are then assembled to derive a complete atomic molecular structure. Extensive experiments demonstrate that HierDiff consistently improves the quality of molecule generation over existing methods. Bo Qiang, Yuxuan Song 0002, Minkai Xu, Jingjing Gong, Hao Zhou 0012, Wei-Ying Ma, Yanyan Lan |
ICML | 7 |
| 2023 | DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening
Bo Qiang, Haichuan Tan, Yinjun Jia, Minsi Ren, Minsi Lu, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 8 |
| 2023 | Equivariant Flow Matching with Hybrid Probability Transport for 3D Molecule GenerationabstractThe generation of 3D molecules requires simultaneously deciding the categorical features (atom types) and continuous features (atom coordinates). Deep generative models, especially Diffusion Models (DMs), have demonstrated effectiveness in generating feature-rich geometries. However, existing DMs typically suffer from unstable probability dynamics with inefficient sampling speed. In this paper, we introduce geometric flow matching, which enjoys the advantages of both equivariant modeling and stabilized probability dynamics.
More specifically, we propose a hybrid probability path where the coordinates probability path is regularized by an equivariant optimal transport, and the information between different modalities is aligned. Experimentally, the proposed method could consistently achieve better performance on multiple molecule generation benchmarks with 4.75$\times$ speed up of sampling on average. Yuxuan Song 0002, Jingjing Gong, Minkai Xu, Ziyao Cao, Yanyan Lan, Stefano Ermon, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 8 |
| 2023 | Controllable Image Synthesis With Attribute-Decomposed GANabstractThis paper proposes Attribute-Decomposed GAN (ADGAN) and its enhanced version (ADGAN++) for controllable image synthesis, which can produce realistic images with desired attributes provided in various source inputs. The core ideas of the proposed ADGAN and ADGAN++ are both to embed component attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. The major difference between them is that ADGAN processes all component attributes simultaneously while ADGAN++ utilizes a serial encoding strategy. More specifically, ADGAN consists of two encoding pathways with style block connections and is capable of decomposing the original hard mapping into multiple more accessible subtasks. In the source pathway, component layouts are extracted via a semantic parser and the segmented components are fed into a shared global texture encoder to obtain decomposed latent codes. This strategy allows for the synthesis of more realistic output images and the automatic separation of un-annotated component attributes. Although the original ADGAN works in a delicate and efficient manner, intrinsically it fails to handle the semantic image synthesizing task when the number of attribute categories is huge. To address this problem, ADGAN++ employs the serial encoding of different component attributes to synthesize each part of the target real-world image, and adopts several residual blocks with segmentation guided instance normalization to assemble the synthesized component images and refine the original synthesis result. The two-stage ADGAN++ is designed to alleviate the massive computational costs required when synthesizing real-world images with numerous attributes while maintaining the disentanglement of different attributes to enable flexible control of arbitrary component attributes of the synthesized images. Experimental results demonstrate the proposed methods' superiority over the state of the art in pose transfer, face style transfer, and semantic image synthesis, as well as their effectiveness in the task of component attribute transfer. Our code and data are publicly available at https://github.com/menyifang/ADGAN. Guo Pu, Yifang Men, Yiming Mao 0006, Yuning Jiang 0001, Wei-Ying Ma, Zhouhui Lian |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | PEMP: Leveraging Physics Properties to Enhance Molecular Property PredictionabstractMolecular property prediction is essential for drug discovery. In recent years, deep learning methods have been introduced to this area and achieved state-of-the-art performances. However, most of existing methods ignore the intrinsic relations between molecular properties which can be utilized to improve the performances of corresponding prediction tasks. In this paper, we propose a new approach, namely Physics properties Enhanced Molecular Property prediction (PEMP), to utilize relations between molecular properties revealed by previous physics theory and physical chemistry studies. Specifically, we enhance the training of the chemical and physiological property predictors with related physics property prediction tasks. We design two different methods for PEMP, respectively based on multi-task learning and transfer learning. Both methods include a model-agnostic molecule representation module and a property prediction module. In our implementation, we adopt both the state-of-the-art molecule embedding models under the supervised learning paradigm and the pretraining paradigm as the molecule representation module of PEMP, respectively. Experimental results on public benchmark MoleculeNet show that the proposed methods have the ability to outperform corresponding state-of-the-art models. Yuancheng Sun, Weizhi Ma, Wenhao Huang 0001, Kang Liu 0001, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
CIKM | 7 |
| 2022 | Energy-Inspired Molecular Conformation Optimization
Jiaqi Guan, Wesley Wei Qian, Qiang Liu 0001, Wei-Ying Ma, Jianzhu Ma, Jian Peng 0001 |
ICLR | 4 |
| 2020 | Controllable Person Image Synthesis With Attribute-Decomposed GANabstractThis paper introduces the Attribute-Decomposed GAN, a novel generative model for controllable person image synthesis, which can produce realistic person images with desired human attributes (e.g., pose, head, upper clothes and pants) provided in various source inputs. The core idea of the proposed model is to embed human attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. Specifically, a new architecture consisting of two encoding pathways with style block connections is proposed to decompose the original hard mapping into multiple more accessible subtasks. In source pathway, we further extract component layouts with an off-the-shelf human parser and feed them into a shared global texture encoder for decomposed latent codes. This strategy allows for the synthesis of more realistic output images and automatic separation of un-annotated attributes. Experimental results demonstrate the proposed method's superiority over the state of the art in pose transfer and its effectiveness in the brand-new task of component attribute transfer. Yifang Men, Yiming Mao 0006, Yuning Jiang 0001, Wei-Ying Ma, Zhouhui Lian |
CVPR | 4 |
| 2020 | Democratizing Content Creation and Dissemination through AI TechnologyabstractWith the rise of mobile video, user-generated content, and social networks, there is a massive opportunity for disruptive innovations in the media and content industry. It is now a fast-changing landscape with rapid advances in AI-powered content creation, dissemination and interaction technologies. I believe the current trends are leading us towards a world where everyone is equally empowered to produce high-quality content in video, music, augmented reality or more – and to share their information, knowledge, and stories with a large global audience. This new AI- powered content platform can further lead to innovations in advertising, e-commerce, online education, and productivity. I will share the current research efforts at ByteDance connected to this emerging new platform through products such as Douyin and TikTok, and discuss the challenges and the direction of our future research. Wei-Ying Ma |
WWW | 1 |
| 2019 | What You Look Matters?: Offline Evaluation of Advertising Creatives for Cold-start ProblemabstractModern online auction-based advertising systems combine item and user features to promote ad creatives with the most revenue.However, new ad creatives have to display for certain initial users before enough click statistics could collected and utilized in later ads ranking and bidding processes. This leads to a well-known challenging cold start problem.In this paper, we argue that the content of the creatives intrinsically determines their performance (e.g. ctr, cvr), and we add a pre-ranking stage based on the content. The stage prunes inferior creatives and thus makes online impressions more effective. Since the pre-ranking stage can be executed offline, we can use deep features and take their well generalization to navigate the cold start problem.Specifically, we propose Pre Evaluation Ad Creation Model (PEAC), a novel method to evaluate creatives even before they were shown in the online ads system. Our proposed PEAC only utilizes ads information such as verbal and visual content, but requires no user data as features. During the online A/B testing, PEAC shows significant improvement in revenue. The method has been implemented and deployed in the large scale online advertising system at ByteDance. Furthermore, we provide detailed analysis on what the model learns, which also gives suggestions for ad creative design. Zhichen Zhao, Lei Li 0005, Bowen Zhang 0007, Yuning Jiang 0001, Fengkun Wang, Wei-Ying Ma |
CIKM | 8 |
| 2019 | Unified Visual-Semantic Embeddings: Bridging Vision and Language With Structured Meaning RepresentationsabstractWe propose the Unified Visual-Semantic Embeddings (Unified VSE) for learning a joint space of visual representation and textual semantics. The model unifies the embeddings of concepts at different levels: objects, attributes, relations, and full scenes. We view the sentential semantics as a combination of different semantic components such as objects and relations; their embeddings are aligned with different image regions. A contrastive learning approach is proposed for the effective learning of this fine-grained alignment from only image-caption pairs. We also present a simple yet effective approach that enforces the coverage of caption embeddings on the semantic components that appear in the sentence. We demonstrate that the Unified VSE outperforms baselines on cross-modal retrieval tasks; the enforcement of the semantic coverage improves the model's robustness in defending text-domain adversarial attacks. Moreover, our model empowers the use of visual cues to accurately resolve word dependencies in novel sentences. Hao Wu 0011, Jiayuan Mao, Yuning Jiang 0001, Lei Li 0005, Weiwei Sun 0008, Wei-Ying Ma |
CVPR | 7 |
| 2018 | Automatic Data Augmentation from Massive Web Images for Deep Visual RecognitionabstractLarge-scale image datasets and deep convolutional neural networks (DCNNs) are the two primary driving forces for the rapid progress in generic object recognition tasks in recent years. While lots of network architectures have been continuously designed to pursue lower error rates, few efforts are devoted to enlarging existing datasets due to high labeling costs and unfair comparison issues. In this article, we aim to achieve lower error rates by augmenting existing datasets in an automatic manner. Our method leverages both the web and DCNN, where the web provides massive images with rich contextual information, and DCNN replaces humans to automatically label images under the guidance of web contextual information. Experiments show that our method can automatically scale up existing datasets significantly from billions of web pages with high accuracy. The performance on object recognition tasks and transfer learning tasks have been significantly improved by using the automatically augmented datasets, which demonstrates that more supervisory information has been automatically gathered from the web. Both the dataset and models trained on the dataset have been made publicly available. Yalong Bai, Kuiyuan Yang, Tao Mei 0001, Wei-Ying Ma, Tiejun Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Topic Aware Neural Response GenerationabstractWe consider incorporating topic information into a sequence-to-sequence framework to generate informative and interesting responses for chatbots. To this end, we propose a topic aware sequence-to-sequence (TA-Seq2Seq) model. The model utilizes topics to simulate prior human knowledge that guides them to form informative and interesting responses in conversation, and leverages topic information in generation by a joint attention mechanism and a biased generation probability. The joint attention mechanism summarizes the hidden vectors of an input message as context vectors by message attention and synthesizes topic vectors by topic attention from the topic words of the message obtained from a pre-trained LDA model, with these vectors jointly affecting the generation of words in decoding. To increase the possibility of topic words appearing in responses, the model modifies the generation probability of topic words by adding an extra probability item to bias the overall distribution. Empirical studies on both automatic evaluation metrics and human annotations show that TA-Seq2Seq can generate more informative and interesting responses, significantly outperforming state-of-the-art response generation models. Chen Xing, Wei Wu 0014, Yu Wu 0012, Jie Liu 0007, Yalou Huang, Ming Zhou 0001, Wei-Ying Ma |
AAAI | 7 |
| 2017 | Beyond the Words: Predicting User Personality from Heterogeneous InformationabstractAn incisive understanding of user personality is not only essential to many scientific disciplines, but also has a profound business impact on practical applications such as digital marketing, personalized recommendation, mental diagnosis, and human resources management. Previous studies have demonstrated that language usage in social media is effective in personality prediction. However, except for single language features, a less researched direction is how to leverage the heterogeneous information on social media to have a better understanding of user personality. In this paper, we propose a Heterogeneous Information Ensemble framework, called HIE, to predict users' personality traits by integrating heterogeneous information including self-language usage, avatar, emoticon, and responsive patterns. In our framework, to improve the performance of personality prediction, we have designed different strategies extracting semantic representations to fully leverage heterogeneous information on social media. We evaluate our methods with extensive experiments based on a real-world data covering both personality survey results and social media usage from thousands of volunteers. The results reveal that our approaches significantly outperform several widely adopted state-of-the-art baseline methods. To figure out the utility of HIE in a real-world interactive setting, we also present DiPsy, a personalized chatbot to predict user personality through heterogeneous information in digital traces and conversation logs. Honghao Wei, Nicholas Jing Yuan, Chuan Cao, Hao Fu 0015, Xing Xie 0001, Yong Rui, Wei-Ying Ma |
WSDM | 8 |
| 2016 | Hashtag-Based Sub-Event Discovery Using Mutually Generative LDA in TwitterabstractSub-event discovery is an effective method for social event analysis in Twitter. It can discover sub-events from large amount of noisy event-related information in Twitter and semantically represent them. The task is challenging because tweets are short, informal and noisy. To solve this problem, we consider leveraging event-related hashtags that contain many locations, dates and concise sub-event related descriptions to enhance sub-event discovery. To this end, we propose a hashtag-based mutually generative Latent Dirichlet Allocation model(MGe-LDA). In MGe-LDA, hashtags and topics of a tweet are mutually generated by each other. The mutually generative process models the relationship between hashtags and topics of tweets, and highlights the role of hashtags as a semantic representation of the corresponding tweets. Experimental results show that MGe-LDA can significantly outperform state-of-the-art methods for sub-event discovery. Chen Xing, Jie Liu 0007, Yalou Huang, Wei-Ying Ma |
AAAI | 5 |
| 2016 | How well do Computers Solve Math Word Problems? Large-Scale Dataset Construction and EvaluationabstractRecently a few systems for automatically solving math word problems have reported promising results. However, the datasets used for evaluation have limitations in both scale and diversity. In this paper, we build a large-scale dataset which is more than 9 times the size of previous ones, and contains many more problem types. Problems in the dataset are semi-automatically obtained from community question-answering (CQA) web pages. A ranking SVM model is trained to automatically extract problem answers from the answer text provided by CQA users, which significantly reduces human annotation cost. Experiments conducted on the new dataset lead to interesting and surprising results. Danqing Huang, Shuming Shi 0001, Chin-Yew Lin, Jian Yin 0001, Wei-Ying Ma |
ACL (1) | 5 |
| 2016 | Learning to Extract Conditional Knowledge for Question Answering using DialogueabstractKnowledge based question answering (KBQA) has attracted much attention from both academia and industry in the field of Artificial Intelligence. However, many existing knowledge bases (KBs) are built by static triples. It is hard to answer user questions with different conditions, which will lead to significant answer variances in questions with similar intent. In this work, we propose to extract conditional knowledge base (CKB) from user question-answer pairs for answering user questions with different conditions through dialogue. Given a subject, we first learn user question patterns and conditions. Then we propose an embedding based co-clustering algorithm to simultaneously group the patterns and conditions by leveraging the answers as supervisor information. After that, we extract the answers to questions conditioned on both question pattern clusters and condition clusters as a CKB. As a result, when users ask a question without clearly specifying the conditions, we use dialogues in natural language to chat with users for question specification and answer retrieval. Experiments on real question answering (QA) data show that the dialogue model using automatically extracted CKB can more accurately answer user questions and significantly improve user satisfaction for questions with missing conditions. Pengwei Wang 0004, Lei Ji 0001, Jun Yan 0001, Wei-Ying Ma |
CIKM | 5 |
| 2016 | Collaborative Knowledge Base Embedding for Recommender SystemsabstractAmong different recommendation techniques, collaborative filtering usually suffer from limited performance due to the sparsity of user-item interactions. To address the issues, auxiliary information is usually used to boost the performance. Due to the rapid collection of information on the web, the knowledge base provides heterogeneous information including both structured and unstructured data with different semantics, which can be consumed by various applications. In this paper, we investigate how to leverage the heterogeneous information in a knowledge base to improve the quality of recommender systems. First, by exploiting the knowledge base, we design three components to extract items' semantic representations from structural content, textual content and visual content, respectively. To be specific, we adopt a heterogeneous network embedding method, termed as TransR, to extract items' structural representations by considering the heterogeneity of both nodes and relationships. We apply stacked denoising auto-encoders and stacked convolutional auto-encoders, which are two types of deep learning based embedding techniques, to extract items' textual representations and visual representations, respectively. Finally, we propose our final integrated framework, which is termed as Collaborative Knowledge Base Embedding (CKE), to jointly learn the latent representations in collaborative filtering as well as items' semantic representations from the knowledge base. To evaluate the performance of each embedding component as well as the whole system, we conduct extensive experiments with two real-world datasets from different scenarios. The results reveal that our approaches outperform several widely adopted state-of-the-art recommendation methods. Nicholas Jing Yuan, Defu Lian, Xing Xie 0001, Wei-Ying Ma |
KDD | 5 |
| 2016 | Dual Learning for Machine TranslationabstractWhile neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this training data bottleneck, we develop a dual-learning mechanism, which can enable an NMT system to automatically learn from unlabeled data through a dual-learning game. This mechanism is inspired by the following observation: any machine translation task has a dual task, e.g., English-to-French translation (primal) versus French-to-English translation (dual); the primal and dual tasks can form a closed loop, and generate informative feedback signals to train the translation models, even if without the involvement of a human labeler. In the dual-learning mechanism, we use one agent to represent the model for the primal task and the other agent to represent the model for the dual task, then ask them to teach each other through a reinforcement learning process. Based on the feedback signals generated during this process (e.g., the language-model likelihood of the output of a model, and the reconstruction error of the original sentence after the primal and dual translations), we can iteratively update the two models until convergence (e.g., using the policy gradient methods). We call the corresponding approach to neural machine translation \emph{dual-NMT}. Experiments show that dual-NMT works very well on English$\leftrightarrow$French translation; especially, by learning from monolingual data (with 10\% bilingual data for warm start), it achieves a comparable accuracy to NMT trained from the full bilingual data for the French-to-English translation task. Di He 0001, Yingce Xia, Tao Qin 0001, Liwei Wang 0001, Nenghai Yu, Tie-Yan Liu, Wei-Ying Ma |
NIPS | 7 |
| 2015 | Automatic Image Dataset Construction from Click-through Logs Using Deep Neural NetworkabstractLabelled image datasets are the backbone for high-level image understanding tasks with wide application scenarios, and continuously drive and evaluate the progress of feature designing and supervised learning models. Recently, the million scale labelled image dataset further contributes to the rebirth of deep convolutional neural network and bypass manual designing handcraft features. However, the construction process of image dataset is mainly manual-based and quite labor intensive, which often take years' efforts to construct a million scale dataset with high quality. In this paper, we propose a deep learning based method to construct large scale image dataset in an automatic way. Specifically, word representation and image representation are learned in a deep neural network from large amount of click-through logs, and further used to define word-word similarity and image-word similarity. These two similarities are used to automatize the two labor intensive steps in manual-based image dataset construction: query formation and noisy image removal. With a new proposed cross convolutional filter regularizer, we can construct a million scale image dataset in one week. Finally, two image datasets are constructed to verify the effectiveness of the method. In addition to scale, the automatically constructed dataset has comparable accuracy, diversity and cross-dataset generalization with manually labelled image datasets. Yalong Bai, Kuiyuan Yang, Wei Yu 0004, Chang Xu 0008, Wei-Ying Ma, Tiejun Zhao |
ACM Multimedia | 5 |
| 2015 | LightLDA: Big Topic Models on Modest Computer ClustersabstractWhen building large-scale machine learning (ML) programs, such as massive topic models or deep neural networks with up to trillions of parameters and training examples, one usually assumes that such massive tasks can only be attempted with industrial-sized clusters with thousands of nodes, which are out of reach for most practitioners and academic researchers. We consider this challenge in the context of topic modeling on web-scale corpora, and show that with a modest cluster of as few as 8 machines, we can train a topic model with 1 million topics and a 1-million-word vocabulary (for a total of 1 trillion parameters), on a document collection with 200 billion tokens --- a scale not yet reported even with thousands of machines. Our major contributions include: 1) a new, highly-efficient O(1) Metropolis-Hastings sampling algorithm, whose running cost is (surprisingly) agnostic of model size, and empirically converges nearly an order of magnitude more quickly than current state-of-the-art Gibbs samplers; 2) a model-scheduling scheme to handle the big model challenge, where each worker machine schedules the fetch/use of sub-models as needed, resulting in a frugal use of limited memory capacity and network bandwidth; 3) a differential data-structure for model storage, which uses separate data structures for high- and low-frequency words to allow extremely large models to fit in memory, while maintaining high inference speed. These contributions are built on top of the Petuum open-source distributed ML framework, and we provide experimental evidence showing how this development puts massive data and models within reach on a small cluster, while still enjoying proportional time cost reductions with increasing cluster size. Jinhui Yuan, Fei Gao 0018, Qirong Ho, Wei Dai 0003, Jinliang Wei, Xun Zheng, Eric P. Xing, Tie-Yan Liu, Wei-Ying Ma |
WWW | 9 |
| 2014 | Indoor air quality monitoring system for smart buildingsabstractMany developing countries are suffering from air pollution, especially the Particulate Matter with diameter of 2.5 micrometers or less (PM2.5). While quite a few air quality monitoring stations have been built by governments in a city's public areas, the indoor PM2.5 has not yet been monitored and dealt with effectively. Though many office buildings have an HVAC (heating, ventilation, and air conditioning) system, PM2.5 is not considered as a factor when the system circulates fresh air from outdoors. This paper introduces a real system that we have deployed in the offices of four Microsoft campuses in China. This system instantly monitors indoor air quality on different floors of a building (including office areas, gyms, garages, and restaurants), enabling Microsoft employees to enquire the air quality of a place by using a mobile phone or checking a website. The information can guide a user's decision making, e.g., finding the right time to work out in the gym or turn on individual air filters in her own office. Through analyzing the indoor and outdoor air quality data collected over a long period, our system can even offer actionable and energy-efficient suggestion to HVAC systems, e.g., automatically turning on the system only a few hours earlier than usual if it is a heavily polluted day, or identifying the filters in HVAC system that should be renewed. Xuxu Chen, Yu Zheng 0004, Yubiao Chen, Qiwei Jin, Weiwei Sun 0008, Eric Chang, Wei-Ying Ma |
UbiComp | 7 |
| 2014 | Bag-of-Words Based Deep Neural Network for Image RetrievalabstractThis work targets image retrieval task hold by MSR-Bing Grand Challenge. Image retrieval is considered as a challenge task because of the gap between low-level image representation and high-level textual query representation. Recently further developed deep neural network sheds light on narrowing the gap by learning high-level image representation from raw pixels. In this paper, we proposed a bag-of-words based deep neural network for image retrieval task, which learns high-level image representation and maps images into bag-of-words space. The DNN model is trained on the large scale clickthrough data, and the relevance between query and image is measured by the cosine similarity of query's bag-of-words representation and image's bag-of-words representation predicted by DNN, the visual similarity of images is computed by high-level image representation extracted via the DNN model too. Finally, PageRank algorithm is used to further improve the ranking list by considering visual similarity of images for each query. The experimental results achieved state-of-the-art performance and verified the effectiveness of our proposed method. Yalong Bai, Wei Yu 0004, Tianjun Xiao, Chang Xu 0008, Kuiyuan Yang, Wei-Ying Ma, Tiejun Zhao |
ACM Multimedia | 6 |
| 2012 | Semantic search and a new moore's law effect in knowledge engineeringabstractIn history, the Moore's law effect has been used to describe phenomena of exponential improvement in technology when it has a virtuous cycle that makes technology improvement proportional to technology itself. For example, chip performance had doubled every 18-24 months because better processors support the development of better layout tools that support the development of even better processors. I will describe a new Moore's law effect that is being created in knowledge engineering and is driven by the self-reinforcing nature of three trends and technical advancements: big data, machine learning, and Internet economics. I will explain how we can take advantage of this new effect to develop a new generation of semantic and knowledge-based search engines. Specifically, my presentation will cover the following three areas: Wei-Ying Ma |
KDD | 1 |
| 2012 | Flickr Distance: A Relationship Measure for Visual ConceptsabstractThis paper proposes the Flickr Distance (FD) to measure the visual correlation between concepts. For each concept, a collection of related images are obtained from the Flickr website. We assume that each concept consists of several states, e.g., different views, different semantics, etc., which are considered as latent topics. Then a latent topic visual language model (LTVLM) is built to capture these states. The Flickr distance between two concepts is defined as the Jensen-Shannon (J-S) divergence between their LTVLM. Differently from traditional conceptual distance measurements, which are based on Web textual documents, FD is based on the visual information. Comparing with the WordNet distance, FD can easily scale up with the increasing size of the conceptual corpus. Comparing with the Google Distance (NGD) and Tag Concurrence Distance (TCD), FD uses the visual information and can properly measure the conceptual relations. We apply FD to multimedia-related tasks and find methods based on FD significantly outperform those based on NGD and TCD. With the FD measurement, we also construct a large-scale visual conceptual network (VCNet) to store the knowledge of conceptual relationship. Experiments show that FD is more coherent to human cognition and it also outperforms text-based distances in real-world applications. Lei Wu 0017, Xian-Sheng Hua 0001, Nenghai Yu, Wei-Ying Ma, Shipeng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Statistical Entity Extraction From the WebabstractThere are various kinds of valuable semantic information about real-world entities embedded in webpages and databases. Extracting and integrating these entity information from the Web is of great significance. Comparing to traditional information extraction problems, web entity extraction needs to solve several new challenges to fully take advantage of the unique characteristic of the Web. In this paper, we introduce our recent work on statistical extraction of structured entities, named entities, entity facts and relations from Web. We also briefly introduce iKnoweb, an interactive knowledge mining framework for entity information integration. We will use two novel web applications, Microsoft Academic Search (aka Libra) and EntityCube, as working examples. Zaiqing Nie, Ji-Rong Wen, Wei-Ying Ma |
Proc. IEEE | 3 |
| 2012 | Duplicate-Search-Based Image Annotation Using Web-Scale DataabstractEasy photo-taking and photo-sharing today make image an increasingly important type of media in people's everyday life, which arouses a growing demand for a practical image understanding technique. Traditional computer vision or machine learning methods which learn models based on a set of training data are still in the stage of tackling hundreds of object categories. Such a scale is far from practical usage. In recent years, the technique of search-based image annotation on a large-scale data set has demonstrated great success. Rather than directly mapping visual features to texts which is inevitably hindered by the semantic gap, it understands the content of an image by propagating labels of its similar images in a large-scale data set. Since similarity search is performed among homogenous data, the difficulty is greatly reduced. This paper summarizes the extensive work on web image annotation using the large-scale metadata and social information available on the Web, and introduces the Arista system, which is a nonparametric image annotation platform built upon two billion web images. We propose a highly efficient and scalable duplicate-search technique so that the Arista system can be deployed on a few servers. A few interesting applications such as building large-scale celebrity face database and text-to-image translation are also presented in this paper. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
Proc. IEEE | 3 |
| 2011 | Recommending friends and locations based on individual location historyabstractThe increasing availability of location-acquisition technologies (GPS, GSM networks, etc.) enables people to log the location histories with spatio-temporal data. Such real-world location histories imply, to some extent, users' interests in places, and bring us opportunities to understand the correlation between users and locations. In this article, we move towards this direction and report on a personalized friend and location recommender for the geographical information systems (GIS) on the Web. First, in this recommender system, a particular individual's visits to a geospatial region in the real world are used as their implicit ratings on that region. Second, we measure the similarity between users in terms of their location histories and recommend to each user a group of potential friends in a GIS community. Third, we estimate an individual's interests in a set of unvisited regions by involving his/her location history and those of other users. Some unvisited locations that might match their tastes can be recommended to the individual. A framework, referred to as a hierarchical-graph-based similarity measurement (HGSM), is proposed to uniformly model each individual's location history, and effectively measure the similarity among users. In this framework, we take into account three factors: 1) the sequence property of people's outdoor movements, 2) the visited popularity of a geospatial region, and 3) the hierarchical property of geographic spaces. Further, we incorporated a content-based method into a user-based collaborative filtering algorithm, which uses HGSM as the user similarity measure, to estimate the rating of a user on an item. We evaluated this recommender system based on the GPS data collected by 75 subjects over a period of 1 year in the real world. As a result, HGSM outperforms related similarity measures, namely similarity-by-count, cosine similarity, and Pearson similarity measures. Moreover, beyond the item-based CF method and random recommendations, our system provides users with more attractive locations and better user experiences of recommendation. Yu Zheng 0004, Lizhu Zhang, Zhengxin Ma, Xing Xie 0001, Wei-Ying Ma |
ACM Trans. Web | 5 |
| 2010 | ARISTA - image search to annotation on billions of web photosabstractThough it has cost great research efforts for decades, object recognition is still a challenging problem. Traditional methods based on machine learning or computer vision are still in the stage of tackling hundreds of object categories. In recent years, non-parametric approaches have demonstrated great success, which understand the content of an image by propagating labels of its similar images in a large-scale dataset. However, due to the limited dataset size and imperfect image crawling strategy, previous work can only address a biased small subset of image concepts. Here we introduce the Arista project, which aims to build a practical image annotation engine targeting at popular concepts in the real world. In this project, we are particularly interested in understanding how many image concepts can be addressed by the data-driven annotation approach (coverage) and how good the performance is (precision). This paper reports the first stage of the work. Two billions web images were indexed, and based on simple yet effective near-duplicate detection, the system is capable of automatically generating accurate tags for popular web images having near-duplicates in the database. We found that about 8.1% web images have more than ten near duplicate and the number increases to 28.5% for top images in search results. Further, based on random samples in the latter case, we observed the precision of 57.9% at the point of the highest recall of 28% on ground truth tags. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
CVPR | 5 |
| 2010 | Empower People with Knowledge: The Next Frontier for Web Search
Wei-Ying Ma |
PAKDD (1) | 1 |
| 2010 | Mining adjacent markets from a large-scale ads video collection for image advertisingabstractThe research on image advertising is still in its infancy. Most previous approaches suggest ads by directly matching an ad to a query image, which lacks the power to identify ads from adjacent market. In this paper, we tackle the problem by mining knowledge on adjacent markets from ads videos with a novel Multi-Modal Dirichlet Process Mixture Sets model, which is a unified model of (video frames) clustering and (ads) ranking. Our approach is not only capable of discovering relevant ads (e.g. car ads for a query car image), but also suggesting ads from adjacent markets (e.g. tyre ads). Experimental results show that our proposed approach is fairly effective. Guwen Feng, Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
SIGIR | 4 |
| 2010 | Diversifying landmark image search results by learning interested views from community photosabstractIn this paper, we demonstrate a novel landmark photo search and browsing system: Agate, which ranks landmark image search results considering their relevance, diversity and quality. Agate learns from community photos the most interested aspects and related activities of a landmark, and generates adaptively a Table of Content (TOC) as a summary of the attractions to facilitate the user browsing. Image search results are thus re-ranked with the TOC so as to ensure a quick overview of the attractions of the landmarks. A novel non-parametric TOC generation and set-based ranking algorithm, MoM-DPM Sets, is proposed as the key technology of Agate. Experimental results based on human evaluation show the effectiveness of our model and users' preference for Agate. Yuheng Ren, Mo Yu, Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
WWW | 5 |
| 2010 | A large-scale study on map search logsabstractMap search engines, such as Google Maps, Yahoo! Maps, and Microsoft Live Maps, allow users to explicitly specify a target geographic location, either in keywords or on the map, and to search businesses, people, and other information of that location. In this article, we report a first study on a million-entry map search log. We identify three key attributes of a map search record—the keyword query, the target location and the user location, and examine the characteristics of these three dimensions separately as well as the associations between them. Comparing our results with those previously reported on logs of general search engines and mobile search engines, including those for geographic queries, we discover the following unique features of map search: (1) People use longer queries and modify queries more frequently in a session than in general search and mobile search; People view fewer result pages per query than in general search; (2) The popular query topics in map search are different from those in general search and mobile search; (3) The target locations in a session change within 50 kilometers for almost 80% of the sessions; (4) Queries, search target locations and user locations (both at the city level) all follow the power law distribution; (5) One third of queries are issued for target locations within 50 kilometers from the user locations; (6) The distribution of a query over target locations appears to follow the geographic location of the queried entity. Xiangye Xiao, Qiong Luo 0001, Zhisheng Li, Xing Xie 0001, Wei-Ying Ma |
ACM Trans. Web | 5 |
| 2010 | Understanding transportation modes based on GPS data for web applicationsabstractUser mobility has given rise to a variety of Web applications, in which the global positioning system (GPS) plays many important roles in bridging between these applications and end users. As a kind of human behavior, transportation modes, such as walking and driving, can provide pervasive computing systems with more contextual information and enrich a user's mobility with informative knowledge. In this article, we report on an approach based on supervised learning to automatically infer users' transportation modes, including driving, walking, taking a bus and riding a bike, from raw GPS logs. Our approach consists of three parts: a change point-based segmentation method, an inference model and a graph-based post-processing algorithm. First, we propose a change point-based segmentation method to partition each GPS trajectory into separate segments of different transportation modes. Second, from each segment, we identify a set of sophisticated features, which are not affected by differing traffic conditions (e.g., a person's direction when in a car is constrained more by the road than any change in traffic conditions). Later, these features are fed to a generative inference model to classify the segments of different modes. Third, we conduct graph-based postprocessing to further improve the inference performance. This postprocessing algorithm considers both the commonsense constraints of the real world and typical user behaviors based on locations in a probabilistic manner. The advantages of our method over the related works include three aspects. (1) Our approach can effectively segment trajectories containing multiple transportation modes. (2) Our work mined the location constraints from user-generated GPS logs, while being independent of additional sensor data and map information like road networks and bus stops. (3) The model learned from the dataset of some users can be applied to infer GPS data from others. Using the GPS logs collected by 65 people over a period of 10 months, we evaluated our approach via a set of experiments. As a result, based on the change-point-based segmentation method and Decision Tree-based inference model, we achieved prediction accuracy greater than 71 percent. Further, using the graph-based post-processing algorithm, the performance attained a 4-percent enhancement. Yu Zheng 0004, Quannan Li, Xing Xie 0001, Wei-Ying Ma |
ACM Trans. Web | 5 |
| 2009 | Vocabulary hierarchy optimization for effective and transferable retrievalabstractScalable image retrieval systems usually involve hierarchical quantization of local image descriptors, which produces a visual vocabulary for inverted indexing of images. Although hierarchical quantization has the merit of retrieval efficiency, the resulting visual vocabulary representation usually faces two crucial problems: (1) hierarchical quantization errors and biases in the generation of “visual words”; (2) the model cannot adapt to database variance. In this paper, we describe an unsupervised optimization strategy in generating the hierarchy structure of visual vocabulary, which produces a more effective and adaptive retrieval model for large-scale search. We adopt a novel Density-based Metric Learning (DML) algorithm, which corrects word quantization bias without supervision in hierarchy optimization, based on which we present a hierarchical rejection chain for efficient online search based on the vocabulary hierarchy. We also discovered that by hierarchy optimization, efficient and effective transfer of a retrieval model across different databases is feasible. We deployed a large-scale image retrieval system using a vocabulary tree model to validate our advances. Experiments on UKBench and street-side urban scene databases demonstrated the effectiveness of our hierarchy optimization approach in comparison with state-of-the-art methods. Rongrong Ji, Xing Xie 0001, Hongxun Yao, Wei-Ying Ma |
CVPR | 4 |
| 2009 | Mining correlation between locations using human location historyabstractThe advance of location-acquisition technologies enables people to record their location histories with spatio-temporal datasets, which imply the correlation between geographical regions. This correlation indicates the relationship between locations in the space of human behavior, and can enable many valuable services, such as sales promotion and location recommendation. In this paper, by taking into account a user's travel experience and the sequentiality locations have been visited, we propose an approach to mine the correlation between locations from a large number of users' location histories. We conducted a personalized location recommendation system using the location correlation, and evaluated this system with a large-scale real-world GPS dataset. As a result, our method outperforms the related work using the Pearson correlation. Yu Zheng 0004, Lizhu Zhang, Xing Xie 0001, Wei-Ying Ma |
GIS | 4 |
| 2009 | Efficient indexing for large scale visual searchabstractWith the popularity of “bag of visual terms” representations of images, many text indexing techniques have been applied in large-scale image retrieval systems. However, due to a fundamental difference between an image query (e.g. 1500 visual terms) and a text query (e.g. 3-5 terms), the usages of some text indexing techniques, e.g. inverted list, are misleading. In this work, we develop a novel indexing technique for this problem. The basic idea is to decompose a document-like representation of an image into two components, one for dimension reduction and the other for residual information preservation. The computing of similarity of two images can be transferred to measuring similarities of their components. The decomposition has two major merits: (1) these components have good properties which enable them to be efficiently indexed and retrieved; (2) The decomposition has better generalization ability than other dimension reduction algorithms. The decomposition can be achieved by either a graphical model or a matrix factorization approach. Theoretic analysis and extensive experiments over a 2.3 million image database show that this framework is scalable to index large scale image database to support fast and accurate visual search. Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma, Harry Shum |
ICCV | 4 |
| 2009 | Advertising based on users' photosabstractIn this paper, we tackle the problem of learning a user's interest from his photo collections and suggesting relevant ads. We address two key challenges in this work: 1) understanding a user's photos to detect his interest, and 2) bridging the lexical and semantic gap between the vocabulary of ads and that of general users' photos. We solve the first problem by employing a data-driven image annotation approach to annotate each photo and modeling a group of photos, and tackle the second problem by learning and matching the topics of users' photos and ads. The experiments based on real flicker data showed the effectiveness of the approach. Xin-Jing Wang, Mo Yu, Lei Zhang 0001, Wei-Ying Ma |
ICME | 4 |
| 2009 | Incorporating site-level knowledge for incremental crawling of web forums: a list-wise strategyabstractWe study in this paper the problem of incremental crawling of web forums, which is a very fundamental yet challenging step in many web applications. Traditional approaches mainly focus on scheduling the revisiting strategy of each individual page. However, simply assigning different weights for different individual pages is usually inefficient in crawling forum sites because of the different characteristics between forum sites and general websites. Instead of treating each individual page independently, we propose a list-wise strategy by taking into account the site-level knowledge. Such site-level knowledge is mined through reconstructing the linking structure, called sitemap, for a given forum site. With the sitemap, posts from the same thread but distributed on various pages can be concatenated according to their timestamps. After that, for each thread, we employ a regression model to predict the time when the next post arrives. Based on this model, we develop an efficient crawler which is 260% faster than some state-of-the-art methods in terms of fetching new generated content; and meanwhile our crawler also ensure a high coverage ratio. Experimental results show promising performance of Coverage, Bandwidth utilization, and Timeliness of our crawler on 18 various forums. Jiang-Ming Yang, Rui Cai 0002, Chunsong Wang, Hua Huang 0002, Lei Zhang 0001, Wei-Ying Ma |
KDD | 6 |
| 2009 | GeoLife2.0: A Location-Based Social Networking ServiceabstractGeoLife2.0 is a GPS-data-driven social networking service where people can share life experiences and connect to each other with their location histories. By mining peoplepsilas location history, GeoLife can measure the similarity between users and perform personalized friend recommendation for an individual. Later, we can predict the individualpsilas interest level in the locations visited by their friends while have not been found by them. The locations with relatively high interesting level can be recommended. Therefore, GeoLife2.0 can expand a userpsilas social network, provide them with a trustworthy resource matching their interests and help them sponsor geo-related activities like cycling with minimal effort. Yu Zheng 0004, Xing Xie 0001, Wei-Ying Ma |
Mobile Data Management | 4 |
| 2009 | Mining city landmarks from blogs by graph modelingabstractRecent years have witnessed great prosperity in community-contributed multimedia. Discovering, extracting, and summarizing knowledge from these data enables us to make better sense of the world. In this paper, we report our work on mining famous city landmarks from blogs for personalized tourist suggestions. Our main contribution is a graph modeling framework to discover city landmarks by mining blog photo correlations with community supervision. This modeling fuses context, content, and community information in a style that simulates both static (PageRank) and dynamic (HITS) ranking models to highlight representative data from the consensus of blog users. Rongrong Ji, Xing Xie 0001, Hongxun Yao, Wei-Ying Ma |
ACM Multimedia | 4 |
| 2009 | Argo: intelligent advertising made possible from users' photosabstractThough monetizing user-generated photos has a great potential in image business, this topic is seldom touched due to the difficulties of both image understanding and ads-to-images vocabulary matching. In this technical demonstration, we show case the Argo system, which attempts to monetize UGC (user-generated content) photos by mining a user's interest from a group of his photos and advertising the photos accordingly. Given a page of photos, it first auto-tags each photo by a large-scale search-based image annotation method, then maps both image annotations and the textual descriptions of ads onto an ODP-based topic hierarchy. The mapping produces semantic features which are statistical distributions on ODP topics. Ads are ranked by their similarities to such topic distributions of the photos and the top-ranked ones are output. Xin-Jing Wang, Mo Yu, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 4 |
| 2009 | Incorporating site-level knowledge to extract structured data from web forumsabstractWeb forums have become an important data resource for many web applications, but extracting structured data from unstructured web forum pages is still a challenging task due to both complex page layout designs and unrestricted user created posts. In this paper, we study the problem of structured data extraction from various web forum sites. Our target is to find a solution as general as possible to extract structured data, such as post title, post author, post time, and post content from any forum site. In contrast to most existing information extraction methods, which only leverage the knowledge inside an individual page, we incorporate both page-level and site-level knowledge and employ Markov logic networks (MLNs) to effectively integrate all useful evidence by learning their importance automatically. Site-level knowledge includes (1) the linkages among different object pages, such as list pages and post pages, and (2) the interrelationships of pages belonging to the same object. The experimental results on 20 forums show a very encouraging information extraction performance, and demonstrate the ability of the proposed approach on various forums. We also show that the performance is limited if only page-level knowledge is used, while when incorporating the site-level knowledge both precision and recall can be significantly improved. Jiang-Ming Yang, Rui Cai 0002, Yida Wang 0008, Jun Zhu 0001, Lei Zhang 0001, Wei-Ying Ma |
WWW | 6 |
| 2009 | Mining interesting locations and travel sequences from GPS trajectoriesabstractThe increasing availability of GPS-enabled devices is changing the way people interact with the Web, and brings us a large amount of GPS trajectories representing people's location histories. In this paper, based on multiple users' GPS trajectories, we aim to mine interesting locations and classical travel sequences in a given geospatial region. Here, interesting locations mean the culturally important places, such as Tiananmen Square in Beijing, and frequented public areas, like shopping malls and restaurants, etc. Such information can help users understand surrounding locations, and would enable travel recommendation. In this work, we first model multiple individuals' location histories with a tree-based hierarchical graph (TBHG). Second, based on the TBHG, we propose a HITS (Hypertext Induced Topic Search)-based inference model, which regards an individual's access on a location as a directed link from the user to that location. This model infers the interest of a location by taking into account the following three factors. 1) The interest of a location depends on not only the number of users visiting this location but also these users' travel experiences. 2) Users' travel experiences and location interests have a mutual reinforcement relationship. 3) The interest of a location and the travel experience of a user are relative values and are region-related. Third, we mine the classical travel sequences among locations considering the interests of these locations and users' travel experiences. We evaluated our system using a large GPS dataset collected by 107 users over a period of one year in the real world. As a result, our HITS-based inference model outperformed baseline approaches like rank-by-count and rank-by-frequency. Meanwhile, when considering the users' travel experiences and location interests, we achieved a better performance beyond baselines, such as rank-by-count and rank-by-interest, etc. Yu Zheng 0004, Lizhu Zhang, Xing Xie 0001, Wei-Ying Ma |
WWW | 4 |
| 2009 | Browsing on small displays by transforming Web pages into hierarchically structured subpagesabstractWe propose a new Web page transformation method to facilitate Web browsing on handheld devices such as Personal Digital Assistants (PDAs). In our approach, an original Web page that does not fit on the screen is transformed into a set of subpages, each of which fits on the screen. This transformation is done through slicing the original page into page blocks iteratively, with several factors considered. These factors include the size of the screen, the size of each page block, the number of blocks in each transformed page, the depth of the tree hierarchy that the transformed pages form, as well as the semantic coherence between blocks. We call the tree hierarchy of the transformed pages an SP-tree. In an SP-tree, an internal node consists of a textually enhanced thumbnail image with hyperlinks, and a leaf node is a block extracted from a subpage of the original Web page. We adaptively adjust the fanout and the height of the SP-tree so that each thumbnail image is clear enough for users to read, while at the same time, the number of clicks needed to reach a leaf page is few. Through this transformation algorithm, we preserve the contextual information in the original Web page and reduce scrolling. We have implemented this transformation module on a proxy server and have conducted usability studies on its performance. Our system achieved a shorter task completion time compared with that of transformations from the Opera browser in nine of ten tasks. The average improvement on familiar pages was 44%. The average improvement on unfamiliar pages was 37%. Subjective responses were positive. Xiangye Xiao, Qiong Luo 0001, Dan Hong, Hongbo Fu 0001, Xing Xie 0001, Wei-Ying Ma |
ACM Trans. Web | 6 |
| 2008 | Building Web-Scale Data Mining Infrastructure for Search
Wei-Ying Ma |
APWeb | 1 |
| 2008 | Search-based query suggestionabstractIn this paper, we proposed a unified strategy to combine query log and search results for query suggestion. In this way, we leverage both the users' search intentions for popular queries and the power of search engines for unpopular queries. The suggested queries are also ranked according to their relevance and qualities; and each suggestion is described with a rich snippet including a photo and related description. Jiang-Ming Yang, Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
CIKM | 6 |
| 2008 | What are the high-level concepts with small semantic gaps?abstractConcept-based multimedia search has become more and more popular in Multimedia Information Retrieval (MIR). However, which semantic concepts should be used for data collection and model construction is still an open question. Currently, there is very little research found on automatically choosing multimedia concepts with small semantic gaps. In this paper, we propose a novel framework to develop a lexicon of high-level concepts with small semantic gaps (LCSS) from a large-scale web image dataset. By defining a confidence map and content-context similarity matrix, images with small semantic gaps are selected and clustered. The final concept lexicon is mined from the surrounding descriptions (titles, categories and comments) of these images. This lexicon offers a set of high-level concepts with small semantic gaps, which is very helpful for people to focus for data collection, annotation and modeling. It also shows a promising application potential for image annotation refinement and rejection. The experimental results demonstrate the validity of the developed concepts lexicon. Yijuan Lu, Lei Zhang 0001, Qi Tian 0001, Wei-Ying Ma |
CVPR | 4 |
| 2008 | Mining user similarity based on location historyabstractThe pervasiveness of location-acquisition technologies (GPS, GSM networks, etc.) enable people to conveniently log the location histories they visited with spatio-temporal data. The increasing availability of large amounts of spatio-temporal data pertaining to an individual's trajectories has given rise to a variety of geographic information systems, and also brings us opportunities and challenges to automatically discover valuable knowledge from these trajectories. In this paper, we move towards this direction and aim to geographically mine the similarity between users based on their location histories. Such user similarity is significant to individuals, communities and businesses by helping them effectively retrieve the information with high relevance. A framework, referred to as hierarchical-graph-based similarity measurement (HGSM), is proposed for geographic information systems to consistently model each individual's location history and effectively measure the similarity among users. In this framework, we take into account both the sequence property of people's movement behaviors and the hierarchy property of geographic spaces. We evaluate this framework using the GPS data collected by 65 volunteers over a period of 6 months in the real world. As a result, HGSM outperforms related similarity measures, such as the cosine similarity and Pearson similarity measures. Quannan Li, Yu Zheng 0004, Xing Xie 0001, Wenyu Liu 0001, Wei-Ying Ma |
GIS | 6 |
| 2008 | Density based co-location pattern discoveryabstractCo-location pattern discovery is to find classes of spatial objects that are frequently located together. For example, if two categories of businesses often locate together, they might be identified as a co-location pattern; if several biologic species frequently live in nearby places, they might be a co-location pattern. Most existing co-location pattern discovery methods are generate-and-test methods, that is, generate candidates, and test each candidate to determine whether it is a co-location pattern. In the test step, we identify instances of a candidate to obtain its prevalence. In general, instance identification is very costly. In order to reduce the computational cost of identifying instances, we propose a density based approach. We divide objects into partitions and identifying instances in dense partitions first. A dynamic upper bound of the prevalence for a candidate is maintained. If the current upper bound becomes less than a threshold, we stop identifying its instances in the remaining partitions. We prove that our approach is complete and correct in finding co-location patterns. Experimental results on real data sets show that our method outperforms a traditional approach. Xiangye Xiao, Xing Xie 0001, Qiong Luo 0001, Wei-Ying Ma |
GIS | 4 |
| 2008 | Understanding mobility based on GPS dataabstractBoth recognizing human behavior and understanding a user's mobility from sensor data are critical issues in ubiquitous computing systems. As a kind of user behavior, the transportation modes, such as walking, driving, etc., that a user takes, can enrich the user's mobility with informative knowledge and provide pervasive computing systems with more context information. In this paper, we propose an approach based on supervised learning to infer people's motion modes from their GPS logs. The contribution of this work lies in the following two aspects. On one hand, we identify a set of sophisticated features, which are more robust to traffic condition than those other researchers ever used. On the other hand, we propose a graph-based post-processing algorithm to further improve the inference performance. This algorithm considers both the commonsense constraint of real world and typical user behavior based on location in a probabilistic manner. Using the GPS logs collected by 65 people over a period of 10 months, we evaluated our approach via a set of experiments. As a result, based on the change point-based segmentation method and Decision Tree-based inference model, the new features brought an eight percent improvement in inference accuracy over previous result, and the graph-based post-processing achieve a further four percent enhancement. Yu Zheng 0004, Quannan Li, Xing Xie 0001, Wei-Ying Ma |
UbiComp | 5 |
| 2008 | Visual pattern weighting for near-duplicate image retrievalabstractRecently, there has been growing interest in mining co-location visual patterns from a collection of images. To find a proper usage of visual patterns in near-duplicate image retrieval systems, we study a TF-IDF weighting function for visual patterns. We show usage of TF and IDF respectively in this weighting function. Experiments demonstrate that 1) visual patterns and words should be weighted separately; 2) visual patterns consist of more frequent visual words would be more important; 3) visual patterns should be given lower weight values than visual words in near-duplicate image retrieval tasks. Manni Duan, Xing Xie 0001, Xiuqing Wu, Wei-Ying Ma |
ICME | 4 |
| 2008 | Vocabulary tree incremental indexing for scalable location recognitionabstractThis work aims at developing a scalable vision-based location recognition system where the backend database can be updated incrementally. Our proposed framework enables incremental indexing of vocabulary tree model, which efficiently includes new data into model refinement without re-generating entire model from overall dataset. An adaption trigger criterion is presented to lessen system computational cost, which is achieved by density-based relative entropy estimation between original dataset and newly coming data. Experiments on Seattle urban scene datasets with over 20K street-side images show the effectiveness of our work. Rongrong Ji, Xing Xie 0001, Hongxun Yao, Yongjian Wu 0001, Wei-Ying Ma |
ICME | 5 |
| 2008 | Spatial pyramid mining for logo detection in natural scenesabstractThis work introduces a novel data mining scheme, spatial pyramid mining, to discover association rules at multiple resolutions in order to identify frequent spatial configurations of local features that correspond to classes of logos appearing in real world scenes. By indexing representative examples by the mined rules we can efficiently detect a variety of different lettering or design marks associated with a brand. Features in an image are marked by matching rules to representative examples selected via a weighted cosine similarity measure. Logos are localized in an image via density-based clustering of matched features. Precision vs. recall curves are presented for experiments on a dataset of web images of nearly 1,000 images containing seven popular logo types. Jim Kleban, Xing Xie 0001, Wei-Ying Ma |
ICME | 3 |
| 2008 | Automatic video annotation through search and miningabstractConventional approaches to video annotation predominantly focus on supervised identification of a limited set of concepts, while unsupervised annotation with infinite vocabulary remains unexplored. This work aims to exploit the overlap in content of news video to automatically annotate by mining similar videos that reinforce, filter, and improve the original annotations. The algorithm employs a two-step process of search followed by mining. Given a query video consisting of visual content and speech-recognized transcripts, similar videos are first ranked in a multimodal search. Then, the transcripts associated with these similar videos are mined to extract keywords for the query. We conducted extensive experiments over the TRECVID 2005 corpus and showed the superiority of the proposed approach to using only the mining process on the original video for annotation. This work represents the first attempt at unsupervised automatic video annotation leveraging overlapping video content. Emily Moxley, Tao Mei 0001, Xian-Sheng Hua 0001, Wei-Ying Ma, B. S. Manjunath |
ICME | 4 |
| 2008 | A Flexible Spatio-Temporal Indexing Scheme for Large-Scale GPS Track RetrievalabstractThe increasing popularity of GPS device has boosted many Web applications where people can upload, browse and exchange their GPS tracks. In these applications, spatial or temporal search function could provide an effective way for users to retrieve specific GPS tracks they are interested in. However, existing spatial-temporal index for trajectory data has not exploited the characteristic of user behavior in these online GPS track sharing applications. In most cases, when sharing a GPS track, people are more likely to upload GPS data of the near past than the distant past. Thus, the interval between the end time of a GPS track and the time it is uploaded, if viewed as a random variable, has a skewed distribution. In this paper, we first propose a probabilistic model to simulate user behavior of uploading GPS tracks onto an online sharing application. Then we propose a flexible spatio-temporal index scheme, referred to as Compressed Start-End Tree (CSE-tree), for large-scale GPS track retrieval. The CSE-tree combines the advantages of B+ Tree and dynamic array, and maintains different index structure for data with different update frequency. Experiments using synthetic data show that CSE-tree outperforms other schemes in requiring less index size and less update cost while keeping satisfactory retrieval performance. Longhao Wang, Yu Zheng 0004, Xing Xie 0001, Wei-Ying Ma |
MDM | 4 |
| 2008 | GeoLife: Managing and Understanding Your Past Life over MapsabstractThe increasing popularity of GPS device has boosted many applications where more and more GPS logs have been accumulating continuously. Managing and understanding the collected GPS data are two important issues for these applications. On one hand, by indexing the increasing GPS data, we can provide effective retrieval method for users to find the corresponding GPS data interests them. On the other hand, by understanding user's GPS data, we are more likely to enable novel services which would stimulate people's passion on contributing GPS data in turn. However, so far, GPS data are still used directly without much understanding. In our project, referred to as GeoLife, we focus on visualization, organization, fast retrieval, and effective understanding of GPS track logs for both personal and public use. It not only provides a powerful platform for people to effectively manage their GPS data but also help them well understand a person's past experience from GPS data. Yu Zheng 0004, Longhao Wang, Ruochi Zhang, Xing Xie 0001, Wei-Ying Ma |
MDM | 5 |
| 2008 | Delivering online advertisements inside imagesabstractWe present in this paper a new channel to deliver online advertisements along with Web images and show a new business model to monetize billions of Web images. The idea is intuitively inspired by image displaying processes on the Web, which typically require people to wait a few seconds before they see full resolution images. This is due to large file sizes and limited network bandwidth. To utilize idle time and the display area, we propose an innovative method for non-intrusively embedding ads into images in a visually pleasant manner. To maintain a smooth user experience, we utilize the thumbnail of the full-resolution image because it is small and visually similar to the full-resolution image. At the client side, a rendering engine first enlarges and blurs the thumbnail, and then blends the pre-chosen ads information into the enlarged image. Based on this idea, we propose three typical scenarios that can adopt the proposed image-advertising mode. More importantly, we can encourage providers of images or other users to participate in our online image ads service by tagging or annotating images. We envision revenue sharing with the providers participating in our service, and we expect that a large number of users will actively submit, tag and annotate images using the system. We have implemented a prototype image ads system, and conducted a series of experiments and user studies to evaluate such a new advertisement channel. The experimental results and user studies show that the proposed online image ad delivery is a non-intrusive ads mode, and the proposed solution is practical. This work also opens multiple new research directions ranging from multimedia to web data mining Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2008 | Flickr distanceabstractThis paper presents Flickr distance, which is a novel measurement of the relationship between semantic concepts (objects, scenes) in visual domain. For each concept, a collection of images are obtained from Flickr, based on which the improved latent topic based visual language model is built to capture the visual characteristic of this concept. Then Flickr distance between different concepts is measured by the square root of Jensen-Shannon (JS) divergence between the corresponding visual language models. Comparing with WordNet, Flickr distance is able to handle far more concepts existing on the Web, and it can scale up with the increase of concept vocabularies. Comparing with Google distance, which is generated in textual domain, Flickr distance is more precise for visual domain concepts, as it captures the visual relationship between the concepts instead of their co-occurrence in text search results. Besides, unlike Google distance, Flickr distance satisfies triangular inequality, which makes it a more reasonable distance metric. Both subjective user study and objective evaluation show that Flickr distance is more coherent to human perception than Google distance. We also design several application scenarios, such as concept clustering and image annotation, to demonstrate the effectiveness of this proposed distance in image related applications. Lei Wu 0017, Xian-Sheng Hua 0001, Nenghai Yu, Wei-Ying Ma, Shipeng Li 0001 |
ACM Multimedia | 4 |
| 2008 | Exploring traversal strategy for web forum crawlingabstractIn this paper, we study the problem of Web forum crawling. Web forum has now become an important data source of many Web applications; while forum crawling is still a challenging task due to complex in-site link structures and login controls of most forum sites. Without carefully selecting the traversal path, a generic crawler usually downloads many duplicate and invalid pages from forums, and thus wastes both the precious bandwidth and the limited storage space. To crawl forum data more effectively and efficiently, in this paper, we propose an automatic approach to exploring an appropriate traversal strategy to direct the crawling of a given target forum. In detail, the traversal strategy consists of the identification of the skeleton links and the detection of the page-flipping links. The skeleton links instruct the crawler to only crawl valuable pages and meanwhile avoid duplicate and uninformative ones; and the page-flipping links tell the crawler how to completely download a long discussion thread which is usually shown in multiple pages in Web forums. The extensive experimental results on several forums show encouraging performance of our approach. Following the discovered traversal strategy, our forum crawler can archive more informative pages in comparison with previous related work and a commercial generic crawler. Yida Wang 0008, Jiang-Ming Yang, Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
SIGIR | 6 |
| 2008 | Directly optimizing evaluation measures in learning to rankabstractOne of the central issues in learning to rank for information retrieval is to develop algorithms that construct ranking models by directly optimizing evaluation measures used in information retrieval such as Mean Average Precision (MAP) and Normalized Discounted Cumulative Gain (NDCG). Several such algorithms including SVMmap and AdaRank have been proposed and their effectiveness has been verified. However, the relationships between the algorithms are not clear, and furthermore no comparisons have been conducted between them. In this paper, we conduct a study on the approach of directly optimizing evaluation measures in learning to rank for Information Retrieval (IR). We focus on the methods that minimize loss functions upper bounding the basic loss function defined on the IR measures. We first provide a general framework for the study and analyze the existing algorithms of SVMmap and AdaRank within the framework. The framework is based on upper bound analysis and two types of upper bounds are discussed. Moreover, we show that we can derive new algorithms on the basis of this analysis and create one example algorithm called PermuRank. We have also conducted comparisons between SVMmap, AdaRank, PermuRank, and conventional methods of Ranking SVM and RankBoost, using benchmark datasets. Experimental results show that the methods based on direct optimization of evaluation measures can always outperform conventional methods of Ranking SVM and RankBoost. However, no significant difference exists among the performances of the direct optimization methods themselves. Jun Xu 0001, Tie-Yan Liu, Hang Li 0001, Wei-Ying Ma |
SIGIR | 5 |
| 2008 | Rich media and web 2.0abstractRich media data, such as video, imagery, music, and gaming, do no longer play just a supporting role on the World Wide Web to text data. Thanks to Web 2.0, rich media is the primary content on sites such as Flickr, PicasaWeb, YouTube, and QQ. Because of massive user generated content, the volume of rich media being transmitted on the Internet has surpassed that of text. It is vital to properly manage these data to ensure efficient bandwidth utilization, to support effective indexing and search, and to safeguard copyrights (just to name a few). This panel invites both researchers and practitioners to discuss the challenges of Web-scale media-data management. In particular, the panelists will address issues such as leveraging Rich Media and Web 2.0, indexing, search, and scalability. Edward Y. Chang, Ken Ong, Susanne Boll, Wei-Ying Ma |
WWW | 4 |
| 2008 | A systematic study on parameter correlations in large-scale duplicate document detection
Shaozhi Ye, Ji-Rong Wen, Wei-Ying Ma |
Knowl. Inf. Syst. | 3 |
| 2008 | Annotating Images by Mining Image Search ResultsabstractAlthough it has been studied for years by the computer vision and machine learning communities, image annotation is still far from practical. In this paper, we propose a novel attempt at model-free image annotation, which is a data-driven approach that annotates images by mining their search results. Some 2.4 million images with their surrounding text are collected from a few photo forums to support this approach. The entire process is formulated in a divide-and-conquer framework where a query keyword is provided along with the uncaptioned image to improve both the effectiveness and efficiency. This is helpful when the collected data set is not dense everywhere. In this sense, our approach contains three steps: 1) the search process to discover visually and semantically similar search results, 2) the mining process to identify salient terms from textual descriptions of the search results, and 3) the annotation rejection process to filter out noisy terms yielded by Step 2. To ensure real-time annotation, two key techniques are leveraged-one is to map the high-dimensional image visual features into hash codes, the other is to implement it as a distributed system, of which the search and mining processes are provided as Web services. As a typical result, the entire process finishes in less than 1 second. Since no training data set is required, our approach enables annotating with unlimited vocabulary and is highly scalable and robust to outliers. Experimental results on both real Web images and a benchmark image data set show the effectiveness and efficiency of the proposed algorithm. It is also worth noting that, although the entire approach is illustrated within the divide-and conquer framework, a query keyword is not crucial to our current implementation. We provide experimental results to prove this. Xin-Jing Wang, Lei Zhang 0001, Xirong Li 0001, Wei-Ying Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2008 | Mobile Search With Multimodal QueriesabstractThe popularity of mobile devices, such as PDAs and SmartPhones, has grown rapidly over the last couple of years. Though most users still perform searches using desktop computers, it is expected that more and more people will also search the Web while they are on the move. In addition to text-based keyword queries, mobile devices can support richer and hybrid queries such as images, audio, video, and their combinations. In this paper, we will discuss mobile search systems that support image queries and audio queries, covering typical designs for mobile visual and audio search, as well as the opportunities and challenges. Specifically, we will present an in-depth study of two real systems we have developed: product image categorization and mobile ringtone search, which use image queries and audio queries, respectively. Experimental results on real-life data demonstrate their effectiveness and efficiency. Xing Xie 0001, Lie Lu, Menglei Jia, Hua Li 0001, Frank Seide, Wei-Ying Ma |
Proc. IEEE | 6 |
| 2008 | An active feedback framework for image retrieval
Tao Qin 0001, Xudong Zhang 0001, Tie-Yan Liu, De-Sheng Wang, Wei-Ying Ma, HongJiang Zhang |
Pattern Recognit. Lett. | 5 |
| 2007 | Object-level Vertical Search
Zaiqing Nie, Ji-Rong Wen, Wei-Ying Ma |
CIDR | 3 |
| 2007 | Computing Geographical Serving Area Based on Search Logs and Website Categorization
Qi Zhang 0066, Xing Xie 0001, Lee Wang, Lihua Yue, Wei-Ying Ma |
DEXA | 5 |
| 2007 | Fast Large-Scale Spectral Clustering by Sequential Shrinkage Optimization
Tie-Yan Liu, Huai-Yuan Yang, Tao Qin 0001, Wei-Ying Ma |
ECIR | 5 |
| 2007 | Improve Ranking by Using Image Information
Shuming Shi 0001, Zhiwei Li 0006, Ji-Rong Wen, Wei-Ying Ma |
ECIR | 5 |
| 2007 | Automated Music Video Generation using WEB Image ResourceabstractIn this paper, we proposed a novel prototype of automated music video generation using web image resource. In this prototype, the salient words/phrases of a song's lyrics are first automatically extracted and then used as queries to retrieve related high-quality images from web search engines. To guarantee the coherence among the chosen images' visual representation and the music song, the returned images are further re-ranked and filtered based on their content characteristics such as color, face, landscape, as well as the song's mood type. Finally, those selected images are concatenated to generate a music video using the Photo2Video technique, based on the rhythm information of the music. Preliminary evaluations of the proposed prototype have shown promising results. Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
ICASSP (2) | 5 |
| 2007 | Recent Advances and Challenges of Semantic Image/Video SearchabstractWe present an overview of recent advances and major challenges in image and video search, with a specific focus on large-scale semantic concept detection and indexing. Such semantic indexing paradigm has been driven by the increasing availability of the large resources of corpora, novel labeling approaches, innovative image features, and machine learning techniques for visual content recognition. We will discus key approaches, recent results, and novel applications in text-to-concept semantic search and multi-modal retrieval models. Open issues and major opportunities are also presented. Shih-Fu Chang, Wei-Ying Ma, Arnold W. M. Smeulders |
ICASSP (4) | 2 |
| 2007 | Distributed Architecture for Large Scale Image-Based SearchabstractIn recent years, some computer vision algorithms such as SIFT (Scale Invariant Feature Transform) have been employed in image similarity match to perform image-based search applications. However, with the increasing scale of image databases, centralized image retrieval system no longer provide adequate prompt search. In this paper, we design a scalable distributed architecture, which is analog to web search engine, for efficient large-scale image retrieval. In our distributed architecture, images are partitioned to multiple servers and an index is built. Administrated by a controlling server, each distributed server matches query image in its own image sub-collection in parallel and returns the intermediate search results, a list of images similar to query image, to the controlling server for further re-ranking and merging. An evaluation of the results shows that our distributed architecture removes the limitation of a centralized image retrieval system. By performing reasonable indexing, merging and ranking strategies, the precision level of search is near to that performed on stand-alone retrieval systems indexing all images. Yu Zheng 0004, Xing Xie 0001, Wei-Ying Ma |
ICME | 3 |
| 2007 | MusicSense: contextual music recommendation using emotional allocation modelingabstractIn this paper, we present a novel contextual music recommendation approach, MusicSense, to automatically suggest music when users read Web documents such as Weblogs. MusicSense matches music to a document's content, in terms of the emotions expressed by both the document and the music songs. To achieve this, we propose a generative model - Emotional Allocation Modeling - in which a collection of word terms is considered as generated with a mixture of emotions. This model also integrates knowledge discovering from a Web-scale corpus and guidance from psychological studies of emotion. Music songs are also described using textual information extracted from their meta-data and relevant Web pages. Thus, both music songs and Web documents can be characterized as distributions over the emotion mixtures through the emotional allocation modeling. For a given document, the songs with the most matched emotion distributions are finally selected as the recommendations. Preliminary experiments on Weblogs show promising results on both emotion allocation and music recommendation. Rui Cai 0002, Chong Wang 0002, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 5 |
| 2007 | Scalable music recommendation by searchabstractThe growth of music resources on personal devices and Internet radio has increased the need for music recommendations. In this paper, aiming at providing an efficient and general solution, we present a search-based solution for scalable music recommendations. In this solution a music piece is first transformed to a music signature sequence in which each signature characterizes the timbre of a local music clip. Based on such signatures, a scale-sensitive method is then proposed to index the music pieces for similarity search, using the locality sensitive hashing (LSH). The scale-sensitive method can numerically find the appropriate parameters for indexing various scales of music collections, and thus can guarantee a proper number of nearest neighbors are found in search. In the recommendation stage, representative signatures from snippets of a seed piece are extracted as query terms, to retrieve pieces with similar melodies for suggestions. We also design a relevance-ranking function to sort the search results, based on the criteria that include matching ratio, temporal order, term weight, and matching confidence. Finally, with the search results, we propose a strategy to generate a dynamic playlist which can automatically expand with time. Evaluations of several music collections at various scales showed that our approach achieves encouraging results in terms of recommendation satisfaction and system scalability. Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 4 |
| 2007 | Dual cross-media relevance model for image annotationabstractImage annotation has been an active research topic in recent years due to its potential impact on both image understanding and web image retrieval. Existing relevance-model-based methods perform image annotation by maximizing the joint probability of images and words, which is calculated by the expectation over training images. However, the semantic gap and the dependence on training data restrict their performance and scalability. In this paper, a dual cross-media relevance model (DCMRM) is proposed for automatic image annotation, which estimates the joint probability by the expectation over words in a pre-defined lexicon. DCMRM involves two kinds of critical relations in image annotation. One is the word-to-image relation and the other is the word-to-word relation. Both relations can be estimated by using search techniques on the web data as well as available training data. Experiments conducted on the Corel dataset and a web image dataset demonstrate the effectiveness of the proposed model. Jing Liu 0001, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma, Hanqing Lu, Songde Ma |
ACM Multimedia | 5 |
| 2007 | Bipartite graph reinforcement model for web image annotationabstractAutomatic image annotation is an effective way for managing and retrieving abundant images on the internet. In this paper, a bipartite graph reinforcement model (BGRM) is proposed for web image annotation. Given a web image, a set of candidate annotations is extracted from its surrounding text and other textual information in the hosting web page. As this set is often incomplete, it is extended to include more potentially relevant annotations by searching and mining a large-scale image database. All candidates are modeled as a bipartite graph. Then a reinforcement algorithm is performed on the bipartite graph to re-rank the candidates. Only those with the highest ranking scores are reserved as the final annotations. Experimental results on real web images demonstrate the effectiveness of the proposed model. Xiaoguang Rui, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma, Nenghai Yu |
ACM Multimedia | 4 |
| 2007 | Dual-Space Pyramid Matching for Medical Image Classification
Yang Hu 0006, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma |
MMM (1) | 4 |
| 2007 | FRank: a ranking method with fidelity lossabstractRanking problem is becoming important in many fields, especially in information retrieval (IR). Many machine learning techniques have been proposed for ranking problem, such as RankSVM, RankBoost, and RankNet. Among them, RankNet, which is based on a probabilistic ranking framework, is leading to promising results and has been applied to a commercial Web search engine. In this paper we conduct further study on the probabilistic ranking framework and provide a novel loss function named fidelity loss for measuring loss of ranking. The fidelity loss notonly inherits effective properties of the probabilistic ranking framework in RankNet, but possesses new properties that are helpful for ranking. This includes the fidelity loss obtaining zero for each document pair, and having a finite upper bound that is necessary for conducting query-level normalization. We also propose an algorithm named FRank based on a generalized additive model for the sake of minimizing the fedelity loss and learning an effective ranking function. We evaluated the proposed algorithm for two datasets: TREC dataset and real Web search dataset. The experimental results show that the proposed FRank algorithm outperforms other learning-based ranking methods on both conventional IR problem and Web search. Ming-Feng Tsai, Tie-Yan Liu, Tao Qin 0001, Hsin-Hsi Chen, Wei-Ying Ma |
SIGIR | 5 |
| 2007 | Webstudio: building infrastructure for web data managementabstractTo explore various ideas and algorithms for improving relevance of a search engine, we found it necessary to build an infrastructure to provide large-scale data management and data processing capabilities. WebStudio is an infrastructure we have constructed to provide an integrated development environment (IDE) for researchers and developers to use in quickly building prototypes and conducting experiments at Web-scale. It is also a Web data management system to allow users to easily store, access, and manipulate Web data. Ji-Rong Wen, Wei-Ying Ma |
SIGMOD Conference | 2 |
| 2007 | Web object retrievalabstractThe primary function of current Web search engines is essentially relevance ranking at the document level. However, myriad structured information about real-world objects embedded in static Web pages and online Web databases. In this paper, we propose a paradigm shift to enable searching at the object level. In traditional information retrieval models, documents are taken as the retrieval units and the content of a document is considered reliable. However, this reliability assumption is no longer valid in the object retrieval context when multiple copies of information about the same object typically exist. These copies may be inconsistent because of diversity of Web site qualities and the limited performance of current information extraction techniques. In this paper, we propose several language models for Web object retrieval. We test these models on our academic search engine called Libra and compare their performances. 1. Zaiqing Nie, Yunxiao Ma, Shuming Shi 0001, Ji-Rong Wen, Wei-Ying Ma |
WWW | 5 |
| 2007 | Topic distillation via sub-site retrieval
Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, De-Sheng Wang, Wei-Ying Ma |
Inf. Process. Manag. | 6 |
| 2007 | A survey of content-based image retrieval with high-level semantics
Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma |
Pattern Recognit. | 4 |
| 2007 | Clustering and searching WWW images using link and page layout analysisabstractDue to the rapid growth of the number of digital images on the Web, there is an increasing demand for an effective and efficient method for organizing and retrieving the available images. This article describes iFind, a system for clustering and searching WWW images. By using a vision-based page segmentation algorithm, a Web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. The textual information is used for image indexing. By extracting the page-to-block, block-to-image, block-to-page relationships through link structure and page layout analysis, we construct an image graph. Our method is less sensitive to noisy links than previous methods like PageRank, HITS, and PicASHOW, and hence the image graph can better reflect the semantic relationship between images. Using the notion of Markov Chain, we can compute the limiting probability distributions of the images, ImageRanks, which characterize the importance of the images. The ImageRanks are combined with the relevance scores to produce the final ranking for image search. With the graph models, we can also use techniques from spectral graph theory for image clustering and embedding, or 2-D visualization. Some experimental results on 11.6 million images downloaded from the Web are provided in the article. Xiaofei He 0001, Deng Cai 0001, Ji-Rong Wen, Wei-Ying Ma, HongJiang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2006 | Adaptive User Profile Model and Collaborative Filtering for Personalized News
Zhiwei Li 0006, Jinyi Yao, Zengqi Sun, Mingjing Li, Wei-Ying Ma |
APWeb | 6 |
| 2006 | Ranking web objects from multiple communitiesabstractVertical search is a promising direction as it leverages domain-specific knowledge and can provide more precise information for users. In this paper, we study the Web object-ranking problem, one of the key issues in building a vertical search engine. More specifically, we focus on this problem in cases when objects lack relationships between different Web communities, and take high-quality photo search as the test bed for this investigation. We proposed two score fusion methods that can automatically integrate as many Web communities (Web forums) with rating information as possible. The proposed fusion methods leverage the hidden links discovered by a duplicate photo detection algorithm, and aims at minimizing score differences of duplicate photos in different forums. Both intermediate results and user studies show the proposed fusion methods are practical and efficient solutions to Web object ranking in cases we have described. Though the experiments were conducted on high-quality photo ranking, the proposed algorithms are also applicable to other ranking problems, such as movie ranking and music ranking. Lei Zhang 0001, Kefeng Deng, Wei-Ying Ma |
CIKM | 5 |
| 2006 | A comparative study on classifying the functions of web page blocksabstractIn this paper, we study the problem of learning block classification models to estimate block functions. We distinguish general models, which are learned across multiple sites, and site-specific models, which are learned within individual sites. We further consider several factors that affect the learning process and model effectiveness. These factors include the layout features, the content features, the classifiers, and the term selection methods. We have empirically evaluated the performance of the models when the factors are varied. Our main results are that layout features do better than content features for learning both general and site-specific models. Xiangye Xiao, Qiong Luo 0001, Xing Xie 0001, Wei-Ying Ma |
CIKM | 4 |
| 2006 | Learning Distance Metrics with Contextual Constraints for Image RetrievalabstractRelevant Component Analysis (RCA) has been proposed for learning distance metrics with contextual constraints for image retrieval. However, RCA has two important disadvantages. One is the lack of exploiting negative constraints which can also be informative, and the other is its incapability of capturing complex nonlinear relationships between data instances with the contextual information. In this paper, we propose two algorithms to overcome these two disadvantages, i.e., Discriminative Component Analysis (DCA) and Kernel DCA. Compared with other complicated methods for distance metric learning, our algorithms are rather simple to understand and very easy to solve. We evaluate the performance of our algorithms on image retrieval in which experimental results show that our algorithms are effective and promising in learning good quality distance metrics for image retrieval. Steven C. H. Hoi, Wei Liu 0005, Michael R. Lyu, Wei-Ying Ma |
CVPR (2) | 4 |
| 2006 | AnnoSearch: Image Auto-Annotation by SearchabstractAlthough it has been studied for several years by computer vision and machine learning communities, image annotation is still far from practical. In this paper, we present AnnoSearch, a novel way to annotate images using search and data mining technologies. Leveraging the Web-scale images, we solve this problem in two-steps: 1) searching for semantically and visually similar images on the Web, 2) and mining annotations from them. Firstly, at least one accurate keyword is required to enable text-based search for a set of semantically similar images. Then content-based search is performed on this set to retrieve visually similar images. At last, annotations are mined from the descriptions (titles, URLs and surrounding texts) of these images. It worth highlighting that to ensure the efficiency, high dimensional visual features are mapped to hash codes which significantly speed up the content-based search process. Our proposed approach enables annotating with unlimited vocabulary, which is impossible for all existing approaches. Experimental results on real web images show the effectiveness and efficiency of the proposed algorithm. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
CVPR (2) | 4 |
| 2006 | Exploring URL Hit Priors for Web Search
Ruihua Song, Guomao Xin, Shuming Shi 0001, Ji-Rong Wen, Wei-Ying Ma |
ECIR | 5 |
| 2006 | Ranking Web News Via Homepage Visual Layout and Cross-Site Voting
Jinyi Yao, Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma |
ECIR | 5 |
| 2006 | Fast Spectral Clustering of Data Using Sequential Matrix Compression
Bin Gao 0001, Tie-Yan Liu, Yufu Chen, Wei-Ying Ma |
ECML | 5 |
| 2006 | Automated known problem diagnosis with event tracesabstractComputer problem diagnosis remains a serious challenge to users and support professionals. Traditional troubleshooting methods relying heavily on human intervention make the process inefficient and the results inaccurate even for solved problems, which contribute significantly to user's dissatisfaction. We propose to use system behavior information such as system event traces to build correlations with solved problems, instead of using only vague text descriptions as in existing practices. The goal is to enable automatic identification of the root cause of a problem if it is a known one, which would further lead to its resolution. By applying statistical learning techniques to classifying system call sequences, we show our approach can achieve considerable accuracy of root cause recognition by studying four case examples. Chun Yuan 0004, Ni Lao, Ji-Rong Wen, Zheng Zhang 0001, Yi-Min Wang, Wei-Ying Ma |
EuroSys | 7 |
| 2006 | Extracting Objects from the WebabstractExtracting and integrating object information from the Web is of great significance for Web data management. The existing Web information extraction techniques cannot provide satisfactory solution to the Web object extraction task since objects of the same type are distributed in diverse Web sources, whose structures are highly heterogeneous. In this paper, we propose a novel approach called Object-Level Information Extraction (OLIE) to extract Web objects. This approach extends a classic information extraction algorithm, Conditional Random Fields (CRF), by adding Web-specific information. The experimental results show OLIE can significantly improve the Web object extraction accuracy. Zaiqing Nie, Fei Wu 0011, Ji-Rong Wen, Wei-Ying Ma |
ICDE | 4 |
| 2006 | Query Selection Techniques for Efficient Crawling of Structured Web SourcesabstractThe high quality, structured data from Web structured sources is invaluable for many applications. Hidden Web databases are not directly crawlable by Web search engines and are only accessible through Web query forms or via Web service interfaces. Recent research efforts have been focusing on understanding these Web query forms. A critical but still largely unresolved question is: how to efficiently acquire the structured information inside Web databases through iteratively issuing meaningful queries? In this paper we focus on the central issue of enabling efficient Web database crawling through query selection, i.e. how to select good queries to rapidly harvest data records from Web databases. We model each structured Web database as a distinct attribute-value graph. Under this theoretical framework, the database crawling problem is transformed into a graph traversal one that follows "relational" links. We show that finding an optimal query selection plan is equivalent to finding a Minimum Weighted Dominating Set of the corresponding database graph, a well-known NP-Complete problem. We propose a suite of query selection techniques aiming at optimizing the query harvest rate. Extensive experimental evaluations over real Web sources and simulations over controlled database servers validate the effectiveness of our techniques and provide insights for future efforts in this Ji-Rong Wen, Huan Liu 0001, Wei-Ying Ma |
ICDE | 4 |
| 2006 | Star-Structured High-Order Heterogeneous Data Co-clustering Based on Consistent Information TheoryabstractHeterogeneous object co-clustering has become an important research topic in data mining. In early years of this research, people mainly worked on two types of heterogeneous data (denoted by pair-wise co-clustering); while recently more and more attention was paid to multiple types of heterogeneous data (denoted by high- order co-clustering). In this paper, we studied the high- order co-clustering of objects with star-structured interrelationship, i.e., there is a central type of objects that connects the other types of objects. Actually, this case could be a very good model for many real-world applications, such as the co-clustering of Web images, their low-level visual features, and the surrounding text. We used a tripartite graph to represent the interrelationships among different objects, and proposed a consistent information theory which generates an effective algorithm to obtain the co-clusters of different types of objects. Experiments on a Web image show that our proposed algorithm is a better choice compared with previous work on heterogeneous object co-clustering. Bin Gao 0001, Tie-Yan Liu, Wei-Ying Ma |
ICDM | 3 |
| 2006 | Automatic Classification of Photographs and GraphicsabstractIn general, digital images can be classified into photographs and computer graphics. This taxonomy is very useful in many applications, such as Web image search. However, there are no effective methods to perform this classification automatically. In this paper, we manage to solve this problem from two aspects. At first, we propose some novel low-level features that can reveal perceptional differences between photographs and graphics. Then, we adopt an effective algorithm to perform the classification. The experiments conducted on a large-scale image database indicate the effectiveness of our algorithm Yuanhao Chen, Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma |
ICME | 4 |
| 2006 | Using Implicit Relevane Feedback to Advance Web Image SearchabstractAlthough relevance feedback has been extensively studied in content-based image retrieval in the academic area, no commercial Web image search engine has employed the idea. There are several obstacles for Web image search engines in applying relevance feedback. To overcome these obstacles, we proposed an efficient implicit relevance feedback mechanism. The proposed mechanism shows advantage over traditional relevance feedback methods in the following three aspects. Firstly, instead of enforcing the users to make explicit judgment on the results, our method regards user's click-through data as implicit relevance feedback which release burden from users. Secondly, a hierarchical image search results clustering algorithm is proposed to semantically organize the search results. Using the clustering results as features, our relevance feedback scheme could catch and reflect users' search intention precisely. Lastly, unlike traditional relevance feedback user interface which hardily substitutes subsequent results for previous ones, our method employed friendly recommendation rather than substitution to let the user narrow down on the refined images. To evaluate the implicit relevance feedback mechanism, comprehensive user studies were performed En Cheng, Mingjing Li, Wei-Ying Ma, Hai Jin 0001 |
ICME | 4 |
| 2006 | Large-Scale Duplicate Detection for Web Image SearchabstractFinding visually identical images in large image collections is important for many applications such as intelligence propriety protection and search result presentation. Several algorithms have been reported in the literature, but they are not suitable for large image collections. In this paper, a novel algorithm is proposed to handle the situation, in which each image is compactly represented by a hash code. To detect duplicate images, only the hash codes are required. In addition, a very efficient search method is implemented to quickly group images with similar hash codes for fast detection. The experiments show that our algorithm can be both efficient and effective for duplicate detection in Web image search Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma |
ICME | 4 |
| 2006 | Inquiring of the Sights from the Web via Camera MobilesabstractIn this paper, we presented an image search service for mobile users. It can be used to acquire related information by taking and sending pictures to the server, for example, getting book reviews by a photo of the cover. The key problem here is to find images that contain the same prominent object as that in the query image. In the literature, local feature based image matching has been proven to outperform those based on global features. When using local features, however, one query image may contain thousands of high dimensional feature vectors. Each feature vector needs to match against millions of features in the database. Therefore, it is critical to design an efficient search scheme. Our proposed matching approach was based on identifying semi-local visual parts from multiple query images. Experiments on two real-world datasets showed that this approach was superior to conventional solutions Yinghua Zhou, Xin Fan 0001, Xing Xie 0001, Yuchang Gong, Wei-Ying Ma |
ICME | 5 |
| 2006 | Event detection from evolution of click-through dataabstractPrevious efforts on event detection from the web have focused primarily on web content and structure data ignoring the rich collection of web log data. In this paper, we propose the first approach to detect events from the click-through data, which is the log data of web search engines. The intuition behind event detection from click-through data is that such data is often event-driven and each event can be represented as a set ofquery-page pairs that are not only semantically similar but also have similar evolution pattern over time. Given the click-through data, in our proposed approach, we first segment it into a sequence of bipartite graphs based on theuser-defined time granularity. Next, the sequence of bipartite graphs is represented as a vector-based graph, which records the semantic and evolutionary relationships between queries and pages. After that, the vector-based graph is transformed into its dual graph, where each node is a query-page pair that will be used to represent real world events. Then, the problem of event detection is equivalent to the problem of clustering the dual graph of the vector-based graph. The clustering process is based on a two-phase graph cut algorithm. In the first phase, query-page pairs are clustered based on thesemantic-based similarity such that each cluster in the result corresponds to a specific topic. In the second phase, query-page pairs related to the same topic are further clustered based on the evolution pattern-based similarity such that each cluster is expected to represent a specific event under the specific topic. Experiments with real click-through data collected from a commercial web search engine show that the proposed approach produces high quality results. Qiankun Zhao, Tie-Yan Liu, Sourav S. Bhowmick, Wei-Ying Ma |
KDD | 4 |
| 2006 | Simultaneous record detection and attribute labeling in web data extractionabstractRecent work has shown the feasibility and promise of templateindependent Web data extraction. However, existing approaches use decoupled strategies – attempting to do data record detection and attribute labeling in two separate phases. In this paper, we show that separately extracting data records and attributes is highly ineffective and propose a probabilistic model to perform these two tasks simultaneously. In our approach, record detection can benefit from the availability of semantics required in attribute labeling and, at the same time, the accuracy of attribute labeling can be improved when data records are labeled in a collective manner. The proposed model is called Hierarchical Conditional Random Fields. It can efficiently integrate all useful features by learning their importance, and it can also incorporate hierarchical interactions which are very important for Web data extraction. We empirically compare the proposed model with existing decoupled approaches for product information extraction, and the results show significant improvements in both record detection and attribute labeling. Jun Zhu 0001, Zaiqing Nie, Ji-Rong Wen, Bo Zhang 0010, Wei-Ying Ma |
KDD | 5 |
| 2006 | Detecting The Sufficient Display Resolution For Image BrowsingabstractIn image browsing, the resolution greatly affects user’s experience. If an image is down-scaled too much, a considerable amount of information within it will be lost. In this paper, we studied the problem of "What is a sufficient display resolution or scale for an image or an image region?" This problem arises in many real-life applications including image browsing on mobile devices, image adaptation and progressive image delivery. Kullback-Leibler (K-L) distance is employed to measure the information loss and the sufficient display scale is selected based on the information loss curve during image down-sampling. Since the images are presented to viewers finally, some visual characteristics are also taken into account to ensure the precision of the measurement. A user study was carried out to evaluate the performance of our approach. Experimental results show that the approach is in good accord with human perception. Xin Fan 0001, Xing Xie 0001, Wei-Ying Ma |
MDM | 3 |
| 2006 | Photo-to-Search: Using Camera Phones to Inquire of the Surrounding WorldabstractWith the pervasive use of camera phones, the embedded camera has been considered as a promising HCI manner for mobiles. With necessary technologies, it is possible to become a powerful tool to acquire the information in daily life. We have designed and implemented a system named Photo-to-Search to carry out queries from camera phones simply by taking some photos of interested objects. The captured pictures are compared with a large amount of Web images to select the ones which contain the same prominent object. Consequently, the related information is extracted from the Web pages where the matched images locate. In our demo, data of large buildings, storefronts and products are collected and these kinds of queries are specifically demonstrated to show the efficiency and the effectiveness of our system. Menglei Jia, Xin Fan 0001, Xing Xie 0001, Mingjing Li, Wei-Ying Ma |
MDM | 5 |
| 2006 | IGroup: web image search results clusteringabstractIn this paper, we propose, IGroup, an efficient and effective algorithm that organizes Web image search results into clusters. IGroup is different from all existing Web image search results clustering algorithms that only cluster the top few images using visual or textual features. Our proposed algorithm first identifies several query-related semantic clusters based on a key phrases extraction algorithm originally proposed for clustering general Web search results. Then, all the resulting images are separated and assigned to corresponding clusters. As a result, all the resulting images are organized into a clustering structure with semantic level. To make the best use of the clustering results, a new user interface (UI) is proposed. Different from existing Web image search interfaces, which show only a limited number of suggested query terms or representative image thumbnails of some clusters, the proposed interface displays both representative thumbnails and appropriate titles of semantically coherent image clusters. Comprehensive user studies have been completed to evaluate both the clustering algorithm and the new UI. Changhu Wang, Yuhuan Yao, Kefeng Deng, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 6 |
| 2006 | IGroup: a web image search engine with semantic clustering of search resultsabstractIn this demo, we present IGroup, a Web image search engine that organizes the search results into semantic clusters. Different from all existing Web image search results clustering algorithms that only cluster the top few images using visual or textual features, IGroup first identifies several query-related semantic clusters based on a key phrases extraction algorithm originally proposed for clustering general Web search results. Then, all the resulting images are separated and assigned to corresponding clusters. To make the best use of the clustering results, a new user interface is proposed. Please go to http://igroup.msra.cn for real experience. Changhu Wang, Yuhuan Yao, Kefeng Deng, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 6 |
| 2006 | VirtualTour: an online travel assistant based on high quality imagesabstractWith the popularity of both travel and Web, more and more people use online travel services to facilitate their travel activities or share their travel experiences. Considering that existing services emphasize more on the textual content with the pictorial content only as supplement, we propose the VirtualTour system. It is an online travel service dedicated on high quality images, which helps travelers plan their trip. The images of VirtualTour are from photo forum sites. They have rich and accurate metadata which could be used to extract geographic location information of them and assess the quality of them. A representative sights identification algorithm is also proposed to automatically identify the possible related sights of a region. Based on the map services that seamlessly integrated into the system, a wieldy UI is designed to support several useful features, e.g. query by map, location or path. Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2006 | Towards content-based relevance ranking for video searchabstractMost existing web video search engines index videos by file names, URLs, and surrounding texts. These types of video metadata roughly describe the whole video in an abstract level without taking the rich content, such as semantic content descriptions and speech within the video, into consideration. Therefore the relevance ranking of the video search results is not satisfactory as the details of video contents are ignored. In this paper we propose a novel relevance ranking approach for Web-based video search using both video metadata and the rich content contained in the videos. To leverage real content into ranking, the videos are segmented into shots, which are smaller and more semantic-meaningful retrievable units, and then more detailed information of video content such as semantic descriptions and speech of each shots are used to improve the retrieval and ranking performance. With video metadata and content information of shots, we developed an integrated ranking approach, which achieves improved ranking performance. We also introduce machine learning into the ranking system, and compare them with IR-model (information retrieval model) based method. The evaluation results demonstrate the effectiveness of the proposed ranking methods. Xian-Sheng Hua 0001, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2006 | Image annotation by large-scale content-based image retrievalabstractImage annotation has been an active research topic in recent years due to its potentially large impact on both image understanding and Web image search. In this paper, we target at solving the automatic image annotation problem in a novel search and mining framework. Given an uncaptioned image, first in the search stage, we perform content-based image retrieval (CBIR) facilitated by high-dimensional indexing to find a set of visually similar images from a large-scale image database. The database consists of images crawled from the World Wide Web with rich annotations, e.g. titles and surrounding text. Then in the mining stage, a search result clustering technique is utilized to find most representative keywords from the annotations of the retrieved image subset. These keywords, after salience ranking, are finally used to annotate the uncaptioned image. Based on search technologies, this framework does not impose an explicit training stage, but efficiently leverages large-scale and well-annotated images, and is potentially capable of dealing with unlimited vocabulary. Based on 2.4 million real Web images, comprehensive evaluation of image annotation on Corel and U. Washington image databases show the effectiveness and efficiency of the proposed approach. Xirong Li 0001, Lei Zhang 0001, Fuzong Lin, Wei-Ying Ma |
ACM Multimedia | 5 |
| 2006 | EnjoyPhoto: a vertical image search engine for enjoying high-quality photosabstractIn this paper, we propose building a vertical image search engine called EnjoyPhoto that leverages rich metadata from various photo forum web sites to meet users' requirements for enjoying high-quality photos, which is virtually impossible in traditional image search engines. To solve the ranking problem when aggregating multiple photo forums, we propose a novel rank fusion algorithm that uses duplicate photos to normalize rating scores. To further improve user experiences in enjoying photos, we design an in-place image browsing interface, and compare it with several other interfaces in a user study. With rich metadata and rating information, more attractive user interfaces are enabled, including slideshow authoring and photo recommendations. We conducted experiments and user studies on a 2.5-million image database to evaluate the proposed rank fusion algorithm, investigate the rationale behind building a vertical image search engine, and study user interfaces and preferences for the purpose of enjoying high-quality photos. The experimental results demonstrate the effectiveness of the proposed ranking algorithm. The results also show that the 2.5-million high-quality image database in EnjoyPhoto performs comparably with Google's 1- billion image database for queries related to location, nature, and daily life categories. Finally, our results show that the in-place browsing interface-called Force-Transfer view-is much more convenient for users than traditional interfaces. Lei Zhang 0001, Kefeng Deng, Wei-Ying Ma |
ACM Multimedia | 5 |
| 2006 | Study on texture feature extraction in region-based image retrieval systemabstractTexture is an important feature to describe images. Though lots of work has been done for efficient texture feature extraction from rectangular images, no much effort has been made in texture feature extraction from arbitrary-shaped regions in region-based image retrieval (RBIR) system. In this paper, we present an efficient texture feature extraction algorithm for arbitrary-shaped regions. This algorithm first extends an arbitrary-shaped region into a rectangular area onto which block transformation can be applied. Based on the projection-onto-convex-sets (POCS) theory, a set of coefficients best describing the original region are finally obtained, from which texture feature of the region can be extracted. Via intensive experiments, we select a set of parameters proper for image retrieval purpose. Experimental results on real-world image database demonstrate the effectiveness of the proposed algorithm for image retrieval purpose Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma |
MMM | 4 |
| 2006 | Level-Biased Statistics in the Hierarchical Structure of the Web
Tie-Yan Liu, Xudong Zhang 0001, Wei-Ying Ma |
PAKDD | 4 |
| 2006 | Heterogeneous Information Integration in Hierarchical Text Classification
Huai-Yuan Yang, Tie-Yan Liu, Wei-Ying Ma |
PAKDD | 4 |
| 2006 | A Systematic Study of Parameter Correlations in Large Scale Duplicate Document Detection
Shaozhi Ye, Ji-Rong Wen, Wei-Ying Ma |
PAKDD | 3 |
| 2006 | AggregateRank: bringing order to web sitesabstractSince the website is one of the most important organizational structures of the Web, how to effectively rank websites has been essential to many Web applications, such as Web search and crawling. In order to get the ranks of websites, researchers used to describe the inter-connectivity among websites with a so-called HostGraph in which the nodes denote websites and the edges denote linkages between websites (if and only if there are hyperlinks from the pages in one website to the pages in the other, there will be an edge between these two websites), and then adopted the random walk model in the HostGraph. However, as pointed in this paper, the random walk over such a HostGraph is not reasonable because it is not in accordance with the browsing behavior of web surfers. Therefore, the derivate rank cannot represent the true probability of visiting the corresponding website.In this work, we mathematically proved that the probability of visiting a website by the random web surfer should be equal to the sum of the PageRank values of the pages inside that website. Nevertheless, since the number of web pages is much larger than that of websites, it is not feasible to base the calculation of the ranks of websites on the calculation of PageRank. To tackle this problem, we proposed a novel method named AggregateRank rooted in the theory of stochastic complement, which cannot only approximate the sum of PageRank accurately, but also have a lower computational complexity than PageRank. Both theoretical analysis and experimental evaluation show that AggregateRank is a better method for ranking websites than previous methods. Tie-Yan Liu, Ying Bao, Zhiming Ma, Xudong Zhang 0001, Wei-Ying Ma |
SIGIR | 7 |
| 2006 | Building implicit links from content for forum searchabstractThe objective of Web forums is to create a shared space for open communications and discussions of specific topics and issues. The tremendous information behind forum sites is not fully-utilized yet. Most links between forum pages are automatically created, which means the link-based ranking algorithm cannot be applied efficiently. In this paper, we proposed a novel ranking algorithm which tries to introduce the content information into link-based methods as implicit links. The basic idea is derived from the more focused random surfer: the surfer may more likely jump to a page which is similar to what he is reading currently. In this manner, we are allowed to introduce the content similarities into the link graph as a personalization bias. Our method, named Fine-grained Rank (FGRank), can be efficiently computed based on an automatically generated topic hierarchy. Not like the topic-sensitive PageRank, our method only need to compute single PageRank score for each page. Another contribution of this paper is to present a very efficient algorithm for automatically generating topic hierarchy and map each page in a large-scale collection onto the computed hierarchy. The experimental results show that the proposed method can improve retrieval performance, and reveal that content-based link graph is also important compared with the hyper-link graph. Gu Xu, Wei-Ying Ma |
SIGIR | 2 |
| 2006 | Image annotation using search and mining technologiesabstractIn this paper, we present a novel solution to the image annotation problem which annotates images using search and data mining technologies. An accurate keyword is required to initialize this process, and then leveraging a large-scale image database, it 1) searches for semantically and visually similar images, 2) and mines annotations from them. A notable advantage of this approach is that it enables unlimited vocabulary, while it is not possible for all existing approaches. Experimental results on real web images show the effectiveness and efficiency of the proposed algorithm. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
WWW | 4 |
| 2006 | Time-dependent semantic similarity measure of queries using historical click-through dataabstractIt has become a promising direction to measure similarity of Web search queries by mining the increasing amount of click-through data logged by Web search engines, which record the interactions between users and the search engines. Most existing approaches employ the click-through data for similarity measure of queries with little consideration of the temporal factor, while the click-through data is often dynamic and contains rich temporal information. In this paper we present a new framework of time-dependent query semantic similarity model on exploiting the temporal characteristics of historical click-through data. The intuition is that more accurate semantic similarity values between queries can be obtained by taking into account the timestamps of the log data. With a set of user-defined calendar schema and calendar patterns, our time-dependent query similarity model is constructed using the marginalized kernel technique, which can exploit both explicit similarity and implicit semantics from the click-through data effectively. Experimental results on a large set of click-through data acquired from a commercial search engine show that our time-dependent query similarity model is more accurate than the existing approaches. Moreover, we observe that our time-dependent query similarity model can, to some extent, reflect real-world semantics such as real-world events that are happening over time. Qiankun Zhao, Steven C. H. Hoi, Tie-Yan Liu, Sourav S. Bhowmick, Michael R. Lyu, Wei-Ying Ma |
WWW | 6 |
| 2006 | A scalable supervised algorithm for dimensionality reduction on streaming data
Jun Yan 0001, Benyu Zhang, Shuicheng Yan, Ning Liu 0001, Qiang Yang 0001, Hua Li 0001, Zheng Chen 0001, Wei-Ying Ma |
Inf. Sci. | 9 |
| 2006 | Exploring statistical correlations for image retrieval
Xin-Jing Wang, Wei-Ying Ma, Xing Li 0001 |
Multim. Syst. | 2 |
| 2006 | A probabilistic semantic model for image annotation and multi-modal image retrieval
Ruofei Zhang, Zhongfei Zhang, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
Multim. Syst. | 4 |
| 2006 | Multitype Features Coselection for Web Document ClusteringabstractFeature selection has been widely applied in text categorization and clustering. Compared to unsupervised selection, supervised feature selection is more successful in filtering out noise in most cases. However, due to a lack of label information, clustering can hardly exploit supervised selection. Some studies have proposed to solve this problem by "pseudoclass." As empirical results show, this method is sensitive to selection criteria and data sets. In this paper, we propose a novel feature coselection for Web document clustering, which is called multitype features coselection for clustering (MFCC). MFCC uses intermediate clustering results in one type of feature space to help the selection in other types of feature spaces. Our experiments show that for most selection criteria, MFCC reduces effectively the noise introduced by "pseudoclass," and further improves clustering performance. Shen Huang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2006 | Design and Performance Studies of an Adaptive Scheme for Serving Dynamic Web Content in a Mobile Computing EnvironmentabstractCurrently, people gain easy access to an increasingly diverse range of mobile devices such as personal digital assistants (PDAs), smart phones, and handheld computers. As dynamic content has become dominant on the fast-growing World Wide Web (C. Yuan et al., 2003), it is necessary to provide effective ways for the users to access such prevalent Web content in a mobile computing environment. During a course of browsing dynamic content on mobile devices, the requested content is first dynamically generated by remote Web server, then transmitted over a wireless network, and, finally, adapted for display' on small screens. This leads to considerable latency and processing load on mobile devices. By integrating a novel Web content adaptation algorithm and an enhanced caching strategy, we propose an adaptive scheme called MobiDNA for serving dynamic content in a mobile computing environment. To validate the feasibility and effectiveness of the proposed MobiDNA system, we construct an experimental testbed to investigate its performance. Experimental results demonstrate that this scheme can effectively improve mobile dynamic content browsing, by improving Web content readability on small displays, decreasing mobile browsing latency, and reducing wireless bandwidth consumption Zhigang Hua, Xing Xie 0001, Hao Liu 0007, Hanqing Lu, Wei-Ying Ma |
IEEE Trans. Mob. Comput. | 5 |
| 2006 | Browsing Large Pictures Under Limited Display SizesabstractPictures have become increasingly common and popular in mobile communications. However, due to the limitation of mobile devices, there is a need to develop new technologies to facilitate the browsing of large pictures on the small screen. In this paper, we propose a set of novel approaches which are able to aid or automate common image browsing tasks on mobile devices. All of these approaches are based on an image attention model which is employed to illustrate the information structure within an image. An efficient algorithm to generate the optimal browsing path based on the information foraging theory is presented. Experimental evaluations of the proposed mechanism indicate that our approach is an effective way for viewing large images on small displays Xing Xie 0001, Hao Liu 0007, Wei-Ying Ma, HongJiang Zhang |
IEEE Trans. Multim. | 3 |
| 2006 | TSSP: Multi-features based reinforcement algorithm to find related papers
Shen Huang, Yong Yu 0001, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Wei-Ying Ma |
Web Intell. Agent Syst. | 6 |
| 2005 | Level-Based Link Analysis
Tie-Yan Liu, Xudong Zhang 0001, Tao Qin 0001, Bin Gao 0001, Wei-Ying Ma |
APWeb | 6 |
| 2005 | Supervised Semi-definite Embedding for Email Data Cleaning and Visualization
Ning Liu 0001, Fengshan Bai, Jun Yan 0001, Benyu Zhang, Zheng Chen 0001, Wei-Ying Ma |
APWeb | 6 |
| 2005 | A Similarity Reinforcement Algorithm for Heterogeneous Web Pages
Ning Liu 0001, Jun Yan 0001, Fengshan Bai, Benyu Zhang, Wensi Xi, Weiguo Fan, Zheng Chen 0001, Lei Ji 0001, Chenyong Hu, Wei-Ying Ma |
APWeb | 10 |
| 2005 | Learning user interest for image browsing on small-form-factor devicesabstractMobile devices which can capture and view pictures are becoming increasingly common in our life. The limitation of these small-form-factor devices makes the user experience of image browsing quite different from that on desktop PCs. In this paper, we first present a user study on how users interact with a mobile image browser with basic functions. We found that on small displays, users tend to use more zooming and scrolling actions in order to view interesting regions in detail. From this fact, we designed a new method to detect user interest maps and extract user attention objects from the image browsing log. This approach is more efficient than image-analysis based methods and can better represent users' actual interest. A smart image viewer was then developed based on user interest analysis. A second experiment was carried out to study how users behave with such a viewer. Experimental results demonstrate that the new smart features can improve the browsing efficiency and are a good compliment to traditional image browsers. Xing Xie 0001, Hao Liu 0007, Simon Goumaz, Wei-Ying Ma |
CHI | 4 |
| 2005 | Hybrid index structures for location-based web searchabstractThere is more and more commercial and research interest in location-based web search, i.e. finding web content whose topic is related to a particular place or region. In this type of search, location information should be indexed as well as text information. However, the index of conventional text search engine is set-oriented, while location information is two-dimensional and in Euclidean space. This brings new research problems on how to efficiently represent the location attributes of web pages and how to combine two types of indexes. In this paper, we propose to use a hybrid index structure, which integrates inverted files and R*-trees, to handle both textual and location aware queries. Three different combining schemes are studied: (1) inverted file and R*-tree double index, (2) first inverted file then R*-tree, (3) first R*-tree then inverted file. To validate the performance of proposed index structures, we design and implement a complete location-based web search engine which mainly consists of four parts: (1) an extractor which detects geographical scopes of web pages and represents geographical scopes as multiple MBRs based on geographical coordinates; (2) an indexer which builds hybrid index structures to integrate text and location information; (3) a ranker which ranks results by geographical relevance as well as non-geographical relevance; (4) an interface which is friendly for users to input location-based search queries and to obtain geographical and textual relevant results. Experiments on large real-world web dataset show that both the second and the third structures are superior in query time and the second is slightly better than the third. Additionally, indexes based on R*-trees are proven to be more efficient than indexes based on grid structures. Yinghua Zhou, Xing Xie 0001, Chuang Wang 0001, Yuchang Gong, Wei-Ying Ma |
CIKM | 5 |
| 2005 | A Unified Optimization Based Learning Method for Image RetrievalabstractIn this paper, an optimization based learning method is proposed for image retrieval from graph model point of view. Firstly, image retrieval is formulated as a regularized optimization problem, which simultaneously considers the constraints from low-level feature, online relevance feedback and offline semantic information. Then, the global optimal solution is developed in both closed form and iterative form, providing that the latter converges to the former. The proposed method is unified in the senses that 1) it makes use of the information from various aspects in a global optimization manner so that the retrieval performance might be maximally improved; 2) it provides a natural way to support two typical query scenarios in image retrieval. The proposed method has a solid mathematical ground. Systematic experimental results on a general-purpose image database demonstrate that it achieves significant improvements over existing methods. Hanghang Tong, Jingrui He, Mingjing Li, Wei-Ying Ma, Changshui Zhang, HongJiang Zhang |
CVPR (2) | 4 |
| 2005 | Deriving High-Level Concepts Using Fuzzy-ID3 Decision Tree for Image RetrievalabstractTo improve the retrieval accuracy of content-based image retrieval, an important task is to reduce the 'semantic gap' between low-level image features and the richness of human semantics. We present a region-based image retrieval system using high-level semantic concepts. The contribution of the paper is two-fold. First, salient low-level features are extracted from arbitrarily-shaped regions. Second, a fuzzy-ID3 decision tree learning method is proposed to derive association rules which map low-level image features to high-level concepts. Experimental results prove that, by reducing the 'semantic gap', the proposed system not only improves the retrieval accuracy, but also supports users in query-by-keyword. Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma |
ICASSP (2) | 4 |
| 2005 | A Probabilistic Semantic Model for Image Annotation and Multi-Modal Image RetrievaabstractThis paper addresses automatic image annotation problem and its application to multi-modal image retrieval. The contribution of our work is three-fold. (1) We propose a probabilistic semantic model in which the visual features and the textual words are connected via a hidden layer which constitutes the semantic concepts to be discovered to explicitly exploit the synergy among the modalities. (2) The association of visual features and textual words is determined in a Bayesian framework such that the confidence of the association can be provided. (3) Extensive evaluation on a large-scale, visually and semantically diverse image collection crawled from Web is reported to evaluate the prototype system based on the model. In the proposed probabilistic model, a hidden concept layer which connects the visual feature and the word layer is discovered by fitting a generative model to the training image and annotation words through an Expectation-Maximization (EM) based iterative learning procedure. The evaluation of the prototype system on 17,000 images and 7,736 automatically extracted annotation words from crawled Web pages for multi-modal image retrieval has indicated that the proposed semantic model and the developed Bayesian framework are superior to a state-of-the-art peer system in the literature. Ruofei Zhang, Zhongfei Zhang, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
ICCV | 4 |
| 2005 | Automatic Annotation of Location Information for WWW ImagesabstractCurrently, a crucial challenge is raised on how to manage a large amount of images on the Web. Due to a real synergy between an image and its location, we propose an automatic solution to annotate contextual location information for WWW images. We construct an image importance model to acquire the dominant images in a page that comprise contextual surrounding text. For each acquired image, we develop an effective algorithm to compute location from its contextual text. We apply our approach to 1,000 pages from various Websites for image location annotation. The experiments demonstrated that more than 30% WWW images are related with geographic location information, and our solution can achieve the satisfactory results. Finally, we present some potential applications involving the utilization of image location information Zhigang Hua, Chuang Wang 0001, Xing Xie 0001, Hanqing Lu, Wei-Ying Ma |
ICME | 5 |
| 2005 | Natural Image Retrieval with SketchesabstractIn this paper, we present a method to retrieve natural images by sketch query. To measure the similarity between the sketch and an image, relevant regions are first located in that image through a multi-resolution search, and a normalized local shape similarity is proposed for image retrieval. Efficiency and other implementation issues are discussed. Experimental results show that it is an effective approach for content-based image retrieval Jinyi Yao, Mingjing Li, Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma |
ICME | 5 |
| 2005 | Supervised semi-definite embedding for image manifoldsabstractSemi-definite embedding (SDE) has been a recently proposed to maximize the sum of pair wise squared distances between outputs while the input data and outputs are locally isometric, i.e. it pulls the outputs as far apart as possible, subject to unfolding a manifold without any furling or fold for unsupervised nonlinear dimensionality reduction. The extensions of SDE to supervised feature extraction, named as supervised Semi-definite embedding (SSDE) was proposed by the authors of this paper. Here, the method is unified in a mathematical framework and applied to a number of benchmark data sets. Results show that SSDE performs very well on high-dimensional data, which exhibits a manifold structure. Benyu Zhang, Jun Yan 0001, Ning Liu 0001, Zheng Chen 0001, Wei-Ying Ma |
ICME | 6 |
| 2005 | Auto cropping for digital photographsabstractIn this paper, we propose an effective approach to the nearly untouched problem, still photograph auto cropping, which is one of the important features to automatically enhance photographs. To obtain an optimal result, we first formulate auto cropping as an optimization problem by defining an energy function, which consists of three sub models: composition sub model, conservative sub model, and penalty sub model. Then, particle swarm optimization (PSO) is employed to obtain the optimal solution by maximizing the objective function. Experimental results and user studies over hundreds of photographs show that the proposed approach is effective and accurate in most cases, and can be used in many practical multimedia applications. Mingju Zhang, Lei Zhang 0001, Wei-Ying Ma |
ICME | 5 |
| 2005 | 2D Conditional Random Fields for Web information extractionabstractThe Web contains an abundance of useful semistructured information about real world objects, and our empirical study shows that strong sequence characteristics exist for Web information about objects of the same type across different Web sites. Conditional Random Fields (CRFs) are the state of the art approaches taking the sequence characteristics to do better labeling. However, as the information on a Web page is two-dimensionally laid out, previous linear-chain CRFs have their limitations for Web information extraction. To better incorporate the two-dimensional neighborhood interactions, this paper presents a two-dimensional CRF model to automatically extract object information from the Web. We empirically compare the proposed model with existing linear-chain CRF models for product information extraction, and the results show the effectiveness of our model. Jun Zhu 0001, Zaiqing Nie, Ji-Rong Wen, Bo Zhang 0010, Wei-Ying Ma |
ICML | 5 |
| 2005 | Consistent bipartite graph co-partitioning for star-structured high-order heterogeneous data co-clusteringabstractHeterogeneous data co-clustering has attracted more and more attention in recent years due to its high impact on various applications. While the co-clustering algorithms for two types of heterogeneous data (denoted by pair-wise co-clustering), such as documents and terms, have been well studied in the literature, the work on more types of heterogeneous data (denoted by high-order co-clustering) is still very limited. As an attempt in this direction, in this paper, we worked on a specific case of high-order co-clustering in which there is a central type of objects that connects the other types so as to form a star structure of the inter-relationships. Actually, this case could be a very good abstract for many real-world applications, such as the co-clustering of categories, documents and terms in text mining. In our philosophy, we treated such kind of problems as the fusion of multiple pair-wise co-clustering sub-problems with the constraint of the star structure. Accordingly, we proposed the concept of consistent bipartite graph co-partitioning, and developed an algorithm based on semi-definite programming (SDP) for efficient computation of the clustering results. Experiments on toy problems and real data both verified the effectiveness of our proposed method. Bin Gao 0001, Tie-Yan Liu, Wei-Ying Ma |
KDD | 5 |
| 2005 | Web object indexing using domain knowledgeabstractA web object is defined to represent any meaningful object embedded in web pages (e.g. images, music) or pointed to by hyperlinks (e.g. downloadable files). In many cases, users would like to search for information of a certain 'object', rather than a web page containing the query terms. To facilitate web object searching and organizing, in this paper, we propose a novel approach to web object indexing, by discovering its inherent structure information with existed domain knowledge. In our approach, first, Layered LSI spaces are built for a better representation of the hierarchically structured domain knowledge, in order to emphasize the specific semantics and term space in each layer of the domain knowledge. Meanwhile, the web object representation is constructed by hyperlink analysis, and further pruned to remove the noises. Then an optimal matching between the web object and the domain knowledge is performed, in order to pick out the structure attributes of the web object from the knowledge. Finally, the obtained structure attributes are used to re-organize and index the web objects. Our approach also indicates a new promising way to use trust-worthy Deep Web knowledge to help organize dispersive information of Surface Web. Muyuan Wang, Zhiwei Li 0006, Lie Lu, Wei-Ying Ma, Naiyao Zhang |
KDD | 4 |
| 2005 | AIRE: an ambient interactive and responsive environment for mobile image managementabstractThis paper proposes an ambient interactive and responsive environment (AIRE) to improve user's image browsing experiences on the small-form-factor devices. This solution is characterized as two aspects: (1) establishing an ambient interactive communication across various devices; and (2) designing a distributed interface that crosses various devices to overcome the display constraint in mobile devices. In the AIRE system, a two-level image browsing scheme is designed to meet users' various image browsing needs. Zhigang Hua, Xiang-Jun Wang, Xing Xie 0001, Qingshan Liu 0001, Hanqing Lu, Wei-Ying Ma |
Mobile HCI | 6 |
| 2005 | Web image clustering by consistent utilization of visual features and surrounding textsabstractImage clustering, an important technology for image processing, has been actively researched for a long period of time. Especially in recent years, with the explosive growth of the Web, image clustering has even been a critical technology to help users digest the large amount of online visual information. However, as far as we know, many previous works on image clustering only used either low-level visual features or surrounding texts, but rarely exploited these two kinds of information in the same framework. To tackle this problem, we proposed a novel method named consistent bipartite graph co-partitioning in this paper, which can cluster Web images based on the consistent fusion of the information contained in both low-level features and surrounding texts. In particular, we formulated it as a constrained multi-objective optimization problem, which can be efficiently solved by semi-definite programming (SDP). Experiments on a real-world Web image collection showed that our proposed method outperformed the methods only based on low-level features or surround texts. Bin Gao 0001, Tie-Yan Liu, Tao Qin 0001, Wei-Ying Ma |
ACM Multimedia | 6 |
| 2005 | Graph based multi-modality learningabstractTo better understand the content of multimedia, a lot of research efforts have been made on how to learn from multi-modal feature. In this paper, it is studied from a graph point of view: each kind of feature from one modality is represented as one independent graph; and the learning task is formulated as inferring from the constraints in every graph as well as supervision information (if available). For semi-supervised learning, two different fusion schemes, namely linear form and sequential form, are proposed. For each scheme, it is derived from optimization point of view; and further justified from two sides: similarity propagation and Bayesian interpretation. By doing so, we reveal the regular optimization nature, transductive learning nature as well as prior fusion nature of the proposed schemes, respectively. Moreover, the proposed method can be easily extended to unsupervised learning, including clustering and embedding. Systematic experimental results validate the effectiveness of the proposed method. Hanghang Tong, Jingrui He, Mingjing Li, Changshui Zhang, Wei-Ying Ma |
ACM Multimedia | 5 |
| 2005 | Iteratively clustering web images based on link and attribute reinforcementsabstractImage clustering is an important research topic which contributes to a wide range of applications. Traditional image clustering approaches are based on image content features only, while content features alone can hardly describe the semantics of the images. In the context of Web, images are no longer assumed homogeneous and "flatdistributed but are richly structured. There are two kinds of reinforcements embedded in such data: 1) the reinforcement between attributes of different data types (intra-type links reinforcements); and 2) the reinforcement between object attributes and the inter-type links (inter-type links reinforcements). Unfortunately, most of the previous works addressing relational data failed to fully explore the reinforcements. In this paper, we propose a reinforcement clustering framework to tackle this problem. It reinforces images and texts' attributes via inter-type links and inversely uses these attributes to update these links. The iterative reinforcing nature of this framework promises the discovery of the semantic structure of images, which is the basis of image clustering. Experimental results show the effectiveness of our proposed framework. Xin-Jing Wang, Wei-Ying Ma, Lei Zhang 0001, Xing Li 0001 |
ACM Multimedia | 2 |
| 2005 | ASAP: A Synchronous Approach for Photo Sharing across Multiple DevicesabstractDigital photos have become increasingly common and popular in mobile communications. However, due to the distribution of these photos captured in various devices, there is a need to develop new technologies to facilitate the sharing of large image collections across these devices for users. In this paper, we propose A Synchronous Approach for Photo sharing across multiple devices (ASAP). The ASAP provides a hierarchical two-level synchronization scheme, namely image-level and region-level. In the ASAP, a user’s interaction with any device automatically leads to a series of synchronous updates in other devices. Thus, the ASAP simultaneously presents similar images across devices in a way that allows automatic synchronization of images based on user interactions. Experimental evaluations indicate that it is effective and useful to improve image browsing experience across devices for users. Zhigang Hua, Xing Xie 0001, Hanqing Lu, Wei-Ying Ma |
MMM | 4 |
| 2005 | Grouping WWW Image Search Results by Novel Inhomogeneous Clustering MethodabstractIn this paper, a novel inhomogeneous clustering method is proposed for grouping web images. It is used to re-organize the search result of web image search engines into a hierarchical structure so that the users can conveniently browse the search result. This method takes into account various features associated with web images, and treats them in different ways. For the surrounding text extracted from the containing web pages, co-clustering approach is adopted; for low-level features of the image content and other features, one-way clustering approach is adopted. The clustering results of different approaches are combined together to produce the final image groups. Experimental results demonstrate the effectiveness of the proposed method. Zhiwei Li 0006, Gu Xu, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
MMM | 4 |
| 2005 | Effective Feature Extraction for Play Detection in American Football VideoabstractThe fact that a typical broadcast can last over 3 hours for a game of 60 minutes makes video summarization of American football games most desirable. In this paper, we present several feature extraction methods for play detection in American football video. Wavelet based motion analysis is used to extract the trend component from the noisy motion vectors; a hybrid field-color model detects field area with both high accuracy and fast speed; and a prior knowledge driven line detection method uses the court information to estimate miss-detections. Based on the so-extracted features, a boosting chain is used for feature selection and decision making. Tested on large-size video data, the detection performance of our work is very promising. Tie-Yan Liu, Wei-Ying Ma, HongJiang Zhang |
MMM | 2 |
| 2005 | Region-Based Image Retrieval with High-Level Semantic Color NamesabstractPerformance of traditional content-based image retrieval systems is far from user’s expectation due to the ‘semantic gap’ between low-level visual features and the richness of human semantics. In attempt to reduce the ‘semantic gap’, this paper introduces a region-based image retrieval system with high-level semantic color names. In this system, database images are segmented into color-texture homogeneous regions. For each region, we define a color name as that used in our daily life. In the retrieval process, images containing regions of same color name as that of the query are selected as candidates. These candidate images are further ranked based on their color and texture features. In this way, the system reduces the ‘semantic gap’ between numerical image features and the rich semantics in the user’s mind. Experimental results show that the proposed system provides promising retrieval results with few features used. Ying Liu 0026, Dengsheng Zhang, Guojun Lu, Wei-Ying Ma |
MMM | 4 |
| 2005 | Sports Video Mining with MosaicabstractVideo is an information-intensive media with much redundancy. Therefore, it is desirable to be able to mine structure or semantics of video data for efficient browsing, summarization and highlight extraction. In this paper, we propose a generic approach to key-event as well as structure mining for sports video analysis. Mosaic is generated for each shot as the representative image of shot content. Based on mosaic, sports video is mined by the method with prior knowledge and without prior knowledge. Without prior knowledge, our system may locate plays by discriminating those segments without essential content, such as breaks. If prior knowledge is available, the key-events in plays are detected using robust features extracted from mosaic. Experimental results have demonstrated the effectiveness and robustness of this sports video mining approach. Tao Mei 0001, Yufei Ma 0006, He-Qin Zhou, Wei-Ying Ma, HongJiang Zhang |
MMM | 4 |
| 2005 | Subspace Clustering and Label Propagation for Active Feedback in Image RetrievalabstractIn recent years, relevance feedback has been studied extensively as a way to improve performance of content-based image retrieval (CBIR). However, since users are usually unwilling to provide many feedbacks, the insufficiency of the training samples limited the success of relevance feedback. To tackle this problem, we propose two coupled algorithms: (i) overlapped subspace clustering to select representative images for user’s feedback; and (ii) multi-subspace label propagation to include unlabeled data in the training process. As these two algorithms are both working on sub feature spaces of the image database, they can not only deal with the insufficient training samples but also well capture the user’s attention during the retrieval process. Experimental results on a large database of general-purposed images demonstrated the high effectiveness of our proposed algorithms. Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, Wei-Ying Ma, HongJiang Zhang |
MMM | 4 |
| 2005 | Learning No-Reference Quality Metric by ExamplesabstractIn this paper, a novel learning based method is proposed for No-Reference image quality assessment. Instead of examining the exact prior knowledge for the given type of distortion and finding a suitable way to represent it, our method aims to directly get the quality metric by means of learning. At first, some training examples are prepared for both high-quality and low-quality classes; then a binary classifier is built on the training set; finally the quality metric of an un-labeled example is denoted by the extent to which it belongs to these two classes. Different schemes to acquire examples from a given image, to build the binary classifier and to model the quality metric are proposed and investigated. While most existing methods are tailored for some specific distortion type, the proposed method might provide a general solution for No-Reference image quality assessment. Experimental results on JPEG and JPEG2000 compressed images validate the effectiveness of the proposed method. Hanghang Tong, Mingjing Li, HongJiang Zhang, Changshui Zhang, Jingrui He, Wei-Ying Ma |
MMM | 6 |
| 2005 | Realistic 3D Face Modeling by Fusing Multiple 2D ImagesabstractIn this paper, we propose a fully automatic and efficient algorithm for realistic 3D face reconstruction by fusing multiple 2D face images. Firstly, an efficient multi-view 2D face alignment algorithm is utilized to localize the facial points of the face images; and then the intrinsic shape and texture models are inferred by the proposed Syncretized Shape Model (SSM) and Syncretized Texture Model (STM), respectively. Compared with other related works, our proposed algorithm has the following characteristics: 1) the inferred shape and texture are more realistic owing to the constraints and co-enhancement among the multiple images; 2) it is fully automatic, without any user interaction; and 3) the shape and pose parameter estimation is efficient via EM approach and unit quaternion based pose representation, and is also robust as a result of the dynamic correspondence approach. The experimental results show the effectiveness of our proposed algorithm for 3D face reconstruction. Changhu Wang, Shuicheng Yan, HongJiang Zhang, Wei-Ying Ma |
MMM | 4 |
| 2005 | Parallel Image Matrix Compression for Face RecognitionabstractThe canonical face recognition algorithm Eigenface and Fisherface are both based on one dimensional vector representation. However, with the high feature dimensions and the small training data, face recognition often suffers from the curse of dimension and the small sample problem. Recent research [4] shows that face recognition based on direct 2D matrix representation, i.e. 2DPCA, obtains better performance than that based on traditional vector representation. However, there are three questions left unresolved in the 2DPCA algorithm: I ) what is the meaning of the eigenvalue and eigenvector of the covariance matrix in 2DPCA; 2) why 2DPCA can outperform Eigenface; and 3) how to reduce the dimension after 2DPCA directly. In this paper, we analyze 2DPCA in a different view and proof that is 2DPCA actually a "localized" PCA with each row vector of an image as object. With this explanation, we discover the intrinsic reason that 2DPCA can outperform Eigenface is because fewer feature dimensions and more samples are used in 2DPCA when compared with Eigenface. To further reduce the dimension after 2DPCA, a two-stage strategy, namely parallel image matrix compression (PIMC), is proposed to compress the image matrix redundancy, which exists among row vectors and column vectors. The exhaustive experiment results demonstrate that PIMC is superior to 2DPCA and Eigenface, and PIMC+LDA outperforms 2DPC+LDA and Fisherface. Dong Xu 0001, Shuicheng Yan, Lei Zhang 0001, Mingjing Li, Wei-Ying Ma, Zhengkai Liu, HongJiang Zhang |
MMM | 5 |
| 2005 | Efficient Browsing of Web Search Results on Mobile Devices Based on Block Importance ModelabstractIt is expected that more and more people would search the Web when they are on the move. Though conventional search engines can be directly visited from mobile devices with Web browsing capabilities, the information is not as conveniently accessible from a handheld device as it is from desktops. Existing information discovery mechanisms for searching the Web are not well-suited to mobile devices. In this paper, a block importance model is employed to assign importance values to different segments of a Web page, in order to extract and present more condensed search results to mobile users. Based on the block importance model, three presentations for displaying the result pages in different levels of detail have been designed to reduce both the number of user interactions and the overall search time. A set of user study experiments have been carried out to compare the three presentations and a commercial service on typical mobile devices. Experimental results show that our approaches can help users to explore Web search results more efficiently. Xing Xie 0001, Gengxin Miao, Ruihua Song, Ji-Rong Wen, Wei-Ying Ma |
PerCom | 5 |
| 2005 | A probabilistic model for retrospective news event detectionabstractRetrospective news event detection (RED) is defined as the discovery of previously unidentified events in historical news corpus. Although both the contents and time information of news articles are helpful to RED, most researches focus on the utilization of the contents of news articles. Few research works have been carried out on finding better usages of time information. In this paper, we do some explorations on both directions based on the following two characteristics of news articles. On the one hand, news articles are always aroused by events; on the other hand, similar articles reporting the same event often redundantly appear on many news sources. The former hints a generative model of news articles, and the latter provides data enriched environments to perform RED. With consideration of these characteristics, we propose a probabilistic model to incorporate both content and time information in a unified framework. This model gives new representations of both news articles and news events. Furthermore, based on this approach, we build an interactive RED system, HISCOVERY, which provides additional functions to present events, Photo Story and Chronicle. Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma |
SIGIR | 4 |
| 2005 | A study of relevance propagation for web searchabstractDifferent from traditional information retrieval, both content and structure are critical to the success of Web information retrieval. In recent years, many relevance propagation techniques have been proposed to propagate content information between web pages through web structure to improve the performance of web search. In this paper, we first propose a generic relevance propagation framework, and then provide a comparison study on the effectiveness and efficiency of various representative propagation models that can be derived from this generic framework. We come to many conclusions that are useful for selecting a propagation model in real-world search applications, including 1) sitemap-based propagation models outperform hyperlink-based models in sense of both effectiveness and efficiency, and 2) sitemap-based term propagation is easier to be integrated into real-world search engines because of its parallel offline implementation and acceptable complexity. Some other more detailed study results are also reported in the paper. Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, Zheng Chen 0001, Wei-Ying Ma |
SIGIR | 5 |
| 2005 | Gravitation-based model for information retrievalabstractThis paper proposes GBM (gravitation-based model), a physical model for information retrieval inspired by Newton's theory of gravitation. A mapping is built in this model from concepts of information retrieval (documents, queries, relevance, etc) to those of physics (mass, distance, radius, attractive force, etc). This model actually provides a new perspective on IR problems. A family of effective term weighting functions can be derived from it, including the well-known BM25 formula. This model has some advantages over most existing ones: First, because it is directly based on basic physical laws, the derived formulas and algorithms can have their explicit physical interpretation. Second, the ranking formulas derived from this model satisfy more intuitive heuristics than most of existing ones, thus have the potential to behave empirically better and to be used safely on various settings. Finally, a new approach for structured document retrieval derived from this model is more reasonable and behaves better than existing ones. Shuming Shi 0001, Ji-Rong Wen, Ruihua Song, Wei-Ying Ma |
SIGIR | 5 |
| 2005 | Detecting dominant locations from search queriesabstractAccurately and effectively detecting the locations where search queries are truly about has huge potential impact on increasing search relevance. In this paper, we define a search query's dominant location (QDL) and propose a solution to correctly detect it. QDL is geographical location(s) associated with a query in collective human knowledge, i.e., one or few prominent locations agreed by majority of people who know the answer to the query. QDL is a subjective and collective attribute of search queries and we are able to detect QDLs from both queries containing geographical location names and queries not containing them. The key challenges to QDL detection include false positive suppression (not all contained location names in queries mean geographical locations), and detecting implied locations by the context of the query. In our solution, a query is recursively broken into atomic tokens according to its most popular web usage for reducing false positives. If we do not find a dominant location in this step, we mine the top search results and/or query logs (with different approaches discussed in this paper) to discover implicit query locations. Our large-scale experiments on recent MSN Search queries show that our query location detection solution has consistent high accuracy for all query frequency ranges. Lee Wang, Chuang Wang 0001, Xing Xie 0001, Joshua J. Forman, Yansheng Lu, Wei-Ying Ma, Ying Li 0012 |
SIGIR | 6 |
| 2005 | OCFS: optimal orthogonal centroid feature selection for text categorizationabstractText categorization is an important research area in many Information Retrieval (IR) applications. To save the storage space and computation time in text categorization, efficient and effective algorithms for reducing the data before analysis are highly desired. Traditional techniques for this purpose can generally be classified into feature extraction and feature selection. Because of efficiency, the latter is more suitable for text data such as web documents. However, many popular feature selection techniques such as Information Gain (IG) andχ2-test (CHI) are all greedy in nature and thus may not be optimal according to some criterion. Moreover, the performance of these greedy methods may be deteriorated when the reserved data dimension is extremely low. In this paper, we propose an efficient optimal feature selection algorithm by optimizing the objective function of Orthogonal Centroid (OC) subspace learning algorithm in a discrete solution space, called Orthogonal Centroid Feature Selection (OCFS). Experiments on 20 Newsgroups (20NG), Reuters Corpus Volume 1 (RCV1) and Open Directory Project (ODP) data show that OCFS is consistently better than IG and CHI with smaller computation time especially when the reduced dimension is extremely small. Jun Yan 0001, Ning Liu 0001, Benyu Zhang, Shuicheng Yan, Zheng Chen 0001, Weiguo Fan, Wei-Ying Ma |
SIGIR | 8 |
| 2005 | Improving web search results using affinity graphabstractIn this paper, we propose a novel ranking scheme named Affinity Ranking (AR) to re-rank search results by optimizing two metrics: (1) diversity -- which indicates the variance of topics in a group of documents; (2) information richness -- which measures the coverage of a single document to its topic. Both of the two metrics are calculated from a directed link graph named Affinity Graph (AG). AG models the structure of a group of documents based on the asymmetric content similarities between each pair of documents. Experimental results in Yahoo! Directory, ODP Data, and Newsgroup data demonstrate that our proposed ranking algorithm significantly improves the search performance. Specifically, the algorithm achieves 31% improvement in diversity and 12% improvement in information richness relatively within the top 10 search results. Benyu Zhang, Hua Li 0001, Yi Liu 0054, Lei Ji 0001, Wensi Xi, Weiguo Fan, Zheng Chen 0001, Wei-Ying Ma |
SIGIR | 8 |
| 2005 | Webpage Importance Analysis Using Conditional Markov Random WalkabstractIn this paper, we propose a novel method to calculate the Web page importance based on a conditional Markov random walk model. The main assumption in this model is that given the hyperlinks in a Web page, users are not really randomly clicking one of them. Instead, many factors may bias their behaviors, for example, the anchor text, the content relevance and the previous experiences when visiting the Web site that a destination page belongs to. As one of the results, the user might tend to visit those pages in high-quality Web sites with higher probability. To implement this idea, we reformulate the Web graph to be a two-layer structure, and the Web page importance is calculated by conditional random walk in this new Web graph. Experiments on the topic distillation task of TREC 2003 Web track showed that our new method can achieve about 18% improvement on mean average precision (MAP) and 16% on precision at 10 (P@10) over the PageRank algorithm. Tie-Yan Liu, Wei-Ying Ma |
Web Intelligence | 2 |
| 2005 | An Editor Labeling Model for Training Set Expansion in Web CategorizationabstractAutomatically classifying Web pages is an effective way to manage the massive information on the Web. However, our experiments show that the state-of-the-art text categorization technologies can not achieve a satisfactory classification performance in this task. The major reason is the existence of large proportion of rare categories in Web taxonomies. The failure in such categories is simply because there is not enough information to train reliable classifiers. To tackle this problem, we propose to expand the training set of the rare categories, by simulating the labeling behavior of the human editors of Web directories. Experimental results show that in such a way, we achieved significant (relatively 93%) improvement in classification accuracy, which is highly encouraging for high performance Web classification. Tie-Yan Liu, Hao Wan 0003, Wei-Ying Ma |
Web Intelligence | 3 |
| 2005 | Object-level ranking: bringing order to Web objectsabstractIn contrast with the current Web search methods that essentially do document-level ranking and retrieval, we are exploring a new paradigm to enable Web search at the object level. We collect Web information for objects relevant for a specific application domain and rank these objects in terms of their relevance and popularity to answer user queries. Traditional PageRank model is no longer valid for object popularity calculation because of the existence of heterogeneous relationships between objects. This paper introduces PopRank, a domain-independent object-level link analysis model to rank the objects within a specific domain. Specifically we assign a popularity propagation factor to each type of object relationship, study how different popularity propagation factors for these heterogeneous relationships could affect the popularity ranking, and propose efficient approaches to automatically decide these factors. Our experiments are done using 1 million CS papers, and the experimental results show that PopRank can achieve significantly better ranking results than naively applying PageRank on the object graph. Zaiqing Nie, Ji-Rong Wen, Wei-Ying Ma |
WWW | 4 |
| 2005 | An adaptive web page layout structure for small devices
Xing Xie 0001, Chong Wang 0002, Li-Qun Chen, Wei-Ying Ma |
Multim. Syst. | 4 |
| 2005 | Hierarchical Taxonomy Preparation for Text Categorization Using Consistent Bipartite Spectral Graph CopartitioningabstractMulticlass classification has been investigated for many years in the literature. Recently, the scales of real-world multiclass classification applications have become larger and larger. For example, there are hundreds of thousands of categories employed in the Open Directory Project (ODP) and the Yahoo! directory. In such cases, the scalability of classification methods turns out to be a major concern. To tackle this problem, hierarchical classification is proposed and widely adopted to get better trade-off between effectiveness and efficiency. Unfortunately, many data sets are not explicitly organized in hierarchical forms and, therefore, hierarchical classification cannot be used directly. In this paper, we propose a novel algorithm to automatically mine a hierarchical structure from the flat taxonomy of a data corpus as a preparation for the adoption of hierarchical classification. In particular, we first compute matrices to represent the relations among categories, documents, and terms. And, then, we cocluster the three substances at different scales through consistent bipartite spectral graph copartitioning, which is formulated as a generalized singular value decomposition problem. At last, a hierarchical taxonomy is constructed from the category clusters. Our experiments showed that the proposed algorithm could discover very reasonable taxonomy hierarchy and help improve the classification accuracy. Bin Gao 0001, Tie-Yan Liu, Tao Qin 0001, Wei-Ying Ma |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2004 | A Query-Dependent Duplicate Detection Approach for Large Scale Search Engines
Shaozhi Ye, Ruihua Song, Ji-Rong Wen, Wei-Ying Ma |
APWeb | 4 |
| 2004 | Learning similarity measures in non-orthogonal spaceabstractMany machine learning and data mining algorithms crucially rely on the similarity metrics. The Cosine similarity, which calculates the inner product of two normalized feature vectors, is one of the most commonly used similarity measures. However, in many practical tasks such as text categorization and document clustering, the Cosine similarity is calculated under the assumption that the input space is an orthogonal space which usually could not be satisfied due to synonymy and polysemy. Various algorithms such as Latent Semantic Indexing (LSI) were used to solve this problem by projecting the original data into an orthogonal space. However LSI also suffered from the high computational cost and data sparseness. These shortcomings led to increases in computation time and storage requirements for large scale realistic data. In this paper, we propose a novel and effective similarity metric in the non-orthogonal input space. The basic idea of our proposed metric is that the similarity of features should affect the similarity of objects, and vice versa. A novel iterative algorithm for computing non-orthogonal space similarity measures is then proposed. Experimental results on a synthetic data set, a real MSN search click-thru logs, and 20NG dataset show that our algorithm outperforms the traditional Cosine similarity and is superior to LSI. Ning Liu 0001, Benyu Zhang, Jun Yan 0001, Qiang Yang 0001, Shuicheng Yan, Zheng Chen 0001, Fengshan Bai, Wei-Ying Ma |
CIKM | 8 |
| 2004 | Web page clustering enhanced by summarizationabstractTraditional Web page clustering algorithms use the full-text in the documents to generate feature vectors. Such methods often produce unsatisfactory results because there is much noisy information, such as decoration, interaction, and advertisement, in Web pages. The varying-length problem of the Web pages is also a significant negative factor affecting the performance. In this paper, we investigate the use of several summarization techniques to tackle these issues when clustering Web pages. Compared with the full-text representation of the Web pages, our experimental results indicate that our proposed approach effectively solves the problems of noisy information and varying-length, and thus significantly boosts the clustering performance. Xuanhui Wang, Dou Shen, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma |
CIKM | 5 |
| 2004 | Optimizing web search using web click-through dataabstractThe performance of web search engines may often deteriorate due to the diversity and noisy information contained within web pages. User click-through data can be used to introduce more accurate description (metadata) for web pages, and to improve the search performance. However, noise and incompleteness, sparseness, and the volatility of web pages and are three major challenges for research work on user click-through log mining. In this paper, we propose a novel iterative reinforced algorithm to utilize the user click-through data to improve search performance. The algorithm fully explores the interrelations between and web pages, and effectively finds virtual queries for web pages and overcomes the challenges discussed above. Experiment results on a large set of MSN click-through log data show a significant improvement on search performance over the naive query log mining algorithm as well as the baseline search engine. Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Weiguo Fan |
CIKM | 5 |
| 2004 | MRSSA: an iterative algorithm for similarity spreading over interrelated objectsabstractWe introduce the Multiple Relationship Similarity Spreading Algorithm (MRSSA) to enhance IR effectiveness. This method has similarity computed in an iterative "spreading" fashion for multiple object types, combining both inter- and intra-object relationships. We demonstrate the value of this approach in the context of the WWW, where the key objects are web pages and queries, Relationships considered are derived from hyperlinks (in- and out-links) and click-through logs. Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Edward A. Fox |
CIKM | 5 |
| 2004 | Mining Ratio Rules Via Principal Sparse Non-Negative Matrix FactorizationabstractAssociation rules are traditionally designed to capture statistical relationship among itemsets in a given database. To additionally capture the quantitative association knowledge, Korn et al. (1998) proposed a paradigm named ratio rules for quantifiable data mining. However, their approach is mainly based on principle component analysis (PCA) and as a result, it cannot guarantee that the ratio coefficient is nonnegative. This may lead to serious problems in the rules' application. In this paper, we propose a method, called principal sparse nonnegative matrix factorization (PSNMF), for learning the associations between itemsets in the form of ratio rules. In addition, we provide a support measurement to weigh the importance of each rule for the entire dataset. Chenyong Hu, Benyu Zhang, Shuicheng Yan, Qiang Yang 0001, Jun Yan 0001, Zheng Chen 0001, Wei-Ying Ma |
ICDM | 7 |
| 2004 | Improving Text Classification using Local Latent Semantic IndexingabstractLatent semantic indexing (LSI) has been shown to be extremely useful in information retrieval, but it is not an optimal representation for text classification. It always drops the text classification performance when being applied to the whole training set (global LSI) because this completely unsupervised method ignores class discrimination while only concentrating on representation. Some local LSI methods have been proposed to improve the classification by utilizing class discrimination information. However, their performance improvements over original term vectors are still very limited. In this paper, we propose a new local LSI method called "local relevancy weighted LSI" to improve text classification by performing a separate single value decomposition (SVD) on the transformed local region of each class. Experimental results show that our method is much better than global LSI and traditional local LSI methods on classification within a much smaller LSI dimension. Zheng Chen 0001, Benyu Zhang, Wei-Ying Ma, Gongyi Wu |
ICDM | 4 |
| 2004 | Supervised Latent Semantic Indexing for Document CategorizationabstractLatent semantic indexing (LSI) is a successful technology in information retrieval (IR) which attempts to explore the latent semantics implied by a query or a document through representing them in a dimension-reduced space. However, LSI is not optimal for document categorization tasks because it aims to find the most representative features for document representation rather than the most discriminative ones. In this paper, we propose supervised LSI (SLSI) which selects the most discriminative basis vectors using the training data iteratively. The extracted vectors are then used to project the documents into a reduced dimensional space for better classification. Experimental evaluations show that the SLSI approach leads to dramatic dimension reduction while achieving good classification results. Jian-Tao Sun, Zheng Chen 0001, Hua-Jun Zeng, Yuchang Lu, Chun-Yi Shi, Wei-Ying Ma |
ICDM | 6 |
| 2004 | IRC: An Iterative Reinforcement Categorization Algorithm for Interrelated Web ObjectsabstractMost existing categorization algorithms deal with homogeneous Web data objects, and consider interrelated objects as additional features when taking the interrelationships with other types of objects into account. However, focusing on any single aspects of these interrelationships and objects does not fully reveal their true categories. In this paper, we propose a categorization algorithm, the iterative reinforcement categorization algorithm (IRC), to exploit the full interrelationships between the heterogeneous objects on the Web. IRC attempts to classify the interrelated Web objects by iterative reinforcement between individual classification results of different types via the interrelationships. Experiments on a clickthrough log dataset from MSN search engine show that, with the Fl measures, IRC achieves a 26.4% improvement over a pure content-based classification method, a 21% improvement over a query metadata-based method, and a 16.4% improvement over a virtual document-based method. Furthermore, our experiments show that IRC converges rapidly. Gui-Rong Xue, Dou Shen, Qiang Yang 0001, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wensi Xi, Wei-Ying Ma |
ICDM | 8 |
| 2004 | Organizing WWW images based on the analysis of page layout and Web link structureabstractDue to the rapid growth of the number of digital images on the Web, there is an increasing demand for an effective and efficient method of organizing and retrieving the images available. This paper describes a method for clustering and embedding WWW images. By using a vision-based page segmentation algorithm, a Web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. By extracting the page-to-block, block-to-image, block-to-page relationships through a link structure and page layout analysis, we construct an image graph. With the image graph model, we use techniques from spectral graph theory for image clustering and embedding. Some experimental results are given in the paper. Deng Cai 0001, Xiaofei He 0001, Wei-Ying Ma, Ji-Rong Wen, HongJiang Zhang |
ICME | 3 |
| 2004 | Attention model based progressive image transmissionabstractProgressive image transmission provides a convenient user interface when images are transmitted slowly. However, most of the existing PIT techniques only consider the objective quality of the reconstructed image. Here we present an attention model based progressive image transmission approach to improve the subjective quality of the transmission process. We use both bottom-up image features and top-down semantic information to extract the regions of interest and also propose a new ROI coding scheme based on JPEG2000 to control the trade-off between the transmission of ROI and background. Experiments have shown the efficiency of our approach. Yusuo Hu, Xing Xie 0001, Zonghai Chen, Wei-Ying Ma |
ICME | 4 |
| 2004 | A content-based bit allocation model for video streamingabstractTraditional video coding and streaming algorithms have not been fully utilized by using video content information for bit allocation and adaptation. This paper proposes a content-based video streaming method, based on a visual attention model to better utilize network bandwidth and achieve better subjective video quality. First, the visual attention model is exploited to segment the regions of interest (ROI) in video frames. Then, considering that the ROI is more sensitive to coding error than other regions, a region-weighted rate-distortion model is developed to allocate suitable bits for all ROI and non-ROI regions. The evaluation results indicate that more than 60% of the test video sequences encoded by the proposed method can obtain better subjective visual quality compared to the video encoded by classical methods under the same bandwidth, and about 20% of them can not be distinguished. Renhua Wang, Wei-Ying Ma, HongJiang Zhang |
ICME | 4 |
| 2004 | Extracting texture features from arbitrary-shaped regions for image retrievalabstractLots of work has been done in texture feature extraction for rectangular images, but not as much attention has been paid to the arbitrary-shaped regions available in region-based image retrieval (RBIR) systems. In This work, we present a texture feature extraction algorithm, based on projection onto convex sets (POCS) theory. POCS iteratively concentrates more and more energy into the selected coefficients from which texture features of an arbitrary-shaped region can be extracted. Experimental results demonstrate the effectiveness of the proposed algorithm for image retrieval purposes. Ying Liu 0026, Xiaofang Zhou 0001, Wei-Ying Ma |
ICME | 3 |
| 2004 | Data-driven approach for bridging the cognitive gap in image retrievalabstractBridging the cognitive gap in image retrieval has been an active research direction in recent years. Existing solutions typically require a large volume of training data that could be difficult to obtain in practice. In this paper, we propose a data-driven approach that uses Web images and their surrounding textual annotations as the source of training data to bridge the cognitive gap. We construct an image thesaurus that contains a set of codewords, each representing a semantically related subspace in the feature space. We also explore the use of query expansion based on the constructed image thesaurus for improving image retrieval performance. Xin-Jing Wang, Wei-Ying Ma, Xing Li 0001 |
ICME | 2 |
| 2004 | Maximizing information throughput for multimedia browsing on small displaysabstractAs a great many new devices with diverse capabilities are having a population boom, their limited display sizes become the major obstacle that has undermined their usefulness for information access. We introduce our recent research on adapting multimedia content, including images, videos and Web pages, for browsing on small-form-factor devices. A theoretical framework as well as a set of novel methods for presenting and rendering multimedia under limited screen sizes is proposed to improve the user experience. The content modeling and processing are provided as subscription-based Web services on the Internet. Experiments show that our approach is extensible and able to achieve satisfactory results with high efficiency. Xing Xie 0001, Wei-Ying Ma, HongJiang Zhang |
ICME | 2 |
| 2004 | IMMC: incremental maximum margin criterionabstractSubspace learning approaches have attracted much attention in academia recently. However, the classical batch algorithms no longer satisfy the applications on streaming data or large-scale data. To meet this desirability, Incremental Principal Component Analysis (IPCA) algorithm has been well established, but it is an unsupervised subspace learning approach and is not optimal for general classification tasks, such as face recognition and Web document categorization. In this paper, we propose an incremental supervised subspace learning algorithm, called Incremental Maximum Margin Criterion (IMMC), to infer an adaptive subspace by optimizing the Maximum Margin Criterion. We also present the proof for convergence of the proposed algorithm. Experimental results on both synthetic dataset and real world datasets show that IMMC converges to the similar subspace as that of batch approach. Jun Yan 0001, Benyu Zhang, Shuicheng Yan, Qiang Yang 0001, Hua Li 0001, Zheng Chen 0001, Wensi Xi, Weiguo Fan, Wei-Ying Ma |
KDD | 9 |
| 2004 | Combining High Level Symptom Descriptions and Low Level State Information for Configuration Fault Diagnosis
Ni Lao, Ji-Rong Wen, Wei-Ying Ma, Yi-Min Wang |
LISA | 3 |
| 2004 | Hierarchical clustering of WWW image search results using visual, textual and link informationabstractWe consider the problem of clustering Web image search results. Generally, the image search results returned by an image search engine contain multiple topics. Organizing the results into different semantic clusters facilitates users' browsing. In this paper, we propose a hierarchical clustering method using visual, textual and link analysis. By using a vision-based page segmentation algorithm, a web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. By using block-level link analysis techniques, an image graph can be constructed. We then apply spectral techniques to find a Euclidean embedding of the images which respects the graph structure. Thus for each image, we have three kinds of representations, i.e. visual feature based representation, textual feature based representation and graph based representation. Using spectral clustering techniques, we can cluster the search results into different semantic clusters. An image search example illustrates the potential of these techniques. Deng Cai 0001, Xiaofei He 0001, Zhiwei Li 0006, Wei-Ying Ma, Ji-Rong Wen |
ACM Multimedia | 4 |
| 2004 | Learning an image manifold for retrievalabstractWe consider the problem of learning a mapping function from low-level feature space to high-level semantic space. Under the assumption that the data lie on a submanifold embedded in a high dimensional Euclidean space, we propose a relevance feedback scheme which is naturally conducted only on the image manifold in question rather than the total ambient space. While images are typically represented by feature vectors in Rn, the natural distance is often different from the distance induced by the ambient space Rn. The geodesic distances on manifold are used to measure the similarities between images. However, when the number of data points is small, it is hard to discover the intrinsic manifold structure. Based on user interactions in a relevance feedback driven query-by-example system, the intrinsic similarities between images can be accurately estimated. We then develop an algorithmic framework to approximate the optimal mapping function by a Radial Basis Function (RBF) neural network. The semantics of a new image can be inferred by the RBF neural network. Experimental results show that our approach is effective in improving the performance of content-based image retrieval systems. Xiaofei He 0001, Wei-Ying Ma, HongJiang Zhang |
ACM Multimedia | 2 |
| 2004 | Intuitive and effective interfaces for WWW image search enginesabstractWeb image search engine has become an important tool to organize digital images on the Web. However, most commercial search engines still use a list presentation while little effort has been placed on improving their usability. How to present the image search results in a more intuitive and effective way is still an open question to be carefully studied. In this demo, we present iFind, a scalable Web image search engine, in which we integrated two kinds of search result browsing interfaces. User study results have proved that our interfaces are superior to traditional interfaces. Zhiwei Li 0006, Xing Xie 0001, Hao Liu 0007, Xiaoou Tang, Mingjing Li, Wei-Ying Ma |
ACM Multimedia | 6 |
| 2004 | Grouping web image search resultabstractIn this paper, we propose a Web image search result organizing method to facilitate user browsing. We formalize this problem as a salient image region pattern extraction problem. Given the images returned by Web search engine, we first segment the images into homogeneous regions and quantize the environmental regions into image codewords. The salient codeword "phrases" are then extracted and ranked based on a regression model learned from human labeled training data. According to the salient "phrases", images are assigned to different clusters, with the one nearest to the centroid as the entry for the corresponding cluster. Satisfying experimental results show the effectiveness of our proposed method. Xin-Jing Wang, Wei-Ying Ma, Qi-Cai He, Xing Li 0001 |
ACM Multimedia | 2 |
| 2004 | Multi-model similarity propagation and its application for web image retrievalabstractIn this paper, we propose an iterative similarity propagation approach to explore the inter-relationships between Web images and their textual annotations for image retrieval. By considering Web images as one type of objects, their surrounding texts as another type, and constructing the links structure between them via webpage analysis, we can iteratively reinforce the similarities between images. The basic idea is that if two objects of the same type are both related to one object of another type, these two objects are similar; likewise, if two objects of the same type are related to two different, but similar objects of another type, then to some extent, these two objects are also similar. The goal of our method is to fully exploit the mutual reinforcement between images and their textual annotations. Our experiments based on 10,628 images crawled from the Web show that our proposed approach can significantly improve Web image retrieval performance. Xin-Jing Wang, Wei-Ying Ma, Gui-Rong Xue, Xing Li 0001 |
ACM Multimedia | 2 |
| 2004 | Efficient propagation for face annotation in family albumsabstractIn this paper, we propose and investigate a new user scenario for face annotation, in which users are allowed to multi-select a group of photographs and assign names to these photographs. The system will then attempt to propagate names from photograph level to face level, i.e. to infer the correspondence between name and face. Given the face similarity measure which combines methodologies from face recognition and content-based image retrieval, we formulate name propagation as an optimization problem. We define the objective function as the sum of similarities between each pair of faces of the same individual in different photographs, and propose an iterative optimization algorithm to infer the optimal correspondence. To make the propagation result reliable, a reject scheme is adopted to reject those with low confidence scores. Furthermore, we investigate the combination and alternation of browsing mode for propagation and viewer mode for annotation, so that each mode can benefit from additional inputs from the other mode. The experimental evaluation has been conducted within a typical family album of over one thousand photographs and the results show that the proposed approach is effective and efficient in automated face annotation in family albums. Lei Zhang 0001, Yuxiao Hu 0001, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
ACM Multimedia | 4 |
| 2004 | Locality preserving clustering for image databaseabstractIt is important and challenging to make the growing image repositories easy to search and browse. Image clustering is a technique that helps in several ways, including image data preprocessing, user interface designing, and search result representation. Spectral clustering method has been one of the most promising clustering methods in the last few years, because it can cluster data with complex structure, and the (near) global optimum is guaranteed. However, existing spectral clustering algorithms, like Normalized Cut, are difficult to handle data points out of training set. In this paper, we propose a clustering algorithm named Locality Preserving Clustering (LPC), which shares many of the data representation properties of nonlinear spectral method. Yet LPC provides an explicit mapping function which is defined everywhere, both on training data points and testing points. Experimental results show that LPC is more accurate than both "direct Kmeans" and "PCA + Kmeans". We also show that LPC produces in general comparable results with Normalized Cut, yet is more efficient than Normalized Cut. Deng Cai 0001, Xiaofei He 0001, Wei-Ying Ma, Xueyin Lin |
ACM Multimedia | 4 |
| 2004 | Adaptive Content Delivery on Mobile Internet across Multiple Form FactorabstractSummary form only given. In the PC+ era, a variety of new mobile devices, such as Web-enabled watch, Smartphone, Pocket PC, Tablet PC, etc, are making a population boom. As most of current documents and multimedia contents are designed for applications on desktop PC, we are facing a big challenge - the problem of multiple form factors resulted from the diverse set of emerging new devices. For example, browsing conventional Web pages on Pocket PC is a pain. Browsing a large picture on Smartphone (e.g. picture messaging application) is no different to a thumbnail view, even though the users are now able to capture high-resolution pictures on the latest models of camera-equipped cell phones. Furthermore, enterprise software applications have been aiming to take the advantage of ubiquitous network connection by mobile Internet to increase the productivity of information workers. As the current document-related applications are highly desktop-PC focused, to unleash the power of mobile Internet in the PC+ era, we need to innovate more aggressively and develop a new generation of document and content technologies to cope with the problem of multiple form factors. In this paper, we briefly review the past research and development on adapting multimedia content for heterogeneous device capability, user context and network environment, including the MPEG-7 standard. Then, we focus on the challenge of limited display and input capability in mobile devices, we believe that in order to successfully tackle this challenge, we need to (a) develop new content representation which is scalable and adaptable, (b) develop content analysis techniques to structure content for the new representation, (c) develop corresponding rendering algorithms and innovative user interfaces to maximize the information throughput of the content on small devices, and (d) make content adaptation service part of mobile Internet architecture. We also introduce some specific research works that we have been conducting towards this direction in Microsoft Research Asia. These works include Scalable Web Document which is a new Web content representation for various display sizes, Smart Picture which uses image attention model to facilitate the browsing of large pictures on mobile devices, and Media Companion which uses proxy augmentation on the edge of mobile Internet to construct a content service overlay for dynamic adaptation and value-added services. HongJiang Zhang, Wei-Ying Ma |
MMM | 2 |
| 2004 | Block-level link analysisabstractLink Analysis has shown great potential in improving the performance of web search. PageRank and HITS are two of the most popular algorithms. Most of the existing link analysis algorithms treat a web page as a single node in the web graph. However, in most cases, a web page contains multiple semantics and hence the web page might not be considered as the atomic node. In this paper, the web page is partitioned into blocks using the vision-based page segmentation algorithm. By extracting the page-to-block, block-to-page relationships from link structure and page layout analysis, we can construct a semantic graph over the WWW such that each node exactly represents a single semantic topic. This graph can better describe the semantic structure of the web. Based on block-level link analysis, we proposed two new algorithms, Block Level PageRank and Block Level HITS, whose performances we study extensively using web data. Deng Cai 0001, Xiaofei He 0001, Ji-Rong Wen, Wei-Ying Ma |
SIGIR | 4 |
| 2004 | Block-based web searchabstractMultiple-topic and varying-length of web pages are two negative factors significantly affecting the performance of web search. In this paper, we explore the use of page segmentation algorithms to partition web pages into blocks and investigate how to take advantage of block-level evidence to improve retrieval performance in the web context. Because of the special characteristics of web pages, different page segmentation method will have different impact on web search performance. We compare four types of methods, including fixed-length page segmentation, DOM-based page segmentation, vision-based page segmentation, and a combined method which integrates both semantic and fixed-length properties. Experiments on block-level query expansion and retrieval are performed. Among the four approaches, the combined method achieves the best performance for web search. Our experimental results also show that such a semantic partitioning of web pages effectively deals with the problem of multiple drifting topics and mixed lengths, and thus has great potential to boost up the performance of current web search engines. Deng Cai 0001, Shipeng Yu, Ji-Rong Wen, Wei-Ying Ma |
SIGIR | 4 |
| 2004 | Locality preserving indexing for document representationabstractDocument representation and indexing is a key problem for document analysis and processing, such as clustering, classification and retrieval. Conventionally, Latent Semantic Indexing (LSI) is considered effective in deriving such an indexing. LSI essentially detects the most representative features for document representation rather than the most discriminative features. Therefore, LSI might not be optimal in discriminating documents with different semantics. In this paper, a novel algorithm called Locality Preserving Indexing (LPI) is proposed for document indexing. Each document is represented by a vector with low dimensionality. In contrast to LSI which discovers the global structure of the document space, LPI discovers the local structure and obtains a compact document representation subspace that best detects the essential semantic structure. We compare the proposed LPI approach with LSI on two standard databases. Experimental results show that LPI provides better representation in the sense of semantic structure. Xiaofei He 0001, Deng Cai 0001, Haifeng Liu 0001, Wei-Ying Ma |
SIGIR | 4 |
| 2004 | Web-page classification through summarizationabstractWeb-page classification is much more difficult than pure-text classification due to a large variety of noisy information embedded in Web pages. In this paper, we propose a new Web-page classification algorithm based on Web summarization for improving the accuracy. We first give empirical evidence that ideal Web-page summaries generated by human editors can indeed improve the performance of Web-page classification algorithms. We then propose a new Web summarization-based classification algorithm and evaluate it along with several other state-of-the-art text summarization algorithms on the LookSmart Web directory. Experimental results show that our proposed summarization-based classification algorithm achieves an approximately 8.8% improvement as compared to pure-text-based classification algorithm. We further introduce an ensemble classifier using the improved summarization algorithm and show that it achieves about 12.9% improvement over pure-text based methods. Dou Shen, Zheng Chen 0001, Qiang Yang 0001, Hua-Jun Zeng, Benyu Zhang, Yuchang Lu, Wei-Ying Ma |
SIGIR | 7 |
| 2004 | Probabilistic model for contextual retrievalabstractContextual retrieval is a critical technique for facilitating many important applications such as mobile search, personalized search, PC troubleshooting, etc. Despite of its importance, there is no comprehensive retrieval model to describe the contextual retrieval process. We observed that incompatible context, noisy context and incomplete query are several important issues commonly existing in contextual retrieval applications. However, these issues have not been previously explored and discussed. In this paper, we propose probabilistic models to address these problems. Our study clearly shows that query log is the key to build effective contextual retrieval models. We also conduct a case study in the PC troubleshooting domain to testify the performance of the proposed models and experimental results show that the models can achieve very good retrieval precision. Ji-Rong Wen, Ni Lao, Wei-Ying Ma |
SIGIR | 3 |
| 2004 | Learning to cluster web search resultsabstractOrganizing Web search results into clusters facilitates users' quick browsing through search results. Traditional clustering techniques are inadequate since they don't generate clusters with highly readable names. In this paper, we reformalize the clustering problem as a salient phrase ranking problem. Given a query and the ranked list of documents (typically a list of titles and snippets) returned by a certain Web search engine, our method first extracts and ranks salient phrases as candidate cluster names, based on a regression model learned from human labeled training data. The documents are assigned to relevant salient phrases to form candidate clusters, and the final clusters are generated by merging these candidate clusters. Experimental results verify our method's feasibility and effectiveness. Hua-Jun Zeng, Qi-Cai He, Zheng Chen 0001, Wei-Ying Ma, Jinwen Ma |
SIGIR | 4 |
| 2004 | Collapse-to-zoom: viewing web pages on small screen devices by interactively removing irrelevant contentabstractOverview visualizations for small-screen web browsers were designed to provide users with visual context and to allow them to rapidly zoom in on tiles of relevant content. Given that content in the overview is reduced, however, users are often unable to tell which tiles hold the relevant material, which can force them to adopt a time-consuming hunt-and-peck strategy. Collapse-to-zoom addresses this issue by offering an alternative exploration strategy. In addition to allowing users to zoom into relevant areas, collapse-to-zoom allows users to collapse areas deemed irrelevant, such as columns containing menus, archive material, or advertising. Collapsing content causes all remaining content to expand in size causing it to reveal more detail, which increases the user's chance of identifying relevant content. Collapse-to-zoom navigation is based on a hybrid between a marquee selection tool and a marking menu, called marquee menu. It offers four commands for collapsing content areas at different granularities and to switch to a full-size reading view of what is left of the page. Patrick Baudisch, Xing Xie 0001, Chong Wang 0002, Wei-Ying Ma |
UIST | 4 |
| 2004 | Instance-based Schema Matching for Web Databases by Domain-specific Query Probing
Jiying Wang, Ji-Rong Wen, Frederick H. Lochovsky, Wei-Ying Ma |
VLDB | 4 |
| 2004 | TSSP: A Reinforcement Algorithm to Find Related PapersabstractContent analysis and citation analysis are two common methods in recommending system. Compared with content analysis, citation analysis can discover more implicitly related papers. However, the citation-based methods may introduce more noise in citation graph and cause topic drift. Some work combine content with citation to improve similarity measurement. The problem is that the two features are not used to reinforce each other to get better result. To solve the problem, we propose a new algorithm, Topic Sensitive Similarity Propagation (TSSP), to effectively integrate content similarity into similarity propagation. TSSP has two parts: citation context based propagation and iterative reinforcement. First, citation contexts provide clues for which papers are topic related to and filter out less irrelevant citations. Second, iteratively integrating content and citation similarity enable them to reinforce each other during the propagation. The experimental results of a user study show TSSP outperforms other algorithms in almost all cases. Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma |
Web Intelligence | 6 |
| 2004 | GE-CKO: A Method to Optimize Composite Kernels for Web Page ClassificationabstractMost of current researches on Web page classification focus on leveraging heterogeneous features such as plain text, hyperlinks and anchor texts in an effective and efficient way. Composite kernel method is one topic of interest among them. It first selects a bunch of initial kernels, each of which is determined separately by a certain type of features. Then a classifier is trained based on a linear combination of these kernels. In this paper, we propose an effective way to optimize the linear combination of kernels. We proved that this problem is equivalent to solving a generalized eigenvalue problem. And the weight vector of the kernels is the eigenvector associated with the largest eigen-value. A support vector machine (SVM) classifier is then trained based on this optimized combination of kernels. Our experiment on the WebKB dataset has shown the effectiveness of our proposed method. Jian-Tao Sun, Benyu Zhang, Zheng Chen 0001, Yuchang Lu, Chunyi Shi, Wei-Ying Ma |
Web Intelligence | 6 |
| 2004 | Multi-type Features Based Web Document Clustering
Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma |
WISE | 6 |
| 2004 | Exploiting PageRank at Different Block Level
Xue-Mei Jiang, Gui-Rong Xue, Wen-Guan Song, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma |
WISE | 6 |
| 2004 | Towards Next Generation Web Information Retrieval
Wei-Ying Ma, HongJiang Zhang, Hsiao-Wuen Hon |
WISE | 1 |
| 2004 | Optimizing Web Search Using Spreading Activation on the Clickthrough Data
Gui-Rong Xue, Shen Huang, Yong Yu 0001, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma |
WISE | 6 |
| 2004 | Learning block importance models for web pagesabstractPrevious work shows that a web page can be partitioned into multiple segments or blocks, and often the importance of those blocks in a page is not equivalent. Also, it has been proven that differentiating noisy or unimportant blocks from pages can facilitate web mining, search and accessibility. However, no uniform approach and model has been presented to measure the importance of different segments in web pages. Through a user study, we found that people do have a consistent view about the importance of blocks in web pages. In this paper, we investigate how to find a model to automatically assign importance values to blocks in a web page. We define the block importance estimation as a learning problem. First, we use a vision-based page segmentation algorithm to partition a web page into semantic blocks with a hierarchical structure. Then spatial features (such as position and size) and content features (such as the number of images and links) are extracted to construct a feature vector for each block. Based on these features, learning algorithms are used to train a model to assign importance to different segments in the web page. In our experiments, the best model can achieve the performance with Micro-F1 79% and Micro-Accuracy 85.9%, which is quite close to a person's view. Ruihua Song, Haifeng Liu 0001, Ji-Rong Wen, Wei-Ying Ma |
WWW | 4 |
| 2004 | Link fusion: a unified link analysis framework for multi-type interrelated data objectsabstractWeb link analysis has proven to be a significant enhancement for quality based web search. Most existing links can be classified into two categories: intra-type links (e.g., web hyperlinks), which represent the relationship of data objects within a homogeneous data type (web pages), and inter-type links (e.g., user browsing log) which represent the relationship of data objects across different data types (users and web pages). Unfortunately, most link analysis research only considers one type of link. In this paper, we propose a unified link analysis framework, called "link fusion", which considers both the inter- and intra- type link structure among multiple-type inter-related data objects and brings order to objects in each data type at the same time. The PageRank and HITS algorithms are shown to be special cases of our unified link analysis framework. Experiments on an instantiation of the framework that makes use of the user data and web pages extracted from a proxy log show that our proposed algorithm could improve the search effectiveness over the HITS and DirectHit algorithms by 24.6% and 38.2% respectively. Wensi Xi, Benyu Zhang, Zheng Chen 0001, Yizhou Lu, Shuicheng Yan, Wei-Ying Ma, Edward A. Fox |
WWW | 6 |
| 2003 | Extracting Content Structure for Web Pages Based on Visual Representation
Deng Cai 0001, Shipeng Yu, Ji-Rong Wen, Wei-Ying Ma |
APWeb | 4 |
| 2003 | CBC: Clustering Based Text Classification Requiring Minimal Labeled DataabstractSemisupervised learning methods construct classifiers using both labeled and unlabeled training data samples. While unlabeled data samples can help to improve the accuracy of trained models to certain extent, existing methods still face difficulties when labeled data is not sufficient and biased against the underlying data distribution. We present a clustering based classification (CBC) approach. Using this approach, training data, including both the labeled and unlabeled data, is first clustered with the guidance of the labeled data. Some of unlabeled data samples are then labeled based on the clusters obtained. Discriminative classifiers can subsequently be trained with the expanded labeled dataset. The effectiveness of the proposed method is justified analytically. Our experimental results demonstrated that CBC outperforms existing algorithms when the size of labeled dataset is very small. Hua-Jun Zeng, Xuanhui Wang, Zheng Chen 0001, Hongjun Lu, Wei-Ying Ma |
ICDM | 5 |
| 2003 | Visual attention based image browsing on mobile devicesabstractImages have become more and more common in mobile communications. People now can easily take and exchange pictures on the move using their mobile devices and digital cameras. However, a crucial challenge is to provide a better user experience for browsing large images on limited and heterogeneous screen sizes of mobile devices. In this paper, we propose a novel image viewing technique based on an adaptive attention shifting model. A presentation technique named rapid serial visual presentation (RSVP), borrowed from the UI community, is used to simulate the attention shifting process. We show a prototype image viewer developed for pocket PC and conduct some evaluations to demonstrate the effectiveness of our approach. Xin Fan 0001, Xing Xie 0001, Wei-Ying Ma, HongJiang Zhang, He-Qin Zhou |
ICME | 3 |
| 2003 | Imagerank: spectral techniques for structural analysis of image databaseabstractDrawing on the correspondence between spectral clustering, spectral dimensionality reduction, and the connections to the Markov chain theory, we present a novel unified framework for structural analysis of image database using spectral techniques. The framework provides a computationally efficient approach to both clustering and dimensionality reduction, or 2-D visualization. Within this framework, we can also infer the semantic degrees of the images, i.e. imagerank, which characterize the richness of semantics contained in the images. Some illustrative examples are discussed. Xiaofei He 0001, Wei-Ying Ma, HongJiang Zhang |
ICME | 2 |
| 2003 | An Evaluation on Feature Selection for Text Clustering
Shengping Liu, Zheng Chen 0001, Wei-Ying Ma |
ICML | 4 |
| 2003 | Looking into video frames on small displaysabstractWith the growing popularity of personal digital assistants and smart phones, people have become enthusiastic to watch videos through these mobile devices. However, a crucial challenge is to provide a better user experience for browsing videos on the limited and heterogeneous screen sizes. In this paper, we present a novel approach which allows users to overcome the display constraints by zooming into video frames while browsing. An automatic approach for detecting the focus regions is introduced to minimize the amount of user interaction. In order to improve the quality of output stream, virtual camera control is employed in the system. Preliminary evaluation shows that this approach is an effective way for video browsing on small displays. Xin Fan 0001, Xing Xie 0001, He-Qin Zhou, Wei-Ying Ma |
ACM Multimedia | 4 |
| 2003 | Automatic browsing of large pictures on mobile devicesabstractPictures have become increasingly common and popular in mobile communications. However, due to the limitation of mobile devices, there is a need to develop new technologies to facilitate the browsing of large pictures on the small screen. In this paper, we propose a novel approach which is able to automate the scrolling and navigation of a large picture with a minimal amount of user interaction on mobile devices. An image attention model is employed to illustrate the information structure within an image. An optimal image browsing path is then calculated based on the image attention model to simulate the human browsing behaviors. Experimental evaluations of the proposed mechanism indicate that our approach is an effective way for viewing large images on small displays. Hao Liu 0007, Xing Xie 0001, Wei-Ying Ma, HongJiang Zhang |
ACM Multimedia | 3 |
| 2003 | MobiPicture: browsing pictures on mobile devicesabstractPictures have become increasingly common and popular in mobile communication. However, due to the limitation of mobile devices, there is a need to develop new technologies to facilitate the browsing of pictures on the small screen. MobiPicture is a prototype system which includes a set of novel features to aid or automate a set of common image browsing tasks such as the thumbnail view, set-as-background, zooming and scrolling. Xing Xie 0001, Wei-Ying Ma, HongJiang Zhang |
ACM Multimedia | 3 |
| 2003 | Video summarization based on user log enhanced link analysisabstractEfficient video data management calls for intelligent video summarization tools that automatically generate concise video summaries for fast skimming and browsing. Traditional video summarization techniques are based on low-level feature analysis, which generally fails to capture the semantics of video content. Our vision is that users unintentionally embed their understanding of the video content in their interaction with computers. This valuable knowledge, which is difficult for computers to learn autonomously, can be utilized for video summarization process. In this paper, we present an intelligent video browsing and summarization system that utilizes previous viewers' browsing log to facilitate future viewers. Specifically, a novel ShotRank notion is proposed as a measure of the subjective interestingness and importance of each video shot. A ShotRank computation framework is constructed to seamlessly unify low-level video analysis and user browsing log mining. The resulting ShotRank is used to organize the presentation of video shots and generate video skims. Experimental results from user studies have strongly confirmed that ShotRank indeed represents the subjective notion of interestingness and importance of each video shot, and it significantly improves future viewers' browsing experience. Bin Yu 0010, Wei-Ying Ma, Klara Nahrstedt, HongJiang Zhang |
ACM Multimedia | 2 |
| 2003 | Knowing a tree from the forest: art image retrieval using a society of profilesabstractThis paper aims to address the problem of art image retrieval (AIR), which aims to help users find their favorite painting images. AIR is of great interests to us because of its application potentials and interesting research challenges---the retrieval is not only based on painting contents or styles, but also heavily based on user preference profiles. This paper describes the collaborative ensemble learning, a novel statistical learning approach to this task. It at first applies probabilistic support vector machines (SVMs) to model each individual user's profile based on given examples, i.e. liked or disliked paintings. Due to the high complexity of profile modelling, the SVMs can be rather weak in predicting preferences for new paintings. To overcome this problem, we combine a society of users' profiles, represented by their respective SVM models, to predict a given user's preferences for painting images. We demonstrate that the combination scheme is embedded in a Bayesian framework and retains intuitive interpretations---like-minded users are likely to share similar preferences. We report extensive empirical studies based on two experimental settings. The first one includes some controlled simulations performed on 4533 painting images. In the second setting, we report evaluations based on user preferences collected through an online web-based survey. Both experiments demonstrate that the proposed approach achieves excellent performance in terms of capturing a user's diverse preferences. Kai Yu 0001, Wei-Ying Ma, Volker Tresp, Zhao Xu 0001, Xiaofei He 0001, HongJiang Zhang, Hans-Peter Kriegel |
ACM Multimedia | 2 |
| 2003 | Image Adaptation Based on Attention Model for Small-Form-Factor Device
Li-Qun Chen, Xing Xie 0001, Wei-Ying Ma, HongJiang Zhang, He-Qin Zhou |
MMM | 3 |
| 2003 | Building a web thesaurus from web link structureabstractThesaurus has been widely used in many applications, including information retrieval, natural language processing, and question answering. In this paper, we propose a novel approach to automatically constructing a domain-specific thesaurus from the Web using link structure information. The proposed approach is able to identify new terms and reflect the latest relationship between terms as the Web evolves. First, a set of high quality and representative websites of a specific domain is selected. After filtering out navigational links, link analysis is applied to each website to obtain its content structure. Finally, the thesaurus is constructed by merging the content structures of the selected websites. The experimental results on automatic query expansion based on our constructed thesaurus show 20% improvement in search precision compared to the baseline. Zheng Chen 0001, Shengping Liu, Wenyin Liu, Geguang Pu, Wei-Ying Ma |
SIGIR | 5 |
| 2003 | ReCoM: reinforcement clustering of multi-type interrelated data objectsabstractMost existing clustering algorithms cluster highly related data objects such as Web pages and Web users separately. The interrelation among different types of data objects is either not considered, or represented by a static feature space and treated in the same ways as other attributes of the objects. In this paper, we propose a novel clustering approach for clustering multi-type interrelated data objects, ReCoM (Reinforcement Clustering of Multi-type Interrelated data objects). Under this approach, relationships among data objects are used to improve the cluster quality of interrelated data objects through an iterative reinforcement clustering process. At the same time, the link structure derived from relationships of the interrelated data objects is used to differentiate the importance of objects and the learned importance is also used in the clustering process to further improve the clustering results. Experimental results show that the proposed approach not only effectively overcomes the problem of data sparseness caused by the high dimensional relationship space but also significantly improves the clustering accuracy. Hua-Jun Zeng, Zheng Chen 0001, Hongjun Lu, Wei-Ying Ma |
SIGIR | 6 |
| 2003 | Implicit link analysis for small web searchabstractCurrent Web search engines generally impose link analysis-based re-ranking on web-page retrieval. However, the same techniques, when applied directly to small web search such as intranet and site search, cannot achieve the same performance because their link structures are different from the global Web. In this paper, we propose an approach to constructing implicit links by mining users' access patterns, and then apply a modified PageRank algorithm to re-rank web-pages for small web search. Our experimental results indicate that the proposed method outperforms content-based method by 16%, explicit link-based PageRank by 20% and DirectHit by 14%, respectively. Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma, HongJiang Zhang, Chao-Jun Lu |
SIGIR | 4 |
| 2003 | Collaborative Ensemble Learning: Combining Collaborative and Content-Based Information Filtering via Hierarchical Bayes
Kai Yu 0001, Anton Schwaighofer, Volker Tresp, Wei-Ying Ma, HongJiang Zhang |
UAI | 4 |
| 2003 | Detecting web page structure for adaptive viewing on small form factor devicesabstractMobile devices have already been widely used to access the Web. However, because most available web pages are designed for desktop PC in mind, it is inconvenient to browse these large web pages on a mobile device with a small screen. In this paper, we propose a new browsing convention to facilitate navigation and reading on a small-form-factor device. A web page is organized into a two level hierarchy with a thumbnail representation at the top level for providing a global view and index to a set of sub-pages at the bottom level for detail information. A page adaptation technique is also developed to analyze the structure of an existing web page and split it into small and logically related units that fit into the screen of a mobile device. For a web page not suitable for splitting, auto-positioning or scrolling-by-block is used to assist the browsing as an alterative. Our experimental results show that our proposed browsing convention and developed page adaptation scheme greatly improve the user's browsing experiences on a device with a small display. Wei-Ying Ma, HongJiang Zhang |
WWW | 2 |
| 2003 | Improving pseudo-relevance feedback in web information retrieval using web page segmentationabstractIn contrast to traditional document retrieval, a web page as a whole is not a good information unit to search because it often contains multiple topics and a lot of irrelevant information from navigation, decoration, and interaction part of the page. In this paper, we propose a VIsion-based Page Segmentation (VIPS) algorithm to detect the semantic content structure in a web page. Compared with simple DOM based segmentation method, our page segmentation scheme utilizes useful visual cues to obtain a better partition of a page at the semantic level. By using our VIPS algorithm to assist the selection of query expansion terms in pseudo-relevance feedback in web information retrieval, we achieve 27% performance improvement on Web Track dataset. Shipeng Yu, Deng Cai 0001, Ji-Rong Wen, Wei-Ying Ma |
WWW | 4 |
| 2003 | A visual attention model for adapting images on small displays
Li-Qun Chen, Xing Xie 0001, Xin Fan 0001, Wei-Ying Ma, HongJiang Zhang, He-Qin Zhou |
Multim. Syst. | 4 |
| 2003 | Ubiquitous media agents: a framework for managing personally accumulated multimedia files
Wenyin Liu, Zheng Chen 0001, Fan Lin, HongJiang Zhang, Wei-Ying Ma |
Multim. Syst. | 5 |
| 2003 | Alternating Feature Spaces in Relevance Feedback
Fang Qian, Mingjing Li, HongJiang Zhang, Wei-Ying Ma, Bo Zhang 0010 |
Multim. Tools Appl. | 4 |
| 2003 | Learning a semantic space from user's relevance feedback for image retrievalabstractAs current methods for content-based retrieval are incapable of capturing the semantics of images, we experiment with using spectral methods to infer a semantic space from user's relevance feedback, so that our system will gradually improve its retrieval performance through accumulated user interactions. In addition to the long-term learning process, we also model the traditional approaches to query refinement using relevance feedback as a short-term learning process. The proposed short- and long-term learning frameworks have been integrated into an image retrieval system. Experimental results on a large collection of images have shown the effectiveness and robustness of our proposed algorithms. Xiaofei He 0001, Oliver King, Wei-Ying Ma, Mingjing Li, HongJiang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2003 | Query Expansion by Mining User LogsabstractQueries to search engines on the Web are usually short. They do not provide sufficient information for an effective selection of relevant documents. Previous research has proposed the utilization of query expansion to deal with this problem. However, expansion terms are usually determined on term co-occurrences within documents. In this study, we propose a new method for query expansion based on user interactions recorded in user logs. The central idea is to extract correlations between query terms and document terms by analyzing user logs. These correlations are then used to select high-quality expansion terms for new queries. Compared to previous query expansion methods, ours takes advantage of the user judgments implied in user logs. The experimental results show that the log-based query expansion method can produce much better results than both the classical search method and the other query expansion methods. Hang Cui 0002, Ji-Rong Wen, Jian-Yun Nie, Wei-Ying Ma |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2002 | Learning and inferring a semantic space from user's relevance feedback for image retrievalabstractAs current methods for content-based retrieval are incapable of capturing the semantics of images, we experiment with using spectral methods to infer a semantic space from user's relevance feedback, so the system will gradually improve its retrieval performance through accumulated user interactions. In addition to the long-term learning process, we also model the traditional approaches to query refinement using relevance feedback as a short-term learning process. The proposed short- and long-term learning frameworks have been integrated into an image retrieval system. Experimental results on a large collection of images have shown the effectiveness and robustness of our proposed algorithms. Xiaofei He 0001, Wei-Ying Ma, Oliver King, Mingjing Li, HongJiang Zhang |
ACM Multimedia | 2 |
| 2002 | Enabling personalization services on the edgeabstractIn this paper, we describe our work on enabling personalization services on the edge of the Internet. In contrast to traditional approaches, this method offers many advantages. First, content providers can make their content people-aware while the content is delivered to the user. Second, since the personalization work is distributed to an overlay network of edge servers, it greatly improves the scalability and availability of the services. Third, the sufficiency of user data collected in edge servers enables us to build more accurate user models which further enhance the performance of personalization services. We describe the implementation of a prototype system named Avatar and present some experimental result in this paper. Xing Xie 0001, Hua-Jun Zeng, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2002 | A Unified Framework for Web Link AnalysisabstractWeb link analysis has been proved to significantly enhance the precision of Web searching in practice. Among existing approaches, Kleinberg's (1998) HITS and Google's PageRank are the two most representative algorithms that employ explicit hyperlink structure among Web pages to conduct link analysis, and DirectHit represents the other extreme that takes the user's access frequency as an implicit link to the Web page for assessing its importance. We propose a novel link analysis algorithm which puts both explicit and implicit link structures under a unified framework, and show that HITS and DirectHit are essentially two extreme instances of our proposed method. One important advantage of our method is its ability to analyze not only the hyperlinks between Web pages but also the interactions between users and the Web at the same time. The importance of Web pages and users can reinforce each other to improve Web link analysis. Compared with traditional HITS and DirectHit algorithms, our method further improves the search precision by 11.8% and 25.3%. Zheng Chen 0001, Wenyin Liu, Wei-Ying Ma |
WISE | 5 |
| 2002 | A Unified Framework for Clustering Heterogeneous Web ObjectsabstractWe introduce a novel framework for clustering Web data which is often heterogeneous in nature. As most existing methods often integrate heterogeneous data into a unified feature space, their flexibilities to explore and adjust contributing effects from different heterogeneous information are compromised. In contrast, our framework enables separate clustering of homogeneous data in the entire process based on their respective features, and a layered structure with link information is used to iteratively project and propagate the clustered results between layers until it converges. Our experimental results show that such a scheme not only effectively overcomes the problem of data sparseness caused by the high dimensional link space but also improves the clustering accuracy significantly. We achieve 19% and 41% performance increases when clustering Web-pages and users based on a semi-synthetic Web log. Finally, we show a real clustering result based on UC Berkeley's Web log. Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma |
WISE | 3 |
| 2002 | Probabilistic query expansion using query logsabstractQuery expansion has long been suggested as an effective way to resolve the short query and word mismatching problems. A number of query expansion methods have been proposed in traditional information retrieval. However, these previous methods do not take into account the specific characteristics of web searching; in particular, of the availability of large amount of user interaction information recorded in the web query logs. In this study, we propose a new method for query expansion based on query logs. The central idea is to extract probabilistic correlations between query terms and document terms by analyzing query logs. These correlations are then used to select high-quality expansion terms for new queries. The experimental results show that our log-based probabilistic query expansion method can greatly improve the search performance and has several advantages over other existing methods. Hang Cui 0002, Ji-Rong Wen, Jian-Yun Nie, Wei-Ying Ma |
WWW | 4 |
| 2002 | An interactive video delivery and caching system using video summarization
Sung-Ju Lee 0001, Wei-Ying Ma, Bo Shen 0003 |
Comput. Commun. | 2 |
| 2002 | Learning similarity measure for natural image retrieval with relevance feedbackabstractA new scheme of learning similarity measure is proposed for content-based image retrieval (CBIR). It learns a boundary that separates the images in the database into two clusters. Images inside the boundary are ranked by their Euclidean distances to the query. The scheme is called constrained similarity measure (CSM), which not only takes into consideration the perceptual similarity between images, but also significantly improves the retrieval performance of the Euclidean distance measure. Two techniques, support vector machine (SVM) and AdaBoost from machine learning, are utilized to learn the boundary. They are compared to see their differences in boundary learning. The positive and negative examples used to learn the boundary are provided by the user with relevance feedback. The CSM metric is evaluated in a large database of 10009 natural images with an accurate ground truth. Experimental results demonstrate the usefulness and effectiveness of the proposed similarity measure for image retrieval. Guodong Guo, Anil K. Jain 0001, Wei-Ying Ma, HongJiang Zhang |
IEEE Trans. Neural Networks | 3 |
| 2002 | User Intention Modelling in Web Applications Using Data Mining
Zheng Chen 0001, Fan Lin, Huan Liu 0001, Wei-Ying Ma, Wenyin Liu |
World Wide Web | 4 |
| 2001 | Learning Similarity Measure for Natural Image Retrieval with Relevance FeedbackabstractA new scheme of learning similarity measure is proposed for content-based image retrieval (CBIR). It learns a boundary that separates the images in the database into two parts. Images on the positive side of the boundary are ranked by their Euclidean distances to the query. The scheme is called restricted similarity measure (RSM), which not only takes into consideration the perceptual similarity between images, but also significantly improves the retrieval performance based on the Euclidean distance measure. Two techniques, support vector machine and AdaBoost, are utilized to learn the boundary, and compared with respect to their performance in boundary learning. The positive and negative examples used to learn the boundary are provided by the user with relevance feedback. The RSM metric is evaluated on a large database of 10,009 natural images with an accurate ground truth. Experimental results demonstrate the usefulness and effectiveness of the proposed similarity measure for image retrieval. Guodong Guo, Anil K. Jain 0001, Wei-Ying Ma, HongJiang Zhang |
CVPR (1) | 3 |
| 2000 | EdgeFlow: a technique for boundary detection and image segmentationabstractA novel boundary detection scheme based on "edge flow" is proposed in this paper. This scheme utilizes a predictive coding model to identify the direction of change in color and texture at each image location at a given scale, and constructs an edge flow vector. By propagating the edge flow vectors, the boundaries can be detected at image locations which encounter two opposite directions of flow in the stable state. A user defined image scale is the only significant control parameter that is needed by the algorithm. The scheme facilitates integration of color and texture into a single framework for boundary detection. Segmentation results on a large and diverse collections of natural images are provided, demonstrating the usefulness of this method to content based image retrieval. Wei-Ying Ma, B. S. Manjunath |
IEEE Trans. Image Process. | 1 |
| 1999 | Blur Determination in the Compressed Domain Using DTC InformationabstractThe paper presents a simple yet robust measure of image quality in terms of global (camera) blur. It is based on histogram computation of non-zero DCT coefficients. The technique is directly applicable to images and video frames in compressed (MPEG or JPEG) domain and to all types of MPEG frames (I-, P- or B-frames). The resulting quality measure is proved to be in concordance with subjective testing and is therefore suitable for quick qualitative characterization of images and video frames. Xavier Marichal, Wei-Ying Ma, HongJiang Zhang |
ICIP (2) | 2 |
| 1999 | NeTra: A Toolbox for Navigating Large Image Databases
Wei-Ying Ma, B. S. Manjunath |
Multim. Syst. | 1 |
| 1998 | A Texture Thesaurus for Browsing Large Aerial PhotographsabstractA texture-based image retrieval system for browsing large-scale aerial photographs is presented. The salient components of this system include texture feature extraction, image segmentation and grouping, learning similarity measure, and a texture thesaurus model for fast search and indexing. The texture features are computed by filtering the image with a bank of Gabor filters. This is followed by a texture gradient computation to segment each large airphoto into homogeneous regions. A hybrid neural network algorithm is used to learn the visual similarity by clustering patterns in the feature space. With learning similarity, the retrieval performance improves significantly. Finally, a texture image thesaurus is created by combining the learning similarity algorithm with a hierarchical vector quantization scheme. This thesaurus facilitates the indexing process while maintaining a good retrieval performance. Experimental results demonstrate the robustness of the overall system in searching over a large collection of airphotos and in selecting a diverse collection of geographic features such as housing developments, parking lots, highways, and airports. © 1998 John Wiley & Sons, Inc. Wei-Ying Ma, B. S. Manjunath |
J. Am. Soc. Inf. Sci. | 1 |
| 1997 | Edge Flow: A Framework of Boundary Detection and Image SegmentationabstractA novel boundary detection scheme based on "edge flow" is proposed in this paper. This scheme utilizes a predictive coding model to identify the direction of change in color and texture at each image location at a given scale, and constructs an edge flow vector. By iteratively propagating the edge flow, the boundaries can be detected at image locations which encounter two opposite directions of flow in the stable state. A user defined image scale is the only significant control parameter that is needed by the algorithm. The scheme facilitates integration of color and texture into a single framework for boundary detection. Wei-Ying Ma, B. S. Manjunath |
CVPR | 1 |
| 1997 | NeTra: A Toolbox for Navigating Large Image DatabasesabstractWe present an implementation of NeTra, a prototype image retrieval system that uses color texture, shape and spatial location information in segmented image database. A distinguishing aspect of this system is its incorporation of a robust automated image segmentation algorithm that allows object or region based search. Image segmentation significantly improves the quality of image retrieval when images contain multiple complex objects. Other important components of the system include an efficient color representation, and indexing of color, texture, and shape features for fast search and retrieval. This representation allows the user to compose interesting queries such as "retrieve all images that contain regions that have the color of object A, texture of object B, shape of object C, and lie in the upper one-third of the image" where the individual objects could be regions belonging to different images. Wei-Ying Ma, B. S. Manjunath |
ICIP (1) | 1 |
| 1996 | Texture Features and Learning SimilarityabstractThis paper addresses two important issues related to texture pattern retrieval: feature extraction and similarity search. A Gabor feature representation for textured images is proposed, and its performance in pattern retrieval is evaluated on a large texture image database. These features compare favorably with other existing texture representations. A simple hybrid neural network algorithm is used to learn the similarity by simple clustering in the texture feature space. With learning similarity the performance of similar pattern retrieval improves significantly. An important aspect of this work is its application to real image data. Texture feature extraction with similarity learning is used to search through large aerial photographs. Feature clustering enables efficient search of the database as our experimental results indicate. Wei-Ying Ma, B. S. Manjunath |
CVPR | 1 |
| 1996 | Browsing large satellite and aerial photographsabstractImage content based retrieval in the Alexandria digital library project has focussed on texture and color features for querying the database. A robust texture feature extraction algorithm and a fast segmentation scheme have been developed. The texture features are computed by filtering the image with a bank of Gabor filters. This is followed by a clustering scheme to create a texture based feature dictionary, which is then used to search and retrieve similar looking patterns from other images. Experimental results demonstrate the robustness of the overall system in searching over a large collection of airphotos and in selecting a surprisingly diverse collection of geographic features such as housing developments, parking lots, highways, and airports. B. S. Manjunath, Wei-Ying Ma |
ICIP (2) | 2 |
| 1996 | Texture-Based Pattern Retrieval from Image Databases
Wei-Ying Ma, B. S. Manjunath |
Multim. Tools Appl. | 1 |
| 1996 | Texture Features for Browsing and Retrieval of Image DataabstractImage content based retrieval is emerging as an important research area with application to digital libraries and multimedia databases. The focus of this paper is on the image processing aspects and in particular using texture information for browsing and retrieval of large image data. We propose the use of Gabor wavelet features for texture analysis and provide a comprehensive experimental evaluation. Comparisons with other multiresolution texture features using the Brodatz texture database indicate that the Gabor features provide the best pattern retrieval accuracy. An application to browsing large air photos is illustrated. B. S. Manjunath, Wei-Ying Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1995 | A comparison of wavelet transform features for texture image annotationabstractA comparison of different wavelet transform based texture features for content based search and retrieval is made. These include the conventional orthogonal and bi-orthogonal wavelet transforms, tree-structured decompositions, and the Gabor wavelet transforms. Issues discussed include image processing complexity, texture classification and discrimination, and suitability for developing indexing techniques. Wei-Ying Ma, B. S. Manjunath |
ICIP | 1 |