VLDB 2026 Research / reviewers in the wild / expert
Ruihua Song
dblp:s/RuihuaSong
· DBLP profile ↗
94ranked-venue papers
9as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 43 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 40 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 31 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative WritingabstractQian Cao, Yahui Liu, Wei Bi, Yi Zhao, Ruihua Song, Xiting Wang, Ruiming Tang, Guorui Zhou, Han Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qian Cao 0001, Wei Bi, Ruihua Song, Xiting Wang, Ruiming Tang, Guorui Zhou, Han Li 0005 |
ACL (1) | 5 |
| 2025 | EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed DynamicsabstractImitating how humans move their gaze in a visual scene is a vital research problem for both visual understanding and psychology, kindling crucial applications such as building alive virtual characters. Previous studies aim to predict gaze trajectories when humans are free-viewing an image, searching for required targets, or looking for clues to answer questions in an image. While these tasks focus on visual-centric scenarios, humans move their gaze also along with audio signal inputs in more common scenarios. To fill this gap, we introduce a new task that predicts human gaze trajectories in a visual scene with synchronized audio inputs and provide a new dataset containing 20k gaze points from 8 subjects. To effectively integrate audio information and simulate the dynamic process of human gaze motion, we propose a novel learning framework called EyEar (Eye moving while Ear listening) based on physics-informed dynamics, which considers three key factors to predict gazes: eye inherent motion tendency, vision salient attraction, and audio semantic attraction. We also propose a probability density score to overcome the high individual variability of gaze trajectories, thereby improving the stabilization of optimization and the reliability of the evaluation. Experimental results show that EyEar outperforms all the baselines in the context of all evaluation metrics, thanks to the proposed components in the learning model. Xin Cheng 0008, Yuchong Sun, Ruihua Song, Hao Sun 0002, Denghao Zhang |
AAAI | 5 |
| 2025 | Enhancing Audiovisual Speech Recognition Through Bifocal Preference OptimizationabstractAudiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic environments, spontaneous speech, and the uncertain use of visual information. Most previous works fine-tune audio-only ASR models on audiovisual datasets, optimizing them for conventional ASR objectives. However, they often neglect visual features and common errors in unconstrained video scenarios. In this paper, we propose using a preference optimization strategy to improve speech recognition accuracy for real-world videos. First, we create preference data via simulating common errors that occurred in AV-ASR from two focals: manipulating the audio or vision input and rewriting the output transcript. Second, we propose BPO-AVASR, a Bifocal Preference Optimization method to improve AV-ASR models by leveraging both input-side and output-side preference. Extensive experiments demonstrate that our approach significantly improves speech recognition accuracy across various domains, outperforming previous state-of-the-art models on real-world video speech recognition. Yihan Wu 0008, Yifan Peng 0003, Xihua Wang 0002, Ruihua Song, Shinji Watanabe 0001 |
AAAI | 5 |
| 2025 | Towards Effective and Efficient Continual Pre-training of Large Language ModelsabstractContinual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly improve the Chinese language ability and scientific reasoning ability of LLMs. To achieve it, we design specific data mixture and curriculum strategies based on existing datasets and synthetic high-quality data. Concretely, we synthesize multidisciplinary scientific QA pairs based on related web pages to guarantee the data quality, and also devise the performance tracking and data mixture adjustment strategy to ensure the training stability. For the detailed designs, we conduct preliminary studies on a relatively small model, and summarize the findings to help optimize our CPT method. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of Llama-3 (8B), including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval). Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE. Jie Chen 0007, Zhipeng Chen 0001, Kun Zhou 0002, Yutao Zhu 0001, Jinhao Jiang, Yingqian Min, Wayne Xin Zhao, Zhicheng Dou, Jiaxin Mao, Yankai Lin 0001, Ruihua Song, Jun Xu 0001, Xu Chen 0017, Rui Yan 0001, Zhewei Wei, Di Hu 0001, Wenbing Huang 0001, Ji-Rong Wen |
ACL (1) | 12 |
| 2025 | What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction TuningabstractVisual instruction tuning is crucial for enhancing the zero-shot generalization capability of Multi-modal Large Language Models (MLLMs). In this paper, we aim to investigate a fundamental question: “what makes for good visual instructions”. Through a comprehensive empirical study, we find that instructions focusing on complex visual reasoning tasks are particularly effective in improving the performance of MLLMs, with results correlating to instruction complexity. Based on this insight, we develop a systematic approach to automatically create high-quality complex visual reasoning instructions. Our approach employs a synthesize-complicate-reformulate paradigm, leveraging multiple stages to gradually increase the complexity of the instructions while guaranteeing quality. Based on this approach, we create the ComVint dataset with 32K examples, and fine-tune four MLLMs on it. Experimental results consistently demonstrate the enhanced performance of all compared MLLMs, such as a 27.86% and 27.60% improvement for LLaVA on MME-Perception and MME-Cognition, respectively. Our code and data are publicly available at the link: https://github.com/RUCAIBox/ComVint. Yifan Du 0002, Hangyu Guo, Kun Zhou 0002, Wayne Xin Zhao, Jinpeng Wang 0001, Chuyuan Wang, Mingchen Cai, Ruihua Song, Ji-Rong Wen |
COLING | 8 |
| 2025 | MuKA: Multimodal Knowledge Augmented Visual Information-SeekingabstractThe visual information-seeking task aims to answer visual questions that require external knowledge, such as “On what date did this building officially open?”. Existing methods using retrieval-augmented generation framework primarily rely on textual knowledge bases to assist multimodal large language models (MLLMs) in answering questions. However, the text-only knowledge can impair information retrieval for the multimodal query of image and question, and also confuse MLLMs in selecting the most relevant information during generation. In this work, we propose a novel framework MuKA which leverages a multimodal knowledge base to address these limitations. Specifically, we construct a multimodal knowledge base by automatically pairing images with text passages in existing datasets. We then design a fine-grained multimodal interaction to effectively retrieve multimodal documents and enrich MLLMs with both retrieved texts and images. MuKA outperforms state-of-the-art methods by 38.7% and 15.9% on the InfoSeek and E-VQA benchmark respectively, demonstrating the importance of multimodal knowledge in enhancing both retrieval and answer generation. Lianghao Deng, Yuchong Sun, Shizhe Chen, Ruihua Song |
COLING | 6 |
| 2025 | Animate and Sound an ImageabstractThis paper addresses a promising yet underexplored task, Image-to-Sounding-Video (I2SV) generation, which animates a static image and generates synchronized sound simultaneously. Despite advances in video and audio generation models, challenges remain to develop a unified model for generating naturally sounding videos. In this work, we propose a novel approach that leverages two separate pretrained diffusion models and makes vision and audio influence each other during generation based on the Diffusion Transformer (DiT) architecture. First, the individual video and audio pretrained generation models are decomposed into input, output, and expert sub-modules. We propose using a unified joint DiT block to integrate the expert sub-modules to effectively model the interaction between the two modalities, resulting in high-quality I2SV generation. Then, we introduce a joint classifier-free guidance technique to boost the performance during joint generation. Finally, we conduct extensive experiments on three popular benchmark datasets, and in both objective and subjective evaluation our method surpass all the baseline methods in almost all metrics. Case studies show our generated sounding videos are high quality and synchronized between video and audio. Xihua Wang 0002, Ruihua Song, Chongxuan Li, Xin Cheng 0008, Yihan Wu 0008, Yuyue Wang 0003, Hongteng Xu |
CVPR | 2 |
| 2025 | LoVA: Long-form Video-to-Audio GenerationabstractVideo-to-audio (V2A) generation is important for video editing and post-processing, enabling the creation of semantics-aligned audio for silent video. However, most existing methods focus on generating short-form audio for short video segment (less than 10 seconds), while giving little attention to the scenario of long-form video inputs. For current UNet-based diffusion V2A models, an inevitable problem when handling long-form audio generation is the inconsistencies within the final concatenated audio. In this paper, we first highlight the importance of long-form V2A problem. Besides, we propose LoVA, a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models. Extensive objective and subjective experiments demonstrate that LoVA achieves comparable performance on 10second V2A benchmark and outperforms all other baselines on a benchmark with long-form video input. Xin Cheng 0008, Xihua Wang 0002, Yihan Wu 0008, Yuyue Wang 0003, Ruihua Song |
ICASSP | 5 |
| 2025 | Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion TransformerabstractMulti-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual motions and generating two motions with huge differences from one text condition. Current research often employs separate module branches for individual motions, leading to a loss of interaction information and increased computational demands. To address these challenges, we propose a novel, unified approach that models multi-person motions and their interactions within a single latent space. Our approach streamlines the process by treating interactive motions as an integrated data point, utilizing a Variational AutoEncoder (VAE) for compression into a unified latent space, and performing a diffusion process within this space, guided by the natural language conditions. Experimental results demonstrate our method’s superiority over existing approaches in generation quality, performing text condition in particular when motions have significant asymmetry, and accelerating the generation efficiency while preserving high quality. Xihua Wang 0002, Ruihua Song, Wenbing Huang 0001 |
ICASSP | 3 |
| 2025 | ETVA: Evaluation of Text-to-Video Alignment via Fine-Grained Question Generation and AnsweringabstractPrecisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human preference. To address this limitation, we propose ETVA, a novel Evaluation method of Text-to-Video Alignment via fine-grained question generation and answering. First, a multi-agent system parses prompts into semantic scene graphs to generate atomic questions. Then we design a knowledge-augmented multi-stage reasoning framework for question answering, where an auxiliary LLM first retrieves relevant common-sense knowledge (e.g., physical laws), and then video LLM answers the generated questions through a multi-stage reasoning mechanism. Extensive experiments demonstrate that ETVA achieves a Spearman's correlation coefficient of 58.47, showing a much higher correlation with human judgment than existing metrics which attain only 31.0. We also construct a comprehensive benchmark specifically designed for text-to-video alignment evaluation, featuring 2k diverse prompts and 12k atomic questions spanning 10 categories. Through a systematic evaluation of 15 existing text-to-video models, we identify their key capabilities and limitations, paving the way for next-generation T2V generation. Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Kieran Liu, Ruihua Song |
ICCV | 8 |
| 2025 | VAFlow: Video-to-Audio Generation with Cross-Modality Flow Matching
Xihua Wang 0002, Xin Cheng 0008, Yuyue Wang 0003, Ruihua Song |
ICCV | 4 |
| 2025 | Think Then React: Towards Unconstrained Action-to-Reaction Motion GenerationabstractModeling human-like action-to-reaction generation has significant real-world applications, like human-robot interaction and games.
Despite recent advancements in single-person motion generation, it is still challenging to well handle action-to-reaction generation, due to the difficulty of directly predicting reaction from action sequence without prompts, and the absence of a unified representation that effectively encodes multi-person motion.
To address these challenges, we introduce Think-Then-React (TTR), a large language-model-based framework designed to generate human-like reactions.
First, with our fine-grained multimodal training strategy, TTR is capable to unify two processes during inference: a thinking process that explicitly infers action intentions and reasons corresponding reaction description, which serve as semantic prompts, and a reacting process that predicts reactions based on input action and the inferred semantic prompts.
Second, to effectively represent multi-person motion in language models, we propose a unified motion tokenizer by decoupling egocentric pose and absolute space features, which effectively represents action and reaction motion with same encoding.
Extensive experiments demonstrate that TTR outperforms existing baselines, achieving significant improvements in evaluation metrics, such as reducing FID from 3.988 to 1.942. Wenhui Tan, Chuhao Jin, Wenbing Huang 0001, Xiting Wang, Ruihua Song |
ICLR | 6 |
| 2025 | Uncovering Personality Traits via Multimodal LLM for Personalized Image Emotion AnalysisabstractHuman emotion induced by images is strongly linked to individual personalities. Most existing works focus on analyzing the dominant emotions, i.e., the emotion that most viewers have for an image, leaving personalized image emotion analysis less explored. In this paper, we propose MLLM-PIEA, a framework based on Multimodal Large Language Models (MLLMs) for Personalized Image Emotion Analysis. To better represent the personalities of different viewers, we propose using an MLLM to uncover personality traits from the viewers’ experience data. These personality traits are summarized as structured descriptions and then used to augment another MLLM for emotion analysis. Experimental results show that our method brings a significant relative improvement of 28.8% over the baseline method. Jianzhang Gao, Hao Pu, Yuchong Sun, Ruihua Song |
ICME | 4 |
| 2025 | GALAXY: A Large-Scale Open-Domain Dataset for Multimodal Learning
Yihan Wu 0008, Jiaqi Song, Ruihua Song, Shinji Watanabe 0001 |
INTERSPEECH | 6 |
| 2025 | A Visual Speech Language Model for Visual Text-to-Speech TaskabstractThe task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being consistent with the content of input text and cloning the timbre of a reference speech. Existing VisualTTS models typically adopt lightweight architectures and design specialized modules to achieve the above goals respectively, yet the speech quality is not satisfied due to the model capacity and the limited data in VisualTTS. Recently, speech large language models (SpeechLLM) show the robust ability to generate high-quality speech. But few work has been done to well leverage temporal cues from video input in generating lip-synchronized speech. To generate both high-quality and lip-synchronized speech in VisualTTS tasks, we propose a novel Visual Speech Language Model called VSpeechLM based upon a SpeechLLM. To capture the synchronization relationship between text and video, we propose a text-video aligner. It first learns fine-grained alignment between phonemes and lip movements, and then outputs an expanded phoneme sequence containing lip-synchronization cues. Next, our proposed SpeechLLM based decoders take the expanded phoneme sequence as input and learns to generate lip-synchronized speech. Extensive experiments demonstrate that our VSpeechLM significantly outperforms previous VisualTTS methods in terms of overall quality, speaker similarity, and synchronization metrics. Yuyue Wang 0003, Xin Cheng 0008, Yihan Wu 0008, Xihua Wang 0002, Jinchuan Tian, Ruihua Song |
MMAsia | 6 |
| 2025 | Understanding the Roles of Visual Modality in Multimodal Dialogue: An Empirical Study
Qian Cao 0001, Ruihua Song, Xu Chen 0017 |
MMM (4) | 2 |
| 2025 | RoLD: Robot Latent Diffusion for Multi-task Policy Modeling
Wenhui Tan, Bei Liu 0001, Ruihua Song, Jianlong Fu |
MMM (3) | 4 |
| 2025 | Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning ChainsabstractLarge Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and inefficient.
In this paper, we introduce Compressed Latent Reasoning (CoLaR), a novel framework that dynamically compresses reasoning processes in latent space through a two-stage training approach.
First, during supervised fine-tuning, CoLaR extends beyond next-token prediction by incorporating an auxiliary next compressed embedding prediction objective. This process merges embeddings of consecutive tokens using a compression factor $c$ randomly sampled from a predefined range, and trains a specialized latent head to predict distributions of subsequent compressed embeddings. Second, we enhance CoLaR through reinforcement learning (RL) that leverages the latent head's non-deterministic nature to explore diverse reasoning paths and exploit more compact ones.
This approach enables CoLaR to: i) **perform reasoning at a dense latent level** (i.e., silently), substantially reducing reasoning chain length, and ii) **dynamically adjust reasoning speed** at inference time by simply prompting the desired compression factor.
Extensive experiments across four mathematical reasoning datasets demonstrate that CoLaR achieves 14.1% higher accuracy than latent-based baseline methods at comparable compression ratios, and reduces reasoning chain length by 53.3% with only 4.8% performance degradation compared to explicit CoT method. Moreover, when applied to more challenging mathematical reasoning tasks, our RL-enhanced CoLaR demonstrates performance gains of up to 5.4% while dramatically reducing latent reasoning chain length by 82.8%.
The code and models will be released upon acceptance. Wenhui Tan, Jianzhong Ju, Zhenbo Luo, Ruihua Song, Jian Luan 0001 |
NeurIPS | 5 |
| 2025 | ReGA: Reasoning and Grounding Decoupled GUI Navigation Agents
Feiyue Ni, Yanchu Guan, Yuchong Sun, Dong Wang 0062, Chenyi Zhuang, Jinjie Gu, Ruihua Song |
NLPCC (1) | 7 |
| 2025 | A Mixture-of-Experts Framework Based on Depth Images for Text to Video Storyboard Task
Xu Gu 0003, Feiyue Ni, Ruihua Song |
PRCV (11) | 3 |
| 2025 | Transferring Foundation Models for Generalizable Robotic ManipulationabstractImproving the generalization capabilities of general-purpose robotic manipulation in real world has long been a significant challenge. Existing approaches often rely on collecting large-scale robotic data which is costly and time-consuming. However, due to insufficient diversity of data, they typically suffer from limiting their capability in open-domain scenarios with new objects and diverse environments. In this paper, we propose a novel paradigm that effectively leverages language-reasoning segmentation mask generated by internet-scale foundation models, to condition robot manipulation tasks. By integrating the mask modality, which incorporates semantic, geometric, and temporal correlation priors derived from vision foundation models, into the end-to-end policy model, our approach can effectively and robustly perceive object pose and enable sample-efficient generalization learning, including new object instances, semantic categories, and unseen backgrounds. We first introduce a series of foundation models to ground natural language demands across multiple tasks. Secondly, we develop a two-stream 2D policy model based on imitation learning, which processes raw images and object masks to predict robot actions with a local-global perception manner. Extensive real-world experiments conducted on a Franka Emika robot and a low-cost dual-arm robot demonstrate the effectiveness of our proposed paradigm and policy. Demos can be found in link 1 or link 2 and our code will be released at https://github.com/MCG-NJU/TPM. Jiange Yang, Wenhui Tan, Chuhao Jin, Keling Yao, Bei Liu 0001, Jianlong Fu, Ruihua Song, Gangshan Wu, Limin Wang 0001 |
WACV | 7 |
| 2025 | User Behavior Simulation with Large Language Model-based AgentsabstractSimulating high quality user behavior data has always been a fundamental yet challenging problem in human-centered applications such as recommendation systems, social networks, among many others. The major difficulty of user behavior simulation originates from the intricate mechanism of human cognitive and decision processes. Recently, substantial evidence has suggested that by learning huge amounts of web knowledge, large language models (LLMs) can achieve human-like intelligence and generalization capabilities. Inspired by such capabilities, in this article, we take an initial step to study the potential of using LLMs for user behavior simulation in the recommendation domain. To make LLMs act like humans, we design profile, memory and action modules to equip them, building LLM-based agents to simulate real users. To enable interactions between different agents and observe their behavior patterns, we design a sandbox environment, where each agent can interact with the recommendation system, and different agents can converse with their friends via one-to-one chatting or one-to-many social broadcasting. In the experiments, we first demonstrate the believability of the agent-generated behaviors based on both subjective and objective evaluations. Then, to show the potential applications of our method, we simulate and study two social phenomena including (1) information cocoons and (2) user conformity behaviors. We find that controlling the personalization degree of recommendation algorithms and improving the heterogeneity of user social relations can be two effective strategies for alleviating the problem of information cocoon, and the conformity behaviors can be highly influenced by the amount of user social relations. To advance this direction, we have released our project at https://github.com/RUC-GSAI/YuLan-Rec . Lei Wang 0198, Jingsen Zhang, Hao Yang 0045, Jiakai Tang, Zeyu Zhang 0007, Xu Chen 0017, Yankai Lin 0001, Hao Sun 0002, Ruihua Song, Wayne Xin Zhao, Jun Xu 0001, Zhicheng Dou, Jun Wang 0012, Ji-Rong Wen |
ACM Trans. Inf. Syst. | 10 |
| 2024 | Persuading across Diverse Domains: a Dataset and Persuasion Large Language ModelabstractPersuasive dialogue requires multi-turn following and planning abilities to achieve the goal of persuading users, which is still challenging even for state-of-the-art large language models (LLMs).Previous works focus on retrievalbased models or generative models in a specific domain due to a lack of data across multiple domains.In this paper, we leverage GPT-4 to create the first multi-domain persuasive dialogue dataset DailyPersuasion.Then we propose a general method named PersuGPT to learn a persuasion model based on LLMs through intent-to-strategy reasoning, which summarizes the intent of user's utterance and reasons next strategy to respond.Moreover, we design a simulation-based preference optimization, which utilizes a learned user model and our model to simulate next turns and estimate their rewards more accurately.Experimental results on two datasets indicate that our proposed method outperforms all baselines in terms of automatic evaluation metric Win-Rate and human evaluation.The code and data are available at https://persugpt.github.io. Chuhao Jin, Kening Ren, Lingzhen Kong, Xiting Wang, Ruihua Song |
ACL (1) | 5 |
| 2024 | Parrot: Enhancing Multi-Turn Instruction Following for Large Language ModelsabstractYuchong Sun, Che Liu, Kun Zhou, Jinwen Huang, Ruihua Song, Xin Zhao, Fuzheng Zhang, Di Zhang, Kun Gai. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuchong Sun, Kun Zhou 0002, Jinwen Huang, Ruihua Song, Wayne Xin Zhao, Di Zhang 0026, Kun Gai |
ACL (1) | 5 |
| 2024 | Intelligent Agents with LLM-based Process AutomationabstractWhile intelligent virtual assistants like Siri, Alexa, and Google Assistant have become ubiquitous in modern life, they still face limitations in their ability to follow multi-step instructions and accomplish complex goals articulated in natural language. However, recent breakthroughs in large language models (LLMs) show promise for overcoming existing barriers by enhancing natural language processing and reasoning capabilities. Though promising, applying LLMs to create more advanced virtual assistants still faces challenges like ensuring robust performance and handling variability in real-world user commands. This paper proposes a novel LLM-based virtual assistant that can automatically perform multi-step operations within mobile apps based on high-level user requests. The system represents an advance in assistants by providing an end-to-end solution for parsing instructions, reasoning about goals, and executing actions. LLM-based Process Automation (LLMPA) has modules for decomposing instructions, generating descriptions, detecting interface elements, predicting next actions, and error checking. Experiments demonstrate the system completing complex mobile operation tasks in Alipay based on natural language instructions. This showcases how large language models can enable automated assistants to accomplish real-world tasks. The main contributions are the novel LLMPA architecture optimized for app process automation, the methodology for applying LLMs to mobile apps, and demonstrations of multi-step task completion in a real-world environment. Notably, this work represents the first real-world deployment and extensive evaluation of a large language model-based virtual assistant in a widely used mobile application with an enormous user base numbering in the hundreds of millions. Yanchu Guan, Dong Wang 0062, Zhixuan Chu, Shiyu Wang 0001, Feiyue Ni, Ruihua Song, Chenyi Zhuang |
KDD | 6 |
| 2024 | See or Guess: Counterfactually Regularized Image CaptioningabstractImage captioning, which generates natural language descriptions of images, is a crucial task in vision-language research. Previous models have typically addressed this task by aligning the generative capabilities of machines with humans through statistical fitting existing datasets. While effective for normal images, they may struggle to accurately describe those where certain parts of the image are obscured or edited, unlike humans who excel in such cases. These weaknesses, including hallucinations and limited interpretability, often hinder performance in scenarios with shifted association patterns. In this paper, we present a generic image captioning framework that employs causal inference to make existing models more capable of interventional tasks, and counterfactually explainable. Our approach includes two variants leveraging either total effect or natural direct effect. Integrating them into the training process enables models to handle counterfactual scenarios, increasing their generalizability. Extensive experiments on various datasets show that our method effectively reduces hallucinations and improves the model's faithfulness to images, demonstrating high portability across both small-scale and large-scale image-to-text models. The code is available at https://github.com/Aman-4-Real/See-or-Guess. Qian Cao 0001, Xu Chen 0017, Ruihua Song, Xiting Wang, Xinting Huang, Yuchen Ren 0005 |
ACM Multimedia | 3 |
| 2024 | TiVA: Time-Aligned Video-to-Audio Generation
Xihua Wang 0002, Yuyue Wang 0003, Yihan Wu 0008, Ruihua Song, Xu Tan 0003, Zehua Chen 0005, Hongteng Xu, Guodong Sui |
ACM Multimedia | 4 |
| 2024 | ScaMo: Towards Text to Video Storyboard Generation Using Scale and Movement of Shots
Xu Gu 0003, Xihua Wang 0002, Chuhao Jin, Ruihua Song |
MMAsia | 4 |
| 2024 | ViCo: Engaging Video Comment Generation with Human Preference Rewards
Yuchong Sun, Bei Liu 0001, Xu Chen 0017, Ruihua Song, Jianlong Fu |
MMAsia | 4 |
| 2024 | What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?abstractYan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong, Ruihua Song. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Tao Kong, Ruihua Song |
NAACL-HLT | 9 |
| 2024 | ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs For Audio, Music, and SpeechabstractNeural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with autoregressive language models. However, as extensive downstream applications are investigated, challenges have arisen in ensuring fair comparisons across diverse applications. To address these issues, we present a new open-source platform ESPnet-Codec, which is built on ESPnet and focuses on neural codec training and evaluation. ESPnet-Codec offers various recipes in audio, music, and speech for training and evaluation using several widely adopted codec models. Together with ESPnet-Codec, we present VERSA, a standalone evaluation toolkit, which provides a comprehensive evaluation of codec performance over 20 audio evaluation metrics. Notably, we demonstrate that ESPnet-Codec can be integrated into six ESPnet tasks, supporting diverse applications. Jiatong Shi, Jinchuan Tian, Yihan Wu 0008, Jee-Weon Jung, Jia Qi Yip, Yoshiki Masuyama, Yuning Wu 0001, Yuxun Tang, Massa Baali, Dareen Alharthi, Ruifan Deng, Tejes Srivastava, Alexander H. Liu, Bhiksha Raj, Qin Jin, Ruihua Song, Shinji Watanabe 0001 |
SLT | 19 |
| 2024 | Robust Audiovisual Speech Recognition Models with Mixture-of-ExpertsabstractVisual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities across diverse video scenarios, presenting a significant challenge. In this paper, we introduce EVA, leveraging the mixture-of-Experts for audioVisual ASR to perform robust speech recognition for “in-the-wild” videos. Specifically, we first encode visual information into visual tokens sequence and map them into speech space by a lightweight projection. Then, we build EVA upon a robust pretrained speech recognition model, ensuring its generalization ability. Moreover, to incorporate visual information effectively, we inject visual information into the ASR model through a mixture-of-experts module. Experiments show our model achieves state-of-the-art results on three benchmarks, which demonstrates the generalization ability of EVA across diverse video domains. Yihan Wu 0008, Yifan Peng 0003, Xuankai Chang, Ruihua Song, Shinji Watanabe 0001 |
SLT | 5 |
| 2024 | Understanding Human Preferences: Towards More Personalized Video to Text GenerationabstractWhile previous video to text models have achieved remarkable successes, they mostly focus on how to understand the video contents in a general sense, but fail to capture the human personalized preferences, which is highly demanded for an engaging multimodal chatbots. Different from user modeling in collaborative filtering, there is no other user behaviors in inference as a real-time video stream is coming. In this paper, we formally define the task of personalized video commenting task and design an end-to-end personalized framework for solving this task. In specific, we argue that the personalization for video comment generation can be reflected in two aspects, that is, (1) for the same video, different users may comment on different clips, and (2) for the same clip, different people may also express various opinions with diverse commentary styles. Motivated by these considerations, we design our framework based on two components. The first one is a clip selector, which is responsible for predicting the clips that the user may comment in the video. The second one is a text generator, which aims to produce the comment based on the above predicted clips and the user's preference. In our framework, these two components are optimized in an end-to-end manner to mutually enhance each other, where we design confidence-aware scheduled sampling and iterative inference strategies to solve the problem that the ground truth clips are absent in the inference phase. As the absence of personalized video to text dataset, we collect and release a new dataset for studying this problem. We conduct extensive experiments to demonstrate the effectiveness of our model. Yihan Wu 0008, Ruihua Song, Xu Chen 0017, Hao Jiang 0022, Zhao Cao |
WWW | 2 |
| 2024 | Show Me a Video: A Large-Scale Narrated Video Dataset for Coherent Story IllustrationabstractIllustrating a multi-sentence story with visual content is a significant challenge in multimedia research. While previous works have focused on sequential story-to-visual representations at the image level or representing a single sentence with a video clip, illustrating a long multi-sentence story with coherent videos remains an under-explored area. In this paper, we propose the task of video-based story illustration that focuses on the goal of visually illustrating a story with retrieved video clips. To support this task, we first create a large-scale dataset of coherent video stories in each sample, consisting of 85K narrative stories with 60 pairs of consistent clips and texts. We then propose the Story Context-Enhanced Model, which leverages local and global contextual information within the story, inspired by sequence modeling in language understanding. Through comprehensive quantitative experiments, we demonstrate the effectiveness of our baseline model. In addition, qualitative results and detailed user studies reveal that our method can retrieve coherent video sequences from stories. The dataset and code will be made publicly athttps://nfy-dot.github.io/CVSV-dataset/. Yu Lu 0019, Feiyue Ni, Linchao Zhu, Zongxin Yang, Ruihua Song, Lele Cheng, Yi Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingabstractVideo dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech synthesis. To ensure the translated speech to be well aligned with the corresponding video, the length/duration of the translated speech should be as close as possible to that of the original speech, which requires strict length control. Previous works usually control the number of words or characters generated by the machine translation model to be similar to the source sentence, without considering the isochronicity of speech as the speech duration of words/characters in different languages varies. In this paper, we propose VideoDubber, a machine translation system tailored for the task of video dubbing, which directly considers the speech duration of each token in translation, to match the length of source and target speech. Specifically, we control the speech length of generated sentence by guiding the prediction of each word with the duration information, including the speech duration of itself as well as how much duration is left for the remaining words. We design experiments on four language directions (German -> English, Spanish -> English, Chinese English), and the results show that VideoDubber achieves better length control ability on the generated speech than baseline methods. To make up the lack of real-world datasets, we also construct a real-world test set collected from films to provide comprehensive evaluations on the video dubbing task. Yihan Wu 0008, Junliang Guo, Xu Tan 0003, Chen Zhang 0020, Bohan Li 0003, Ruihua Song, Lei He 0005, Sheng Zhao 0002, Arul Menezes, Jiang Bian 0002 |
AAAI | 6 |
| 2023 | CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment
Hongwei Xue, Yuchong Sun, Bei Liu 0001, Jianlong Fu, Ruihua Song, Houqiang Li, Jiebo Luo 0001 |
ICLR | 5 |
| 2023 | ComedicSpeech: Text To Speech For Stand-up Comedies in Low-Resource ScenariosabstractText to Speech (TTS) models can generate natural and highquality speech, but it is not expressive enough when synthesizing speech with dramatic expressiveness, such as stand-up comedies.Considering comedians have diverse personal speech styles, including personal prosody, rhythm, and fillers, it requires real-world datasets and strong speech style modeling capabilities, which brings challenges.In this paper, we construct a new dataset and develop ComedicSpeech, a TTS system tailored for the stand-up comedy synthesis in low-resource scenarios.First, we extract prosody representation by the prosody encoder and condition it to the TTS model in a flexible way.Second, we enhance the personal rhythm modeling by a conditional duration predictor.Third, we model the personal fillers by introducing comedian-related special tokens.Experiments show that ComedicSpeech achieves better expressiveness than baselines with only ten-minute training data for each comedian.The audio samples are available at https://xh621.github.io/stand-up-comedy-demo/ Yuyue Wang 0003, Yihan Wu 0008, Ruihua Song |
INTERSPEECH | 4 |
| 2023 | TeViS: Translating Text Synopses to Video StoryboardsabstractA video storyboard is a roadmap for video creation which consists of shot-by-shot images to visualize key plots in a text synopsis. Creating video storyboards, however, remains challenging which not only requires cross-modal association between high-level texts and images but also demands long-term reasoning to make transitions smooth across shots. In this paper, we propose a new task called Text synopsis to Video Storyboard (TeViS) which aims to retrieve an ordered sequence of images as the video storyboard to visualize the text synopsis. We construct a MovieNet-TeViS dataset based on the public MovieNet dataset [17]. It contains 10K text synopses each paired with keyframes manually selected from corresponding movies by considering both relevance and cinematic coherence. To benchmark the task, we present strong CLIP-based baselines and a novel VQ-Trans model. VQ-Trans first encodes text synopsis and images into a joint embedding space and uses vector quantization (VQ) to improve the visual representation. Then, it auto-regressively generates a sequence of visual features for retrieval and ordering. Experimental results demonstrate that VQ-Trans significantly outperforms prior methods and the CLIP-based baselines. Nevertheless, there is still a large gap compared to human performance suggesting room for promising future work. The code and data are available at: https://ruc-aimind.github.io/projects/TeViS/ Xu Gu 0003, Yuchong Sun, Feiyue Ni, Shizhe Chen, Xihua Wang 0002, Ruihua Song |
ACM Multimedia | 6 |
| 2023 | TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldabstractTo facilitate the research on intelligent and human-like chatbots with multi-modal context, we introduce a new video-based multi-modal dialogue dataset, called TikTalk. We collect 38K videos from a popular video-sharing platform, along with 367K conversations posted by users beneath them. Users engage in spontaneous conversations based on their multi-modal experiences from watching videos, which helps recreate real-world chitchat context. Compared to previous multi-modal dialogue datasets, the richer context types in TikTalk lead to more diverse conversations, but also increase the difficulty in capturing human interests from intricate multi-modal information to generate personalized responses. Moreover, external knowledge is more frequently evoked in our dataset. These facts reveal new challenges for multi-modal dialogue models. We quantitatively demonstrate the characteristics of TikTalk, propose a video-based multi-modal chitchat task, and evaluate several dialogue baselines. Experimental results indicate that the models incorporating large language models (LLM) can generate more diverse responses, while the model utilizing knowledge graphs to introduce external knowledge performs the best overall. Furthermore, no existing model can solve all the above challenges well. There is still a large room for future improvements, even for LLM with visual extensions. Our dataset is available at https://ruc-aimind.github.io/projects/TikTalk/. Hongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu 0002, Jingyuan Wen, Di Hu 0001, Ruihua Song, Wayne Xin Zhao, Qin Jin, Zhiwu Lu 0001 |
ACM Multimedia | 8 |
| 2023 | Going Beyond Closed Sets: A Multimodal Perspective for Video Emotion Analysis
Hao Pu, Yuchong Sun, Ruihua Song, Xu Chen 0017, Hao Jiang 0022, Zhao Cao |
PRCV (6) | 3 |
| 2023 | Expanding the Horizons: Exploring Further Steps in Open-Vocabulary Segmentation
Xihua Wang 0002, Yuchong Sun, Ruihua Song |
PRCV (10) | 5 |
| 2022 | Text2Poster: Laying Out Stylized Texts on Retrieved ImagesabstractPoster generation is a significant task for a wide range of applications, which is often time-consuming and requires lots of manual editing and artistic experience. In this paper, we propose a novel data-driven framework, called Text2Poster, to automatically generate visually-effective posters from textual information. Imitating the process of manual poster editing, our framework leverages a large-scale pretrained visual-textual model to retrieve background images from given texts, lays out the texts on the images iteratively by cascaded autoencoders, and finally, stylizes the texts by a matching-based method. We learn the modules of the framework by weakly-and self-supervised learning strategies, mitigating the demand for labeled data. Both objective and subjective experiments demonstrate that our Text2Poster outperforms state-of-the-art methods, including academic research and commercial software, on the quality of generated posters. Chuhao Jin, Hongteng Xu, Ruihua Song, Zhiwu Lu 0001 |
ICASSP | 3 |
| 2022 | AdaSpeech 4: Adaptive Text to Speech in Zero-Shot ScenariosabstractAdaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen speakers have diverse characteristics, zero-shot adaptive TTS requires strong generalization ability on speaker characteristics, which brings modeling challenges. In this paper, we develop AdaSpeech 4, a zero-shot adaptive TTS system for high-quality speech synthesis. We model the speaker characteristics systematically to improve the generalization on new speakers. Generally, the modeling of speaker characteristics can be categorized into three steps: extracting speaker representation, taking this speaker representation as condition, and synthesizing speech/mel-spectrogram given this speaker representation. Accordingly, we improve the modeling in three steps: 1) To extract speaker representation with better generalization, we factorize the speaker characteristics into basis vectors and extract speaker representation by weighted combining of these basis vectors through attention. 2) We leverage conditional layer normalization to integrate the extracted speaker representation to TTS model. 3) We propose a novel supervision loss based on the distribution of basis vectors to maintain the corresponding speaker characteristics in generated mel-spectrograms. Without any fine-tuning, AdaSpeech 4 achieves better voice quality and similarity than baselines in multiple datasets. Yihan Wu 0008, Xu Tan 0003, Bohan Li 0003, Lei He 0005, Sheng Zhao 0002, Ruihua Song, Tao Qin 0001, Tie-Yan Liu |
INTERSPEECH | 6 |
| 2022 | Self-supervised Context-aware Style Representation for Expressive Speech SynthesisabstractExpressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction.Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is costly to acquire and difficult to define and annotate accurately.In this paper, we propose a novel framework for learning style representation from abundant plain text in a self-supervised manner.It leverages an emotion lexicon and uses contrastive learning and deep clustering.We further integrate the style representation as a conditioned embedding in a multi-style Transformer TTS.Comparing with multi-style TTS by predicting style tags trained on the same dataset but with human annotations, our method achieves improved results according to subjective evaluations on both in-domain and out-of-domain test sets in audiobook speech.Moreover, with implicit context-aware style representation, the emotion transition of synthesized audio in a long paragraph appears more natural.The audio samples are available on the demo website. Yihan Wu 0008, Xi Wang 0016, Shaofei Zhang, Lei He 0005, Ruihua Song, Jian-Yun Nie |
INTERSPEECH | 5 |
| 2022 | Multi-Modal Experience Inspired AI CreationabstractAI creation, such as poem or lyrics generation, has attracted increasing attention from both industry and academic communities, with many promising models proposed in the past few years. Existing methods usually estimate the outputs based on single and independent visual or textual information. However, in reality, humans usually make creations according to their experiences, which may involve different modalities and be sequentially correlated. To model such human capabilities, in this paper, we define and solve a novel AI creation problem based on human experiences. More specifically, we study how to generate texts based on sequential multi-modal information. Compared with the previous works, this task is much more difficult because the designed model has to well understand and adapt the semantics among different modalities and effectively convert them into the output in a sequential manner. To alleviate these difficulties, we firstly design a multi-channel sequence-to-sequence architecture equipped with a multi-modal attention network. For more effective optimization, we then propose a curriculum negative sampling strategy tailored for the sequential inputs. To benchmark this problem and demonstrate the effectiveness of our model, we manually labeled a new multi-modal experience dataset. With this dataset, we conduct extensive experiments by comparing our model with a series of representative baselines, where we can demonstrate significant improvements in our model based on both automatic and human-centered metrics. The code and data are available at: https://github.com/Aman-4-Real/MMTG. Qian Cao 0001, Xu Chen 0017, Ruihua Song, Hao Jiang 0022, Zhao Cao |
ACM Multimedia | 3 |
| 2022 | Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningabstractLarge-scale video-language pre-training has shown significant improvement in video-language understanding tasks. Previous studies of video-language pretraining mainly focus on short-form videos (i.e., within 30 seconds) and sentences, leaving long-form video-language pre-training rarely explored. Directly learning representation from long-form videos and language may benefit many long-formvideo-language understanding tasks. However, it is challenging due to the difficulty of modeling long-range relationships and the heavy computational burden caused by more frames. In this paper, we introduce a Long-Form VIdeo-LAnguage pre-training model (LF-VILA) and train it on a large-scale long-form video and paragraph dataset constructed from an existing public dataset. To effectively capturethe rich temporal dynamics and to better align video and language in an efficient end-to-end manner, we introduce two novel designs in our LF-VILA model. We first propose a Multimodal Temporal Contrastive (MTC) loss to learn the temporal relation across different modalities by encouraging fine-grained alignment between long-form videos and paragraphs. Second, we propose a Hierarchical Temporal Window Attention (HTWA) mechanism to effectively capture long-range dependency while reducing computational cost in Transformer. We fine-tune the pre-trained LF-VILA model on seven downstream long-form video-language understanding tasks of paragraph-to-video retrieval and long-form video question-answering, and achieve new state-of-the-art performances. Specifically, our model achieves 16.1% relative improvement on ActivityNet paragraph-to-video retrieval task and 2.4% on How2QA task, respectively. We release our code, dataset, and pre-trained models at https://github.com/microsoft/XPretrain. Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu 0001, Huan Yang 0005, Jianlong Fu |
NeurIPS | 3 |
| 2022 | Class-Aware Sounding Objects Localization via Audiovisual CorrespondenceabstractAudiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without category annotations, i.e., localizing the sounding object and recognizing its category. To address this problem, we propose a two-stage step-by-step learning framework to localize and recognize sounding objects in complex audiovisual scenarios using only the correspondence between audio and vision. First, we propose to determine the sounding area via coarse-grained audiovisual correspondence in the single source cases. Then visual features in the sounding area are leveraged as candidate object representations to establish a category-representation object dictionary for expressive visual character extraction. We generate class-aware object localization maps in cocktail-party scenarios and use audiovisual correspondence to suppress silent areas by referring to this dictionary. Finally, we employ category-level audiovisual consistency as the supervision to achieve fine-grained audio and sounding object distribution alignment. Experiments on both realistic and synthesized videos show that our model is superior in localizing and recognizing objects as well as filtering out silent ones. We also transfer the learned audiovisual network into the unsupervised object detection task, obtaining reasonable performance. Di Hu 0001, Yake Wei, Rui Qian 0001, Weiyao Lin, Ruihua Song, Ji-Rong Wen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Leveraging Narrative to Generate Movie ScriptabstractGenerating a text based on a predefined guideline is an interesting but challenging problem. A series of studies have been carried out in recent years. In dialogue systems, researchers have explored driving a dialogue based on a plan, while in story generation, a storyline has also been proved to be useful. In this article, we address a new task—generating movie scripts based on a predefined narrative. As an early exploration, we study this problem in a “retrieval-based” setting. We propose a model (ScriptWriter-CPre) to select the best response (i.e., next script line) among the candidates that fit the context (i.e., previous script lines) as well as the given narrative. Our model can keep track of what in the narrative has been said and what is to be said. Besides, it can also predict which part of the narrative should be paid more attention to when selecting the next line of script. In our study, we find the narrative plays a different role than the context. Therefore, different mechanisms are designed for deal with them. Due to the unavailability of data for this new application, we construct a new large-scale data collection GraphMovie from a movie website where end-users can upload their narratives freely when watching a movie. This new dataset is made available publicly to facilitate other studies in text generation under the guideline. Experimental results on the dataset show that our proposed approach based on narratives significantly outperforms the baselines that simply use the narrative as a kind of context. Yutao Zhu 0001, Ruihua Song, Jian-Yun Nie, Pan Du 0001, Zhicheng Dou |
ACM Trans. Inf. Syst. | 2 |
| 2020 | ScriptWriter: Narrative-Guided Script GenerationabstractIt is appealing to have a system that generates a story or scripts automatically from a storyline, even though this is still out of our reach.In dialogue systems, it would also be useful to drive dialogues by a dialogue plan.In this paper, we address a key problem involved in these applications -guiding a dialogue by a narrative.The proposed model ScriptWriter selects the best response among the candidates that fit the context as well as the given narrative.It keeps track of what in the narrative has been said and what is to be said.A narrative plays a different role than the context (i.e., previous utterances), which is generally used in current dialogue systems.Due to the unavailability of data for this new application, we construct a new large-scale data collection GraphMovie from a movie website where endusers can upload their narratives freely when watching a movie.Experimental results on the dataset show that our proposed approach based on narratives significantly outperforms the baselines that simply use the narrative as a kind of context. Yutao Zhu 0001, Ruihua Song, Zhicheng Dou, Jian-Yun Nie |
ACL | 2 |
| 2020 | Knowledge Enhanced Opinion Generation from an Attitude
Ruihua Song, Hao Fu 0015, Pingping Lin, Jian-Yun Nie |
NLPCC (1) | 2 |
| 2020 | What If Bots Feel Moods?abstractFor social bots, smooth emotional transitions are essential for delivering a genuine conversation experience to users. Yet, the task is challenging because emotion is too implicit and complicated to understand. Among previous studies in the domain of retrieval-based conversational model, they only consider the factors of semantic and functional dependencies of utterances. In this paper, to implement a more empathetic retrieval-based conversation system, we incorporate emotional factors into context-response matching from two aspects: 1) On top of semantic matching, we propose an emotion-aware transition network to model the dynamic emotional flow and enhance context-response matching in retrieval-based dialogue systems with learnt intrinsic emotion features through a multi-task learning framework; 2) We design several flexible controlling mechanisms to customize social bots in terms of emotion. Extensive experiments on two benchmark datasets indicate that the proposed model can effectively track the flow of emotions throughout a human-machine conversation and significantly improve response selection in dialogues over the state-of-the-art baselines. We also empirically validate the emotion-control effects of our proposed model on three different emotional aspects. Finally, we apply such functionalities to a real IoT application. Lisong Qiu, Yingwai Shiu, Pingping Lin, Ruihua Song, Dongyan Zhao 0001, Rui Yan 0001 |
SIGIR | 4 |
| 2019 | Neural Storyboard Artist: Visualizing Stories with Coherent Image SequencesabstractA storyboard is a sequence of images to illustrate a story containing multiple sentences, which has been a key process to create different story products. In this paper, we tackle a new multimedia task of automatic storyboard creation to facilitate this process and inspire human artists. Inspired by the fact that our understanding of languages is based on our past experience, we propose a novel inspire-and-create framework with a story-to-image retriever that selects relevant cinematic images for inspiration and a storyboard creator that further refines and renders images to improve the relevancy and visual consistency. The proposed retriever dynamically employs contextual information in the story with hierarchical attentions and applies dense visual-semantic matching to accurately retrieve and ground images. The creator then employs three rendering steps to increase the flexibility of retrieved images, which include erasing irrelevant regions, unifying styles of images and substituting consistent characters. We carry out extensive experiments on both in-domain and out-of-domain visual story datasets. The proposed model achieves better quantitative performance than the state-of-the-art baselines for storyboard creation. Qualitative visualizations and user studies further verify that our approach can create high-quality storyboards even for stories in the wild. Shizhe Chen, Bei Liu 0001, Jianlong Fu, Ruihua Song, Qin Jin, Pingping Lin, Xiaoyu Qi, Chun-Ting Wang |
ACM Multimedia | 4 |
| 2019 | Neural Response Generation with Relevant Emotions for Short Text Conversation
Zhongxia Chen, Ruihua Song, Xing Xie 0001, Jian-Yun Nie, Xiting Wang, Enhong Chen |
NLPCC (1) | 2 |
| 2019 | Evaluating Image-Inspired Poetry Generation
Chao-Chung Wu, Ruihua Song, Tetsuya Sakai, Wen-Feng Cheng, Xing Xie 0001, Shou-De Lin |
NLPCC (1) | 2 |
| 2019 | From Text to Sound: A Preliminary Study on Retrieving Sound Effects to Radio StoriesabstractSound effects play an essential role in producing high-quality radio stories but require enormous labor cost to add. In this paper, we address the problem of automatically adding sound effects to radio stories with a retrieval-based model. However, directly implementing a tag-based retrieval model leads to high false positives due to the ambiguity of story contents. To solve this problem, we introduce a retrieval-based framework hybridized with a semantic inference model which helps to achieve robust retrieval results. Our model relies on fine-designed features extracted from the context of candidate triggers. We collect two story dubbing datasets through crowdsourcing to analyze the setting of adding sound effects and to train and test our proposed methods. We further discuss the importance of each feature and introduce several heuristic rules for the trade-off between precision and recall. Together with the text-to-speech technology, our results reveal a promising automatic pipeline on producing high-quality radio stories. Songwei Ge, Curtis Xuan, Ruihua Song, Chao Zou |
SIGIR | 3 |
| 2019 | Attitude Detection for One-Round Conversation: Jointly Extracting Target-Polarity PairsabstractWe tackle Attitude Detection, which we define as the task of extracting the replier's attitude, i.e., a target-polarity pair, from a given one-round conversation. While previous studies considered Target Extraction and Polarity Classification separately, we regard them as subtasks of Attitude Detection. Our experimental results show that treating the two subtasks independently is not the optimal solution for Attitude Detection, as achieving high performance in each subtask is not sufficient for obtaining correct target-polarity pairs. Our jointly trained model AD-NET substantially outperforms the separately trained models by alleviating the target-polarity mismatch problem. Moreover, we proposed a method utilising the attitude detection model to improve retrieval-based chatbots by re-ranking the response candidates with attitude features. Human evaluation indicates that with attitude detection integrated, the new responses to the sampled queries from are statistically significantly more consistent, coherent, engaging and informative than the original ones obtained from a commercial chatbot. Zhaohao Zeng, Ruihua Song, Pingping Lin, Tetsuya Sakai |
WSDM | 2 |
| 2019 | Personalized Reason Generation for Explainable Song RecommendationabstractPersonalized recommendation has received a lot of attention as a highly practical research topic. However, existing recommender systems provide the recommendations with a generic statement such as “Customers who bought this item also bought…”. Explainable recommendation, which makes a user aware of why such items are recommended, is in demand. The goal of our research is to make the users feel as if they are receiving recommendations from their friends. To this end, we formulate a new challenging problem called personalized reason generation for explainable recommendation for songs in conversation applications and propose a solution that generates a natural language explanation of the reason for recommending a song to that particular user. For example, if the user is a student, our method can generate an output such as “Campus radio plays this song at noon every day, and I think it sounds wonderful,” which the student may find easy to relate to. In the offline experiments, through manual assessments, the gain of our method is statistically significant on the relevance to songs and personalization to users comparing with baselines. Large-scale online experiments show that our method outperforms manually selected reasons by 8.2% in terms of click-through rate. Evaluation results indicate that our generated reasons are relevant to songs and personalized to users, and they attract users to click the recommendations. Guoshuai Zhao 0001, Hao Fu 0015, Ruihua Song, Tetsuya Sakai, Zhongxia Chen, Xing Xie 0001, Xueming Qian |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | A World of Difference: Divergent Word Interpretations Among People
Tianran Hu, Ruihua Song, Maya Abtahian, Philip Ding, Xing Xie 0001, Jiebo Luo 0001 |
ICWSM | 2 |
| 2017 | Understanding People Lifestyles: Construction of Urban Movement Knowledge Graph from GPS TrajectoryabstractTechnologies are increasingly taking advantage of the explosion in the amount of data generated by social multimedia (e.g., web searches, ad targeting, and urban computing). In this paper, we propose a multi-view learning framework for presenting the construction of a new urban movement knowledge graph, which could greatly facilitate the research domains mentioned above. In particular, by viewing GPS trajectory data from temporal, spatial, and spatiotemporal points of view, we construct a knowledge graph of which nodes and edges are their locations and relations, respectively. On the knowledge graph, both nodes and edges are represented in latent semantic space. We verify its utility by subsequently applying the knowledge graph to predict the extent of user attention (high or low) paid to different locations in a city. Experimental evaluations and analysis of a real-world dataset show significant improvements in comparison to state-of-the-art methods. Chenyi Zhuang, Nicholas Jing Yuan, Ruihua Song, Xing Xie 0001, Qiang Ma 0001 |
IJCAI | 3 |
| 2017 | Search by Screenshots for Universal Article Clipping in Mobile AppsabstractTo address the difficulty in clipping articles from various mobile applications (apps), we propose a novel framework called UniClip, which allows a user to snap a screen of an article to save the whole article in one place. The key task of the framework is search by screenshots , which has three challenges: (1) how to represent a screenshot; (2) how to formulate queries for effective article retrieval; and (3) how to identify the article from search results. We solve these by (1) segmenting a screenshot into structural units called blocks, (2) formulating effective search queries by considering the role of each block, and (3) aggregating the search result lists of multiple queries. To improve efficiency, we also extend our approach with learning-to-rank techniques so that we can find the desired article with only one query. Experimental results show that our approach achieves high retrieval performance ( F 1 = 0.868), which outperforms baselines based on keyword extraction and chunking methods. Learning-to-rank models improve our approach without learning by about 6%. A user study conducted to investigate the usability of UniClip reveals that ours is preferred by 21 out of 22 participants for its simplicity and effectiveness. Kazutoshi Umemoto, Ruihua Song, Jian-Yun Nie, Xing Xie 0001, Katsumi Tanaka, Yong Rui |
ACM Trans. Inf. Syst. | 2 |
| 2016 | Mining Shopping Patterns for Divergent Urban Regions by Incorporating Mobility DataabstractWhat people buy is an important aspect or view of lifestyles. Studying people's shopping patterns in different urban regions can not only provide valuable information for various commercial opportunities, but also enable a better understanding about urban infrastructure and urban lifestyle. In this paper, we aim to predict citywide shopping patterns. This is a challenging task due to the sparsity of the available data -- over 60% of the city regions are unknown for their shopping records. To address this problem, we incorporate another important view of human lifestyles, namely mobility patterns. With information on "where people go", we infer "what people buy". Moreover, to model the relations between regions, we exploit spatial interactions in our method. To that end, Collective Matrix Factorization (CMF) with an interaction regularization model is applied to fuse the data from multiple views or sources. Our experimental results have shown that our model outperforms the baseline methods on two standard metrics. Our prediction results on multiple shopping patterns reveal the divergent demands in different urban regions, and thus reflect key functional characteristics of a city. Furthermore, we are able to extract the connection between the two views of lifestyles, and achieve a better or novel understanding of urban lifestyles. Tianran Hu, Ruihua Song, Yingzi Wang, Xing Xie 0001, Jiebo Luo 0001 |
CIKM | 2 |
| 2016 | UniClip: Leveraging Web Search for Universal Clipping of Articles on MobileabstractIn this paper we address the difficulty of clipping articles from mobile apps. We propose a service called UniClip that allows a user to save the full content of an article by snapping a screenshot part of it. UniClip leverages a huge amount of indexed web data to mine the article by starting with a snapped screenshot. We propose approaches to solve three challenges: (1) how to represent a screenshot; (2) how to formulate effective queries for retrieving a full article; and (3) how to rank the best URL at the top from multiple search result lists. Experimental results indicate that our approach is effective in achieving as high an $$F_1$$ F 1 measure as 0.905, which outperforms the best of three baseline methods by 18 points. Ruihua Song, Kazutoshi Umemoto, Jian-Yun Nie, Xing Xie 0001, Katsumi Tanaka, Yong Rui |
Data Sci. Eng. | 1 |
| 2016 | Enhancing web search with queries of equivalent intents
Ruihua Song, Dingquan Wang, Jian-Yun Nie, Ji-Rong Wen, Yong Yu 0001 |
Inf. Retr. J. | 1 |
| 2016 | Automatically Mining Facets for Queries from Their Search ResultsabstractWe address the problem of finding query facets which are multiple groups of words or phrases that explain and summarize the content covered by a query. We assume that the important aspects of a query are usually presented and repeated in the query’s top retrieved documents in the style of lists, and query facets can be mined out by aggregating these significant lists. We propose a systematic solution, which we refer to as QDMiner, to automatically mine query facets by extracting and grouping frequent lists from free text, HTML tags, and repeat regions within top search results. Experimental results show that a large number of lists do exist and useful query facets can be mined by QDMiner. We further analyze the problem of list duplication, and find better query facets can be mined by modeling fine-grained similarities between lists and penalizing the duplicated lists. Zhicheng Dou, Zhengbao Jiang, Sha Hu 0002, Ji-Rong Wen, Ruihua Song |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2015 | Mobile Query Recommendation via Tensor Function Learning
Zhou Zhao 0001, Ruihua Song, Xing Xie 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 2 |
| 2013 | Summary of the NTCIR-10 INTENT-2 task: subtopic mining and search result diversificationabstractThe NTCIR INTENT task comprises two subtasks: {\em Subtopic Mining}, where systems are required to return a ranked list of {\em subtopic strings} for each given query; and {\em Document Ranking}, where systems are required to return a diversified web search result for each given query. This paper summarises the novel features of the Second INTENT task at NTCIR-10 and its main findings, and poses some questions for future diversified search evaluation. Tetsuya Sakai, Zhicheng Dou, Takehiro Yamamoto, Yiqun Liu 0001, Min Zhang 0006, Makoto P. Kato, Ruihua Song, Mayu Iwata |
SIGIR | 7 |
| 2013 | Diversified search evaluation: lessons from the NTCIR-9 INTENT task
Tetsuya Sakai, Ruihua Song |
Inf. Retr. | 2 |
| 2013 | Mining subtopics from text fragments for a web query
Qinglei Wang, Ya-nan Qian, Ruihua Song, Zhicheng Dou, Fan Zhang 0092, Tetsuya Sakai |
Inf. Retr. | 3 |
| 2012 | Adaptive query suggestion for difficult queriesabstractQuery suggestion is a useful tool to help users formulate better queries. Although this has been found highly useful globally, its effect on different queries may vary. In this paper, we examine the impact of query suggestion on queries of different degrees of difficulty. It turns out that query suggestion is much more useful for difficult queries than easy queries. In addition, the suggestions for difficult queries should rely less on their similarity to the original query. In this paper, we use a learning-to-rank approach to select query suggestions, based on several types of features including a query performance prediction. As query suggestion has different impacts on different queries, we propose an adaptive suggestion approach that makes suggestions only for difficult queries. We carry out experiments on real data from a search engine. Our results clearly indicate that an approach targeting difficult queries can bring higher gain than a uniform suggestion approach. Yang Liu 0005, Ruihua Song, Jian-Yun Nie, Ji-Rong Wen |
SIGIR | 2 |
| 2012 | New assessment criteria for query suggestionabstractQuery suggestion is a useful tool to help users express their information needs by supplying alternative queries. When evaluating the effectiveness of query suggestion algorithms, many previous studies focus on measuring whether a suggestion query is relevant or not to the input query. This assessment criterion is too simple to describe users' requirements. In this paper, we introduce two scenarios of query suggestion. The first scenario represents cases where the search result of the input query is unsatisfactory. The second scenario represents cases where the search result is satisfactory but the user may be looking for alternative solutions. Based on the two scenarios, we propose two assessment criteria. Our labeling results indicate that the new assessment criteria provide finer distinctions among query suggestions than the traditional relevance-based criterion. Zhongrui Ma, Ruihua Song, Tetsuya Sakai, Jiaheng Lu, Ji-Rong Wen |
SIGIR | 3 |
| 2011 | Finding dimensions for queriesabstractWe address the problem of finding multiple groups of words or phrases that explain the underlying query facets, which we refer to as query dimensions. We assume that the important aspects of a query are usually presented and repeated in the query's top retrieved documents in the style of lists, and query dimensions can be mined out by aggregating these significant lists. Experimental results show that a large number of lists do exist in the top results, and query dimensions generated by grouping these lists are useful for users to learn interesting knowledge about the queries. Zhicheng Dou, Sha Hu 0002, Yulong Luo, Ruihua Song, Ji-Rong Wen |
CIKM | 4 |
| 2011 | Evaluating diversified search results using per-intent graded relevanceabstractSearch queries are often ambiguous and/or underspecified. To accomodate different user needs, search result diversification has received attention in the past few years. Accordingly, several new metrics for evaluating diversification have been proposed, but their properties are little understood. We compare the properties of existing metrics given the premises that (1) queries may have multiple intents; (2) the likelihood of each intent given a query is available; and (3) graded relevance assessments are available for each intent. We compare a wide range of traditional and diversified IR metrics after adding graded relevance assessments to the TREC 2009 Web track diversity task test collection which originally had binary relevance assessments. Our primary criterion is discriminative power, which represents the reliability of a metric in an experiment. Our results show that diversified IR experiments with a given number of topics can be as reliable as traditional IR experiments with the same number of topics, provided that the right metrics are used. Moreover, we compare the intuitiveness of diversified IR metrics by closely examining the actual ranked lists from TREC. We show that a family of metrics called D#-measures have several advantages over other metrics such as α-nDCG and Intent-Aware metrics. Tetsuya Sakai, Ruihua Song |
SIGIR | 2 |
| 2011 | Multi-dimensional search result diversificationabstractMost existing search result diversification algorithms diversify search results in terms of a specific dimension. In this paper, we argue that search results should be diversified in a multi-dimensional way, as queries are usually ambiguous at different levels and dimensions. We first explore mining subtopics from four types of data sources, including anchor texts, query logs, search result clusters, and web sites. Then we propose a general framework that explicitly diversifies search results based on multiple dimensions of subtopics. It balances the relevance of documents with respect to the query and the novelty of documents by measuring the coverage of subtopics. Experimental results on the TREC 2009 Web track dataset indicate that combining multiple types of subtopics do help better understand user intents. By incorporating multiple types of subtopics, our models improve the diversity of search results over the sole use of one of them, and outperform two state-of-the-art models. Zhicheng Dou, Sha Hu 0002, Ruihua Song, Ji-Rong Wen |
WSDM | 4 |
| 2011 | Select-the-Best-Ones: A new way to judge relative relevance
Ruihua Song, Qingwei Guo, Ruochi Zhang, Guomao Xin, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon |
Inf. Process. Manag. | 1 |
| 2010 | Learning Query Ambiguity Models by Using Search Logs
Ruihua Song, Zhicheng Dou, Hsiao-Wuen Hon, Yong Yu 0001 |
J. Comput. Sci. Technol. | 1 |
| 2009 | Clustering queries for better document rankingabstractDifferent queries require different ranking methods. It is however challenging to determine what queries are similar, and how to rank documents for them. In this paper, we propose a new method to cluster queries according to the similarity determined based on URLs in their answers. We then train specific ranking models for each query cluster. In addition, a cluster-specific measure of authority is defined to favor documents from authoritative websites on the corresponding topics. The proposed approach is tested using data from a search engine. It turns out that our proposed topic-dependent models can significantly improve the search results of eight most popular categories of queries. Liangjie Zhang, Ruihua Song, Jian-Yun Nie, Ji-Rong Wen |
CIKM | 3 |
| 2009 | Efficient record-level wrapper inductionabstractWeb information is often presented in the form of record, e.g., a product record on a shopping website or a personal profile on a social utility website. Given a host webpage and related information needs, how to identify relevant records as well as their internal semantic structures is critical to many online information systems. Wrapper induction is one of the most effective methods for such tasks. However, most traditional wrapper techniques have issues dealing with web records since they are designed to extract information from a page, not a record. We propose a record-level wrapper system. In our system, we use a novel ``broom'' structure to represent both records and generated wrappers. With such representation, our system is able to effectively extract records and identify their internal semantics at the same time. We test our system on 16 real-life websites from four different domains. Experimental results demonstrate 99\% extraction accuracy in terms of F1-Value. Shuyi Zheng, Ruihua Song, Ji-Rong Wen, C. Lee Giles |
CIKM | 2 |
| 2009 | Using anchor texts with their hyperlink structure for web searchabstractAs a good complement to page content, anchor texts have been extensively used, and proven to be useful, in commercial search engines. However, anchor texts have been assumed to be independent, whether they come from the same Web site or not. Intuitively, an anchor text from unrelated Web sites should be considered as stronger evidence than that from the same site. This paper proposes two new methods to take into account the possible relationships between anchor texts. We consider two relationships in this paper: links from the same site and links from related sites. The importance assigned to the anchor texts in these two situations is discounted. Experimental results show that these two new models outperform the baseline model which assumes independence between hyperlinks. Zhicheng Dou, Ruihua Song, Jian-Yun Nie, Ji-Rong Wen |
SIGIR | 2 |
| 2009 | Identification of ambiguous queries in web search
Ruihua Song, Zhenxiao Luo, Jian-Yun Nie, Yong Yu 0001, Hsiao-Wuen Hon |
Inf. Process. Manag. | 1 |
| 2009 | Evaluating the Effectiveness of Personalized Web SearchabstractAlthough personalized search has been under way for many years and many personalization algorithms have been investigated, it is still unclear whether personalization is consistently effective on different queries for different users and under different search contexts. In this paper, we study this problem and provide some findings. We present a large-scale evaluation framework for personalized search based on query logs and then evaluate five personalized search algorithms (including two click-based ones and three topical-interest-based ones) using 12-day query logs of Windows Live Search. By analyzing the results, we reveal that personalized Web search does not work equally well under various situations. It represents a significant improvement over generic Web search for some queries, while it has little effect and even harms query performance under some situations. We propose click entropy as a simple measurement on whether a query should be personalized. We further propose several features to automatically predict when a query will benefit from a specific personalization algorithm. Experimental results show that using a personalization algorithm for queries selected by our prediction model is better than using it simply for all queries. Zhicheng Dou, Ruihua Song, Ji-Rong Wen, Xiaojie Yuan |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Are click-through data adequate for learning web search rankings?abstractLearning-to-rank algorithms, which can automatically adapt ranking functions in web search, require a large volume of training data. A traditional way of generating training examples is to employ human experts to judge the relevance of documents. Unfortunately, it is difficult, time-consuming and costly. In this paper, we study the problem of exploiting click-through data for learning web search rankings that can be collected at much lower cost. We extract pairwise relevance preferences from a large-scale aggregated click-through dataset, compare these preferences with explicit human judgments, and use them as training examples to learn ranking functions. We find click-through data are useful and effective in learning ranking functions. A straightforward use of aggregated click-through data can outperform human judgments. We demonstrate that the strategies are only slightly affected by fraudulent clicks. We also reveal that the pairs which are very reliable, e.g., the pairs consisting of documents with large click frequency differences, are not sufficient for learning. Zhicheng Dou, Ruihua Song, Xiaojie Yuan, Ji-Rong Wen |
CIKM | 2 |
| 2008 | Viewing Term Proximity from a Different Perspective
Ruihua Song, Michael J. Taylor 0001, Ji-Rong Wen, Hsiao-Wuen Hon, Yong Yu 0001 |
ECIR | 1 |
| 2008 | Pictor: an interactive system for importing data from a websiteabstractWe present a demonstration of an interactive wrapper induction system, called Pictor, which is able to minimize labeling cost, yet extract data with high accuracy from a website. Our demonstration will introduce two proposed technologies: record-level wrappers and a wrapper-assisted labeling strategy. These approaches allow Pictor to exploit previously generated wrappers, in order to predict similar labels in a partially labeled webpage or a completely new webpage. Our experiment results show the effectiveness of the Pictor system. Shuyi Zheng, Matthew R. Scott, Ruihua Song, Ji-Rong Wen |
KDD | 3 |
| 2007 | Template-Independent News Extraction Based on Visual Consistency
Shuyi Zheng, Ruihua Song, Ji-Rong Wen |
AAAI | 2 |
| 2007 | Joint optimization of wrapper generation and template detectionabstractMany websites have large collections of pages generated dynamically from an underlying structured source like a database. The data of a category are typically encoded into similar pages by a common script or template. In recent years, some value-added services, such as comparison shopping and vertical search in a specific domain, have motivated the research of extraction technologies with high accuracy. Almost all previous works assume that input pages of a wrapper induction system conform to a common template and they can be easily identified in terms of a common schema of URL. However, we observed that it is hard to distinguish different templates using dynamic URLs today. Moreover, since extraction accuracy heavily depends on how consistent input pages are, we argue that it is risky to determine whether pages share a common template solely based on URLs. Instead, we propose a new approach that utilizes similarity between pages to detect templates. Our approach separates pages with notable inner differences and then generates wrappers, respectively. Experimental results show that our proposed approach is feasible and effective for improving extraction accuracy. Shuyi Zheng, Ruihua Song, Ji-Rong Wen |
KDD | 2 |
| 2007 | A large-scale evaluation and analysis of personalized search strategiesabstractAlthough personalized search has been proposed for many years and many personalization strategies have been investigated, it is still unclear whether personalization is consistently effective on different queries for different users, and under different search contexts. In this paper, we study this problem and get some preliminary conclusions. We present a large-scale evaluation framework for personalized search based on query logs, and then evaluate five personalized search strategies (including two click-based and three profile-based ones) using 12-day MSN query logs. By analyzing the results, we reveal that personalized search has significant improvement over common web search on some queries but it also has little effect on other queries (e.g., queries with small click entropy). It even harms search accuracy under some situations. Furthermore, we show that straightforward click-based personalization strategies perform consistently and considerably well, while profile-based ones are unstable in our experiments. We also reveal that both long-term and short-term contexts are very important in improving search performance for profile-based personalized search strategies. Zhicheng Dou, Ruihua Song, Ji-Rong Wen |
WWW | 2 |
| 2007 | Identifying ambiguous queries in web searchabstractIt is widely believed that some queries submitted to search engines are by nature ambiguous (e.g., java, apple). However, few studies have investigated the questions of "how many queries are ambiguous?" and "how can we automatically identify an ambiguous query?" This paper deals with these issues. First, we construct the taxonomy of query ambiguity, and ask human annotators to manually classify queries based upon it. From manually labeled results, we find that query ambiguity is to some extent predictable. We then use a supervised learning approach to automatically classify queries as being ambiguous or not. Experimental results show that we can correctly identify 87% of labeled queries. Finally, we estimate that about 16% of queries in a real search log are ambiguous. Ruihua Song, Zhenxiao Luo, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon |
WWW | 1 |
| 2007 | Web page title extraction and its application
Yewei Xue, Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Chin-Yew Lin, Hang Li 0001 |
Inf. Process. Manag. | 4 |
| 2006 | Exploring URL Hit Priors for Web Search
Ruihua Song, Guomao Xin, Shuming Shi 0001, Ji-Rong Wen, Wei-Ying Ma |
ECIR | 1 |
| 2005 | Efficient Browsing of Web Search Results on Mobile Devices Based on Block Importance ModelabstractIt is expected that more and more people would search the Web when they are on the move. Though conventional search engines can be directly visited from mobile devices with Web browsing capabilities, the information is not as conveniently accessible from a handheld device as it is from desktops. Existing information discovery mechanisms for searching the Web are not well-suited to mobile devices. In this paper, a block importance model is employed to assign importance values to different segments of a Web page, in order to extract and present more condensed search results to mobile users. Based on the block importance model, three presentations for displaying the result pages in different levels of detail have been designed to reduce both the number of user interactions and the overall search time. A set of user study experiments have been carried out to compare the three presentations and a commercial service on typical mobile devices. Experimental results show that our approaches can help users to explore Web search results more efficiently. Xing Xie 0001, Gengxin Miao, Ruihua Song, Ji-Rong Wen, Wei-Ying Ma |
PerCom | 3 |
| 2005 | Title extraction from bodies of HTML documents and its application to web page retrievalabstractThis paper is concerned with automatic extraction of titles from the bodies of HTML documents. Titles of HTML documents should be correctly defined in the title fields; however, in reality HTML titles are often bogus. It is desirable to conduct automatic extraction of titles from the bodies of HTML documents. This is an issue which does not seem to have been investigated previously. In this paper, we take a supervised machine learning approach to address the problem. We propose a specification on HTML titles. We utilize format information such as font size, position, and font weight as features in title extraction. Our method significantly outperforms the baseline method of using the lines in largest font size as title (20.9%-32.6% improvement in F1 score). As application, we consider web page retrieval. We use the TREC Web Track data for evaluation. We propose a new method for HTML documents retrieval using extracted titles. Experimental results indicate that the use of both extracted titles and title fields is almost always better than the use of title fields alone; the use of extracted titles is particularly helpful in the task of named page finding (23.1% -29.0% improvements). Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Hang Li 0001 |
SIGIR | 3 |
| 2005 | Gravitation-based model for information retrievalabstractThis paper proposes GBM (gravitation-based model), a physical model for information retrieval inspired by Newton's theory of gravitation. A mapping is built in this model from concepts of information retrieval (documents, queries, relevance, etc) to those of physics (mass, distance, radius, attractive force, etc). This model actually provides a new perspective on IR problems. A family of effective term weighting functions can be derived from it, including the well-known BM25 formula. This model has some advantages over most existing ones: First, because it is directly based on basic physical laws, the derived formulas and algorithms can have their explicit physical interpretation. Second, the ranking formulas derived from this model satisfy more intuitive heuristics than most of existing ones, thus have the potential to behave empirically better and to be used safely on various settings. Finally, a new approach for structured document retrieval derived from this model is more reasonable and behaves better than existing ones. Shuming Shi 0001, Ji-Rong Wen, Ruihua Song, Wei-Ying Ma |
SIGIR | 4 |
| 2004 | A Query-Dependent Duplicate Detection Approach for Large Scale Search Engines
Shaozhi Ye, Ruihua Song, Ji-Rong Wen, Wei-Ying Ma |
APWeb | 2 |
| 2004 | Learning block importance models for web pagesabstractPrevious work shows that a web page can be partitioned into multiple segments or blocks, and often the importance of those blocks in a page is not equivalent. Also, it has been proven that differentiating noisy or unimportant blocks from pages can facilitate web mining, search and accessibility. However, no uniform approach and model has been presented to measure the importance of different segments in web pages. Through a user study, we found that people do have a consistent view about the importance of blocks in web pages. In this paper, we investigate how to find a model to automatically assign importance values to blocks in a web page. We define the block importance estimation as a learning problem. First, we use a vision-based page segmentation algorithm to partition a web page into semantic blocks with a hierarchical structure. Then spatial features (such as position and size) and content features (such as the number of images and links) are extracted to construct a feature vector for each block. Based on these features, learning algorithms are used to train a model to assign importance to different segments in the web page. In our experiments, the best model can achieve the performance with Micro-F1 79% and Micro-Accuracy 85.9%, which is quite close to a person's view. Ruihua Song, Haifeng Liu 0001, Ji-Rong Wen, Wei-Ying Ma |
WWW | 1 |