EDBT 2026 Demo / reviewers in the wild / expert
Mengdi Xu
dblp:87/8693
· DBLP profile ↗
31ranked-venue papers
10as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Meta-Decoupler-inspired Single-Universal Domain Generalization for Intelligent Fault Diagnosis
Mengdi Xu, Biliang Lu, Zhaolin Liu, Qingshuai Sun |
Expert Syst. Appl. | 1 |
| 2025 | A novel domain-private-suppress meta-recognition network based universal domain generalization for machinery fault diagnosis
Mengdi Xu, Biliang Lu, Zhaolin Liu, Qingshuai Sun |
Knowl. Based Syst. | 1 |
| 2024 | Embodied Executable Policy Learning with Language-based Scene SummarizationabstractJielin Qiu, Mengdi Xu, William Han, Seungwhan Moon, Ding Zhao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jielin Qiu, Mengdi Xu, William Jongwon Han, Seungwhan Moon, Ding Zhao |
NAACL-HLT | 2 |
| 2024 | Economic policies assessment and judgement during the pandemic with semantic and social network joint analysisabstractThis paper delves into the economic policies of China during the pandemic and investigates the relationships between policy-issuing institutions. Firstly, we conduct keyword extraction and statistical analysis based on policy texts to understand policy contents and distribution. Then, we calculate the co-occurrence matrix of policy keywords and publishers from social networks and visualise the relationships with UCINET6 and Gephi software. Through a combination of semantic analysis and social network analysis, we examine the content of relevant economic policies, laws, regulations, and the relationships between their publishers. Our findings reveal three stages of China’s economic policies during the crisis, i.e. the shock, stable, and boost periods, which align with the crisis’s impact on China’s economy. Initial policies supported small-sized enterprises (SMEs), followed by a focus on industries like tourism. During the boost phase, policies underscored various support measures, including tax and fee reductions. We also identified a “local agglomeration” characteristic among policy-issuing entities, suggesting potential improvements in cooperation, especially at the provincial level. Our findings provide valuable insights for future policy design in response to public security events. Xing Wan, Tianyou Zhu, Yuyue Wang 0007, Mengdi Xu, Zhenzhen Wu |
Connect. Sci. | 5 |
| 2024 | Functional optimal transport: regularized map estimation and domain adaptation for functional dataabstractWe introduce a formulation of regularized optimal transport problem for distributions on function spaces, where the stochastic map between functional domains can be approximated in terms of an (infinite-dimensional) Hilbert-Schmidt operator mapping a Hilbert space of functions to another. For numerous machine learning applications, data can be naturally viewed as samples drawn from spaces of functions, such as curves and surfaces, in high dimensions. Optimal transport for functional data analysis provides a useful framework of treatment for such domains. Since probability measures in infinite dimensional spaces generally lack absolute continuity (i.e., with respect to non-degenerate Gaussian measures), the Monge map in the standard optimal transport theory for finite dimensional spaces typically does not exist in the functional settings arising in such machine learning applications. This necessitates a suitable notion of approximation for the best pushforward measure to be obtained via a transport map. Indeed, our approach to the transportation problem in functional spaces is by a suitable regularization technique --- we restrict the class of transport maps to be a Hilbert-Schmidt space of operators.Within this regularization framework, we develop an efficient algorithm for finding the stochastic transport map between functional domains and provide theoretical guarantees on the existence, uniqueness, and consistency of our estimate for the Hilbert-Schmidt space of compact linear operators. We validate our method on synthetic datasets and examine the functional properties of the transport map. Experiments on real-world datasets of robot arm trajectories further demonstrate the effectiveness of our method on applications in domain adaptation. Aritra Guha, Dat Do, Mengdi Xu, XuanLong Nguyen, Ding Zhao |
J. Mach. Learn. Res. | 4 |
| 2023 | Group Distributionally Robust Reinforcement Learning with Hierarchical Latent VariablesabstractOne key challenge for multi-task Reinforcement learning (RL) in practice is the absence of task specifications. Robust RL has been applied to deal with task ambiguity but may result in over-conservative policies. To balance the worst-case (robustness) and average performance, we propose Group Distributionally Robust Markov Decision Process (GDR-MDP), a flexible hierarchical MDP formulation that encodes task groups via a latent mixture model. GDR-MDP identifies the optimal policy that maximizes the expected return under the worst-possible qualified belief over task groups within an ambiguity set. We rigorously show that GDR-MDP’s hierarchical structure improves distributional robustness by adding regularization to the worst possible outcomes. We then develop deep RL algorithms for GDR-MDP for both value-based and policy-based RL methods. Extensive experiments on Box2D control tasks, MuJoCo benchmarks, and Google football platforms show that our algorithms outperform classic robust training algorithms across diverse environments in terms of robustness under belief uncertainties. Demos are available on our project page (https://sites.google.com/view/gdr-rl/home). Mengdi Xu, Peide Huang, Yaru Niu, Visak Kumar, Jielin Qiu, Kuan-Hui Lee, Xuewei Qi, Henry Lam, Bo Li 0026, Ding Zhao |
AISTATS | 1 |
| 2023 | Cardiac Disease Diagnosis on Imbalanced Electrocardiography Data Through Optimal Transport AugmentationabstractIn this paper, we focus on a new method of data augmentation to solve the data imbalance problem within imbalanced ECG datasets to improve the robustness and accuracy of heart disease detection. By using Optimal Transport, we augment the ECG disease data from normal ECG beats to balance the data among different categories. We build a Multi-Feature Transformer (MF-Transformer) as our classification model, where different features are extracted from both time and frequency domains to diagnose various heart conditions. Our results demonstrate 1) the classification models’ ability to make competitive predictions on five ECG categories; 2) improvements in accuracy and robustness reflecting the effectiveness of our data augmentation method. Jielin Qiu, Mengdi Xu, Peide Huang, Michael A. Rosenberg, Douglas Weber, Emerson Liu, Ding Zhao |
ICASSP | 3 |
| 2023 | Hyper-Decision Transformer for Efficient Online Policy Adaptation
Mengdi Xu, Yikang Shen, Ding Zhao, Chuang Gan 0001 |
ICLR | 1 |
| 2023 | Adaptive Online Replanning with Diffusion ModelsabstractDiffusion models have risen a promising approach to data-driven planning, and have demonstrated impressive robotic control, reinforcement learning, and video planning performance. Given an effective planner, an important question to consider is replanning -- when given plans should be regenerated due to both action execution error and external environment changes. Direct plan execution, without replanning, is problematic as errors from individual actions rapidly accumulate and environments are partially observable and stochastic. Simultaneously, replanning at each timestep incurs a substantial computational cost, and may prevent successful task execution, as different generated plans prevent consistent progress to any particular goal. In this paper, we explore how we may effectively replan with diffusion models. We propose a principled approach to determine when to replan, based on the diffusion model's estimated likelihood of existing generated plans. We further present an approach to replan existing trajectories to ensure that new plans follow the same goal state as the original trajectory, which may efficiently bootstrap off previously generated plans. We illustrate how a combination of our proposed additions significantly improves the performance of diffusion planners leading to 38\% gains over past diffusion planning approaches on Maze2D and further enables handling of stochastic and long-horizon robotic control tasks. Yilun Du, Mengdi Xu, Yikang Shen, Wei Xiao 0003, Dit-Yan Yeung, Chuang Gan 0001 |
NeurIPS | 4 |
| 2023 | A trajectory is worth three sentences: multimodal transformer for offline reinforcement learningabstractTransformers hold tremendous promise in solving offline reinforcement learning (RL) by formulating it as a sequence modeling problem inspired by language modeling (LM). Prior works using transformers model a sample (trajectory) of RL as one sequence analogous to a sequence of words (one sentence) in LM, despite the fact that each trajectory includes tokens from three diverse modalities: state, action, and reward, while a sentence contains words only. Rather than taking a modality-agnostic approach which uniformly models the tokens from different modalities as one sequence, we propose a multimodal sequence modeling approach in which a trajectory (one “sentence”) of three modalities (state, action, reward) is disentangled into three unimodal ones (three “sentences”). We investigate the correlation of different modalities during sequential decision-making and use the insights to design a multimodal transformer, named Decision Transducer (DTd). DTd outperforms prior art in offline RL on the conducted D4RL benchmarks and enjoys better sample efficiency and algorithm flexibility. Our code is made publicly here. Mengdi Xu, Laixi Shi, Yuejie Chi |
UAI | 2 |
| 2022 | Prompting Decision Transformer for Few-Shot Policy GeneralizationabstractHuman can leverage prior experience and learn novel tasks from a handful of demonstrations. In contrast to offline meta-reinforcement learning, which aims to achieve quick adaptation through better algorithm design, we investigate the effect of architecture inductive bias on the few-shot learning capability. We propose a Prompt-based Decision Transformer (Prompt-DT), which leverages the sequential modeling ability of the Transformer architecture and the prompt framework to achieve few-shot adaptation in offline RL. We design the trajectory prompt, which contains segments of the few-shot demonstrations, and encodes task-specific information to guide policy generation. Our experiments in five MuJoCo control benchmarks show that Prompt-DT is a strong few-shot learner without any extra finetuning on unseen target tasks. Prompt-DT outperforms its variants and strong meta offline RL baselines by a large margin with a trajectory prompt containing only a few timesteps. Prompt-DT is also robust to prompt length changes and can generalize to out-of-distribution (OOD) environments. Project page: \href{https://mxu34.github.io/PromptDT/}{https://mxu34.github.io/PromptDT/}. Mengdi Xu, Yikang Shen, Ding Zhao, Josh Tenenbaum, Chuang Gan 0001 |
ICML | 1 |
| 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial TrainingabstractRobust Reinforcement Learning (RL) focuses on improving performances under model errors or adversarial attacks, which facilitates the real-life deployment of RL agents. Robust Adversarial Reinforcement Learning (RARL) is one of the most popular frameworks for robust RL. However, most of the existing literature models RARL as a zero-sum simultaneous game with Nash equilibrium as the solution concept, which could overlook the sequential nature of RL deployments, produce overly conservative agents, and induce training instability. In this paper, we introduce a novel hierarchical formulation of robust RL -- a general-sum Stackelberg game model called RRL-Stack -- to formalize the sequential nature and provide extra flexibility for robust training. We develop the Stackelberg Policy Gradient algorithm to solve RRL-Stack, leveraging the Stackelberg learning dynamics by considering the adversary's response. Our method generates challenging yet solvable adversarial environments which benefit RL agents' robust learning. Our algorithm demonstrates better training stability and robustness against different testing conditions in the single-agent robotics control and multi-agent highway merging tasks. Peide Huang, Mengdi Xu, Fei Fang 0001, Ding Zhao |
IJCAI | 2 |
| 2022 | Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event SamplingabstractEvaluating rare but high-stakes events is one of the main challenges in obtaining reliable reinforcement learning policies, especially in large or infinite state/action spaces where limited scalability dictates a prohibitively large number of testing iterations. On the other hand, a biased or inaccurate policy evaluation in a safety-critical system could potentially cause unexpected catastrophic failures during deployment. This paper proposes the Accelerated Policy Evaluation (APE) method, which simultaneously uncovers rare events and estimates the rare event probability in Markov decision processes. The APE method treats the environment nature as an adversarial agent and learns towards, through adaptive importance sampling, the zero-variance sampling distribution for the policy evaluation. Moreover, APE is scalable to large discrete or continuous spaces by incorporating function approximators. We investigate the convergence property of APE in the tabular setting. Our empirical studies show that APE can estimate the rare event probability with a smaller bias while only using orders of magnitude fewer samples than baselines in multi-agent and single-agent environments. Mengdi Xu, Peide Huang, Fengpei Li, Xuewei Qi, Kentaro Oguchi 0001, Henry Lam, Ding Zhao |
IROS | 1 |
| 2022 | Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationabstractCurriculum Reinforcement Learning (CRL) aims to create a sequence of tasks, starting from easy ones and gradually learning towards difficult tasks. In this work, we focus on the idea of framing CRL as interpolations between a source (auxiliary) and a target task distribution. Although existing studies have shown the great potential of this idea, it remains unclear how to formally quantify and generate the movement between task distributions. Inspired by the insights from gradual domain adaptation in semi-supervised learning, we create a natural curriculum by breaking down the potentially large task distributional shift in CRL into smaller shifts. We propose GRADIENT which formulates CRL as an optimal transport problem with a tailored distance metric between tasks. Specifically, we generate a sequence of task distributions as a geodesic interpolation between the source and target distributions, which are actually the Wasserstein barycenter. Different from many existing methods, our algorithm considers a task-dependent contextual distance metric and is capable of handling nonparametric distributions in both continuous and discrete context settings. In addition, we theoretically show that GRADIENT enables smooth transfer between subsequent stages in the curriculum under certain conditions. We conduct extensive experiments in locomotion and manipulation tasks and show that our proposed GRADIENT achieves higher performance than baselines in terms of learning efficiency and asymptotic performance. Peide Huang, Mengdi Xu, Laixi Shi, Fei Fang 0001, Ding Zhao |
NeurIPS | 2 |
| 2021 | Context-Aware Safe Reinforcement Learning for Non-Stationary EnvironmentsabstractSafety is a critical concern when deploying reinforcement learning agents for realistic tasks. Recently, safe reinforcement learning algorithms have been developed to optimize the agent’s performance while avoiding violations of safety constraints. However, few studies have addressed the nonstationary disturbances in the environments, which may cause catastrophic outcomes. In this paper, we propose the context-aware safe reinforcement learning (CASRL) method, a metal-earning framework to realize safe adaptation in non-stationary environments. We use a probabilistic latent variable model to achieve fast inference of the posterior environment transition distribution given the context data. Safety constraints are then evaluated with uncertainty-aware trajectory sampling. Prior safety constraints are formulated with domain knowledge to improve safety during exploration. The algorithm is evaluated in realistic safety-critical environments with non-stationary disturbances. Results show that the proposed algorithm significantly outperforms existing baselines in terms of safety and robustness. Baiming Chen, Zuxin Liu, Mengdi Xu, Wenhao Ding, Liang Li 0004, Ding Zhao |
ICRA | 4 |
| 2021 | DeepGAN: Generating Molecule for Drug Discovery Based on Generative Adversarial NetworkabstractAs one of the most core links in the pharmaceutical industry, drug discovery is an important direction for the application of artificial intelligence technology. It is still a huge challenge to accelerate the discovery process. To address it, we have developed a generative model for de novo small-molecule based on Generative Adversarial Network algorithm called DeepGAN. It is worth mentioning that we make DeepSMILES as training object, which has avoided the limitations of SMILES. And the addition of reinforcement learning keeps away from non-differentiable problem of the discriminator. The model is trained to optimize the rewards and adversarial loss in specific areas through strategy gradient. In this way, DeepGAN compares favorably to ORGAN and its derivatives OR(W)GAN and Naive RL which have been already well-tested. The experiments indicate our model can create molecules which can maintain molecular diversity, increase validity and show improvement in the desired metrics. Mengdi Xu, Jiandong Cheng, Yirong Liu |
ISCC | 1 |
| 2021 | Delay-aware model-based reinforcement learning for continuous control
Baiming Chen, Mengdi Xu, Liang Li 0004, Ding Zhao |
Neurocomputing | 2 |
| 2021 | Deep-based Self-refined Face-top CoordinationabstractFace-top coordination, which exists in most clothes-fitting scenarios, is challenging due to varieties of attributes, implicit correlations, and tradeoffs between general preferences and individual preferences. We present a Deep-Based Self-Refined (DBSR) system to simulate face-top coordination based on intuition evaluation. To this end, we first establish a well-coordinated face-top (WCFT) dataset from fashion databases and communities. Then, we use a jointly trained CNN Deep Canonical Correlation Analysis (DCCA) method to bridge the semantic face-top gap based on the WCFT dataset to deal with general preferences. Subsequently, an irrelevance-based Optimum-path Forest (OPF) method is developed to adapt the results to individual preferences iteratively. Experimental results and user study demonstrate the effectiveness of our method. Xiaoyang Mao, Mengdi Xu, Xiaogang Jin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | CMTS: A Conditional Multiple Trajectory Synthesizer for Generating Safety-Critical Driving ScenariosabstractNaturalistic driving trajectory generation is crucial for the development of autonomous driving algorithms. However, most of the data is collected in collision-free scenarios leading to the sparsity of the safety-critical cases. When considering safety, testing algorithms in near-miss scenarios that rarely show up in off-the-shelf datasets and are costly to accumulate is a vital part of the evaluation. As a remedy, we propose a safety-critical data synthesizing framework based on variational Bayesian methods and term it as Conditional Multiple Trajectory Synthesizer (CMTS). We extend a generative model to connect safe and collision driving data by representing their distribution in the latent space and use conditional probability to adapt to different maps. Sampling from the mixed distribution enables us to synthesize the safety-critical data not shown in the safe or collision datasets. Experimental results demonstrate that the generated dataset covers many different realistic scenarios, especially the near-misses. We conclude that the use of data generated by CMTS can improve the accuracy of trajectory predictions and autonomous vehicle safety. Wenhao Ding, Mengdi Xu, Ding Zhao |
ICRA | 2 |
| 2020 | Task-Agnostic Online Reinforcement Learning with an Infinite Mixture of Gaussian ProcessesabstractContinuously learning to solve unseen tasks with limited experience has been extensively pursued in meta-learning and continual learning, but with restricted assumptions such as accessible task distributions, independently and identically distributed tasks, and clear task delineations. However, real-world physical tasks frequently violate these assumptions, resulting in performance degradation. This paper proposes a continual online model-based reinforcement learning approach that does not require pre-training to solve task-agnostic problems with unknown task boundaries. We maintain a mixture of experts to handle nonstationarity, and represent each different type of dynamics with a Gaussian Process to efficiently leverage collected data and expressively model uncertainty. We propose a transition prior to account for the temporal dependencies in streaming data and update the mixture online via sequential variational inference. Our approach reliably handles the task distribution shift by generating new models for never-before-seen dynamics and reusing old models for previously seen dynamics. In experiments, our approach outperforms alternative methods in non-stationary tasks, including classic control with changing dynamics and decision making in different driving scenarios. Mengdi Xu, Wenhao Ding, Zuxin Liu, Baiming Chen, Ding Zhao |
NeurIPS | 1 |
| 2016 | Cast2Face: Assigning Character Names Onto Faces in Movie With Actor-Character CorrespondenceabstractAutomatically identifying characters in movies has attracted researchers' interest and led to several significant and interesting applications. However, due to the vast variation in character appearance as well as the weakness and ambiguity of available annotation, it is still a challenging problem. In this paper, we investigate this problem with the supervision of actor-character name correspondence provided by the movie cast. Our proposed framework, namely, Cast2Face, is featured by: 1) we restrict the assigned names within the set of character names in the cast; 2) for each character, by using the corresponding actor and movie name as keywords, we retrieve from the Google image search and get a group of face images to form the gallery set; 3) the probe face tracks in the movie are then identified as one of the actors by a robust kernel multitask joint sparse representation and classification method; and 4) the conditional random field model with consideration of the constraints between face tracks is introduced to enhance the final labeling. Finally, the assigned actor name of a face track is then mapped to the character name based on the cast again. Besides face naming, we further apply the proposed method to spotlight the summarization of a particular actor in his/her movies. We conduct extensive experiments and empirical evaluations on several feature-length movies to demonstrate the satisfying performance of our method. Guangyu Gao, Mengdi Xu, Jialie Shen 0001, Huadong Ma, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Data-Driven Affective Filtering for Images and VideosabstractIn this paper, a novel system is developed for synthesizing user-specified emotions onto arbitrary input images or videos. Other than defining the visual affective model based on empirical knowledge, a data-driven learning framework is proposed to extract the emotion-related knowledge from a set of emotion-annotated images. In a divide-and-conquer manner, the images are clustered into several emotion-specific scene subgroups for model learning. The visual affection is modeled with Gaussian mixture models based on color features of local image patches. For the purpose of affective filtering, the feature distribution of the target is aligned to the statistical model constructed from the emotion-specific scene subgroup, through a piecewise linear transformation. The transformation is derived through a learning algorithm, which is developed with the incorporation of a regularization term enforcing spatial smoothness, edge preservation, and temporal smoothness for the derived image or video transformation. Optimization of the objective function is sought via standard nonlinear method. Intensive experimental results and user studies demonstrate that the proposed affective filtering framework can yield effective and natural effects for images and videos. Teng Li 0001, Bingbing Ni, Mengdi Xu, Meng Wang 0001, Qingwei Gao, Shuicheng Yan |
IEEE Trans. Cybern. | 3 |
| 2014 | Touch Saliency: Characteristics and PredictionabstractIn this work, we propose an alternative ground truth to the eye fixation map in visual attention study, called touch saliency. As it can be directly collected from the recorded data of users' daily browsing behavior on widely used smart phone devices with touch screens, the touch saliency data is easy to obtain. Due to the limited screen size, smart phone users usually move and zoom in the images, and fix the region of interest on the screen when browsing images. Our studies are two-fold. First, we collect and study the characteristics of these touch screen fixation maps (named touch saliency) by comprehensive comparisons with their counterpart, the eye-fixation maps (namely, visual saliency). The comparisons show that the touch saliency is highly correlated with the eye fixations for the same stimuli, which indicates its utility in data collection for visual attention study. Based on the consistency between both touch saliency and visual saliency, our second task is to propose a unified saliency prediction model for both visual and touch saliency detection. This model utilizes middle-level object category features extracted from pre-segmented image superpixels as input to the recently proposed multitask sparsity pursuit (MTSP) framework for saliency prediction. Extensive evaluations show that the proposed middle-level category features can considerably improve the saliency prediction performance when taking both touch saliency and visual saliency as ground truth. Bingbing Ni, Mengdi Xu, Tam V. Nguyen 0002, Meng Wang 0001, Congyan Lang, ZhongYang Huang, Shuicheng Yan |
IEEE Trans. Multim. | 2 |
| 2013 | Static saliency vs. dynamic saliency: a comparative studyabstractRecently visual saliency has attracted wide attention of researchers in the computer vision and multimedia field. However, most of the visual saliency-related research was conducted on still images for studying static saliency. In this paper, we give a comprehensive comparative study for the first time of dynamic saliency (video shots) and static saliency (key frames of the corresponding video shots), and two key observations are obtained: 1) video saliency is often different from, yet quite related with, image saliency, and 2) camera motions, such as tilting, panning or zooming, affect dynamic saliency significantly. Motivated by these observations, we propose a novel camera motion and image saliency aware model for dynamic saliency prediction. The extensive experiments on two static-vs-dynamic saliency datasets collected by us show that our proposed method outperforms the state-of-the-art methods for dynamic saliency prediction. Finally, we also introduce the application of dynamic saliency prediction for dynamic video captioning, assisting people with hearing impairments to better entertain videos with only off-screen voices, e.g., documentary films, news videos and sports videos. Tam V. Nguyen 0002, Mengdi Xu, Guangyu Gao, Mohan Kankanhalli, Qi Tian 0001, Shuicheng Yan |
ACM Multimedia | 2 |
| 2013 | Learning to Photograph: A Compositional PerspectiveabstractIn this paper, we present an intelligent photography system which can recommend the most user-favored view rectangle for arbitrary camera input, from a photographic compositional perspective. Automating this process is difficult, due to the subjectivity of human's aesthetics judgement and large variations of image contents, where heuristic compositional rules lack generality. Motivated by the recent prevalence of photo-sharing websites, e.g., Flickr.com, we develop a learning-based framework which discovers the underlying aesthetic photographic compositional structures from a large set of user-favored online sharing photographs and utilizes the implicitly shared knowledge among the professional photographers for aesthetically optimal view recommendation. In particular, we propose an Omni-Range Context method which explicitly encodes the spatial and geometric distributions of various visual elements in the photograph as well as cooccurrence characteristics of visual element pairs by using generative mixture models. Searching the optimal view rectangle is then formulated as maximum a posterior by imposing the trained prior distributions along with additional photographic constraints. The proposed system has the potential to operate in near real-time. Comprehensive user studies well demonstrate the effectiveness of the proposed framework for aesthetically optimal view recommendation. Bingbing Ni, Mengdi Xu, Bin Cheng 0001, Meng Wang 0001, Shuicheng Yan, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Omni-range spatial contexts for visual classificationabstractSpatial contexts encode rich discriminative information for visual classification. However, as object shapes and scales vary significantly among images, spatial contexts with manually specified distance ranges are not guaranteed with optimality. In this work, we investigate how to automatically select discriminative and stable distance bin groups for modeling image spatial contexts to improve classification performance. We make two observations. First, the number of distance bins for context modeling can be arbitrarily large, and discriminative contexts are only from a small subset of distance bins. Second, adjacent distance bins for contexts modeling often show similar characteristics, thus encouraging grouping them together can result in more stable representation. Utilizing these two observations, we propose an omni-range spatial context mining framework for image classification. A sparse selection and grouping regularizer is employed along with an empirical risk, to discover discriminative and stable distance bin groups for context modeling. To facilitate efficient optimization, the objective function is approximated by a smooth convex function with theoretically guaranteed error bounds. The selected and grouped image spatial contexts, which are applied in food and national flag recognition, are demonstrated to be discriminative, compact and robust. Bingbing Ni, Mengdi Xu, Jinhui Tang 0001, Shuicheng Yan, Pierre Moulin |
CVPR | 2 |
| 2012 | Touch saliencyabstractIn this work, we propose a new concept of touch saliency, and attempt to answer the question of whether the underlying image saliency map may be implicitly derived from the accumulative touch behaviors (or more specifically speaking, zoom-in and panning manipulations) when many users browse the image on smart mobile devices with multi-touch display of small size. The touch saliency maps are collected for the images of the recently released NUSEF dataset, and the preliminary comparison study demonstrates: 1) the touch saliency map is highly correlated with human eye fixation map for the same stimuli, yet compared to the latter, the touch data collection is much more flexible and requires no cooperation from users; and 2) the touch saliency is also well predictable by popular saliency detection algorithms. This study opens a new research direction of multimedia analysis by harnessing human touch information on increasingly popular multi-touch smart mobile devices. Mengdi Xu, Bingbing Ni, Jian Dong 0011, ZhongYang Huang, Meng Wang 0001, Shuicheng Yan |
ACM Multimedia | 1 |
| 2011 | Video accessibility enhancement for hearing-impaired usersabstractThere are more than 66 million people suffering from hearing impairment and this disability brings them difficulty in video content understanding due to the loss of audio information. If the scripts are available, captioning technology can help them in a certain degree by synchronously illustrating the scripts during the playing of videos. However, we show that the existing captioning techniques are far from satisfactory in assisting the hearing-impaired audience to enjoy videos. In this article, we introduce a scheme to enhance video accessibility using a Dynamic Captioning approach, which explores a rich set of technologies including face detection and recognition, visual saliency analysis, text-speech alignment, etc. Different from the existing methods that are categorized as static captioning, dynamic captioning puts scripts at suitable positions to help the hearing-impaired audience better recognize the speaking characters. In addition, it progressively highlights the scripts word-by-word via aligning them with the speech signal and illustrates the variation of voice volume. In this way, the special audience can better track the scripts and perceive the moods that are conveyed by the variation of volume. We implemented the technology on 20 video clips and conducted an in-depth study with 60 real hearing-impaired users. The results demonstrated the effectiveness and usefulness of the video accessibility enhancement scheme. Richang Hong, Meng Wang 0001, Xiao-Tong Yuan, Mengdi Xu, Shuicheng Yan, Tat-Seng Chua |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2010 | Dynamic captioning: video accessibility enhancement for hearing impairmentabstractThere are more than 66 million people su®ering from hearing impairment and this disability brings them di±culty in the video content understanding due to the loss of audio information. If scripts are available, captioning technology can help them in a certain degree by synchronously illustrating the scripts during the playing of videos. However, we show that the existing captioning techniques are far from satisfactory in assisting hearing impaired audience to enjoy videos. Richang Hong, Meng Wang 0001, Mengdi Xu, Shuicheng Yan, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2010 | Movie2Comics: a feast of multimedia artworkabstractAs a type of artwork, comics are prevalent and popular around the world. However, although there are several assistive software and tools available, the creation of comics is still a tedious and labor intensive process. This paper proposes a scheme that is able to automatically turn a movie to comics with two principles: (1) optimizing the information reservation of movie; and (2) generating outputs following the rules and styles of comics. The scheme mainly contains three components: script-face mapping, key-scene extraction, and cartoonization. Script-face mapping utilizes face recognition and tracking techniques to accomplish the mapping between character's faces and their scripts. Key-scene extraction then combines the frames derived from subshots and the extracted index frames based on subtitle to select a sequence of frames for cartoonization. Finally, the cartoonization is accomplished via four steps: panel scale, stylization, word balloon placement and comics layout. Experiments conducted on a set of movie clips have demonstrates the usefulness and e®ectiveness of the scheme. Richang Hong, Xiao-Tong Yuan, Mengdi Xu, Meng Wang 0001, Shuicheng Yan, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2010 | Cast2Face: character identification in movie with actor-character correspondenceabstractWe investigate the problem of automatically identifying characters in a movie with the supervision of actor-character name correspondence provided by the movie cast. Our proposed framework, namely Cast2Face, is featured by: (i) we restrict the names to assign within the set of character names in the cast; (ii) for each character, by using the corresponding actor's name as a key word, we retrieve from Google image search a group of face images to form the gallery set; and (iii) the probe face tracks in the movie are then identified as one of the actors by robust multi-task joint sparse representation and classification method. The assigned actor name on a face track is then mapped to the character name based on the cast again. In addition to face naming, we further apply the proposed method to spotlights summarization of a particular actor in his/her movies. Empirical evaluations on several feature-length movies demonstrate the satisfying performance of our method. Mengdi Xu, Xiao-Tong Yuan, Jialie Shen 0001, Shuicheng Yan |
ACM Multimedia | 1 |