EDBT 2026 Demo / reviewers in the wild / expert
Yanjun Ma
dblp:79/2225
· DBLP profile ↗
26ranked-venue papers
7as first author
5since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Graph learning · 46% Machine translation · 30% Question answering and dialogue systems · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 50% GPUs and heterogeneous computing · 50% | |
| Human-computer interaction and pervasive computing
1 paper |
Interaction techniques and input · 33% Immersive interaction · 33% Wearable and physiological sensing · 33% | |
| Computer networks
1 paper |
Network optimization and economics · 44% Wireless networking · 44% Cellular and mobile networks · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 64% Information retrieval · 36% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network training |
0.7 | 1 | 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale Recommendation · KDD 2023 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU training |
0.7 | 1 | 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale Recommendation · KDD 2023 |
High-performance computing
performance optimization at scale |
0.7 | 1 | 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale Recommendation · KDD 2023 |
Immersive interaction
augmented reality interaction |
0.4 | 1 | 2019 | ActiTouch: Robust Touch Detection for On-Skin AR/VR Interfaces · UIST 2019 |
Wearable and physiological sensing
on-skin interfaces |
0.4 | 1 | 2019 | ActiTouch: Robust Touch Detection for On-Skin AR/VR Interfaces · UIST 2019 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.3 | 1 | 2018 | Answer-focused and Position-aware Neural Question Generation · EMNLP 2018 |
Recommender systems
large-scale recommendation |
0.2 | 1 | 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale Recommendation · KDD 2023 |
Network optimization and economics › resource allocation › spectrum allocation
dynamic spectrum allocation |
0.2 | 1 | 2013 | Coverage probability driven dynamic spectrum allocation in heterogeneous wireless networks · Sci. China Inf. Sci. 2013 |
Wireless networking
heterogeneous wireless networks |
0.2 | 1 | 2013 | Coverage probability driven dynamic spectrum allocation in heterogeneous wireless networks · Sci. China Inf. Sci. 2013 |
Natural language and speech › Machine translation › computer-assisted translation
translation memory |
0.1 | 1 | 2011 | Consistent Translation using Discriminative Learning - A Translation Memory-inspired Approach · ACL 2011 |
Natural language and speech › Machine translation
statistical machine translation |
0.1 | 1 | 2010 | Bridging SMT and TM with Translation Recommendation · ACL 2010 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.1 | 1 | 2007 | Bootstrapping Word Alignment via Word Packing · ACL 2007 |
Cellular and mobile networks › coverage analysis
coverage probability |
0.0 | 1 | 2013 | Coverage probability driven dynamic spectrum allocation in heterogeneous wireless networks · Sci. China Inf. Sci. 2013 |
Methods — techniques the papers use, named apart from their topics
subgraph sampling · 2.0hierarchical storage · 2.0asynchronous pipeline · 2.0computer vision · 0.4RF waveguide · 0.4sequence-to-sequence model · 0.3answer embedding · 0.3translation recommendation · 0.2stochastic geometry · 0.2discriminative learning · 0.1word packing · 0.1bootstrapping · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale RecommendationabstractWhile having been used widely for large-scale recommendation and online advertising, the Graph Neural Network (GNN) has demonstrated its representation learning capacity to extract embeddings of nodes and edges through passing, transforming, and aggregating information over the graph. In this work, we propose PGLBox1 - a multi-GPU graph learning framework based on PaddlePaddle [24], incorporating with optimized storage, computation, and communication strategies, to train deep GNNs based on web-scale graphs for the recommendation. Specifically, PGLBox adopts a hierarchical storage system with three layers to facilitate I/O, where graphs and embeddings are stored in the HBMs and SSDs, respectively, with MEMs as the cache. To fully utilize multi-GPUs and I/O bandwidth, PGLBox proposes an asynchronous pipeline with three stages - it first samples the subgraphs from the input graph, then pulls & updates embeddings and trains GNNs on the subgraph with parameters updating queued at the end of the pipeline. Thanks to the capacity of PGLBox in handling web-scale graphs, it becomes feasible to unify the view of GNN-based recommendation tasks for multiple advertising verticals and fuse all these graphs into a unified yet huge one. We evaluate PGLBox using a bucket of realistic GNN training tasks for the recommendation, and compare the performance of PGLBox on top of a multi-GPU server (Tesla A100×8) and the legacy training system based on a 40-node MPI cluster at Baidu. The overall comparisons show that PGLBox could save up to 55% monetary cost for training GNN models, and achieve up to 14× training speedup with the same accuracy as the legacy trainer. The open-source implementation of PGLBox is available at https://github.com/PaddlePaddle/PGL/tree/main/apps/PGLBox. Xuewu Jiao, Weibin Li 0004, Xinxuan Wu, Jiang Bian 0003, Siming Dai, Xinsheng Luo, Mingqing Hu, Zhengjie Huang, Danlei Feng, Junchao Yang 0001, Shikun Feng, Haoyi Xiong, Dianhai Yu, Shuanglong Li, Jingzhou He, Yanjun Ma |
KDD | 18 |
| 2023 | HeterPS: Distributed deep learning with reinforcement learning based scheduling in heterogeneous environments
Ji Liu 0003, Danlei Feng, Minxu Zhang, Xinxuan Wu, Xuefeng Yao, Dianhai Yu, Yanjun Ma, Dejing Dou |
Future Gener. Comput. Syst. | 8 |
| 2022 | Sensor Fault Estimation in a Probabilistic Framework for Industrial Processes and its ApplicationsabstractIn this article, a new sensor fault estimation algorithm is proposed for industrial processes described by linear discrete-time systems, where the fault dynamics are modeled as a stochastic process. By performing the variational Bayesian inference, the potential sensor fault, as well as the system states, is estimated simultaneously in a probabilistic framework. It is shown that the target fault signal can be satisfactorily estimated through the proposed method, without knowing the statistics of measurement noise and fault coefficient matrix. The efficiency and superiority of the proposed method are demonstrated through numerical simulations and experimental tests performed on a hybrid tank system. Chen Xu 0009, Shunyi Zhao, Yanjun Ma, Biao Huang 0001, Fei Liu 0001, Xiaoli Luan |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Multi-criteria Confidence Evaluation for Robust Visual Tracking
Siqi Shi, Nanting Li, Yanjun Ma, Liping Zheng |
PRCV (1) | 3 |
| 2021 | Emerging trends: A gentle introduction to fine-tuningabstractAbstract The previous Emerging Trends article (Churchet al., 2021.Natural Language Engineering27(5), 631–645.) introduced deep nets to poets. Poets is an imperfect metaphor, intended as a gesture toward inclusion. The future for deep nets will benefit by reaching out to a broad audience of potential users, including people with little or no programming skills, and little interest in training models. That paper focused on inference, the use of pre-trained models, as is, without fine-tuning. The goal of this paper is to make fine-tuning more accessible to a broader audience. Since fine-tuning is more challenging than inference, the examples in this paper will require modest programming skills, as well as access to a GPU. Fine-tuning starts with a general purpose base (foundation) model and uses a small training set of labeled data to produce a model for a specific downstream application. There are many examples of fine-tuning in natural language processing (question answering (SQuAD) and GLUE benchmark), as well as vision and speech. Kenneth Church 0001, Yanjun Ma |
Nat. Lang. Eng. | 3 |
| 2019 | ActiTouch: Robust Touch Detection for On-Skin AR/VR InterfacesabstractContemporary AR/VR systems use in-air gestures or handheld controllers for interactivity. This overlooks the skin as a convenient surface for tactile, touch-driven interactions, which are generally more accurate and comfortable than free space interactions. In response, we developed ActiTouch, a new electrical method that enables precise on-skin touch segmentation by using the body as an RF waveguide. We combine this method with computer vision, enabling a system with both high tracking precision and robust touch detection. Our system requires no cumbersome instrumentation of the fingers or hands, requiring only a single wristband (e.g., smartwatch) and sensors integrated into an AR/VR headset. We quantify the accuracy of our approach through a user study and demonstrate how it can enable touchscreen-like interactions on the skin. Yang Zhang 0041, Wolf Kienzle, Yanjun Ma, Shiu S. Ng, Hrvoje Benko, Chris Harrison 0001 |
UIST | 3 |
| 2019 | Feature Extraction of Constrained Dynamic Latent VariablesabstractFeature extraction has become an essential prerequisite of developing data-based models, control and monitoring tools from massive industrial data. When the temporal correlation is significant, the latent feature is commonly described by a dynamic model, such as the state-space model. Industrial processes are widely subject to certain boundary constraints. However, most of the existing feature extraction methods have not considered the boundary constraints on the latent features. This study develops a learning approach with consideration of boundary constrained latent features. To retain dynamic behavior with a compact probability description, a novel state transition model is developed by using the Beta distribution for the constrained state. To learn the constrained dynamic feature from regularly observed data, a nonlinear observation function is incorporated, and the variational Bayesian inference is adopted for solving the problem. The effectiveness of the proposed method is demonstrated through numerical simulations along with industrial data sets. Yanjun Ma, Shunyi Zhao, Biao Huang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Robust FIR State Estimation of Dynamic Processes Corrupted by OutliersabstractThe outlier is a common issue in the design of state estimators for an industrial process. In this paper, a robust finite impulse response (FIR) filter is proposed for time-invariant state-space models with its noise following the Student's t distributions. A batch solution is first derived by maximizing the likelihood, and then, an equivalent iterative realization is given to provide a clearer insight into the FIR structure. It shows that the essence of the proposed approach is the convergence of the maximum likelihood estimates in horizon scale through iterations, and the state estimate at each sampling instant is independent of the degree-of-freedom (DOF) parameter of the Student's t distribution. Based on this, a modified algorithm that updates the DOF parameter in each iteration is further proposed. Applications to a moving target tracking example and a 3-DOF helicopter system demonstrate that the proposed methods can exhibit good immunity against outliers during the filtering. Shunyi Zhao, Yanjun Ma, Biao Huang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2018 | Answer-focused and Position-aware Neural Question GenerationabstractIn this paper, we focus on the problem of question generation (QG).Recent neural networkbased approaches employ the sequence-tosequence model which takes an answer and its context as input and generates a relevant question as output.However, we observe two major issues with these approaches: (1) The generated interrogative words (or question words) do not match the answer type.(2) The model copies the context words that are far from and irrelevant to the answer, instead of the words that are close and relevant to the answer.To address these two issues, we propose an answer-focused and position-aware neural question generation model.(1) By answerfocused, we mean that we explicitly model question word generation by incorporating the answer embedding, which can help generate an interrogative word matching the answer type.(2) By position-aware, we mean that we model the relative distance between the context words and the answer.Hence the model can be aware of the position of the context words when copying them to generate a question.We conduct extensive experiments to examine the effectiveness of our model.The experimental results show that our model significantly improves the baseline and outperforms the state-of-the-art system. Xingwu Sun, Jing Liu 0022, Yajuan Lyu, Wei He 0014, Yanjun Ma |
EMNLP | 5 |
| 2014 | 3D Polarization Projection for WINNER Channel SimulationsabstractWireless world initiative new radio (WINNER) delivers a 3D cross-polarized channel model for B3G/4G system designs. Comprehensive radiation characteristics of polarized antenna are crucial in generating channel coefficients. Being currently supported within WINNER channel model, field patterns are technically obtained by chamber measurement. However, in some channel related performance analysis scenarios, design insight can be crystallized better by starting the derivations with theoretical co-pol and cross-pol components. Specifically, these two components are mathematically linked with field patterns through the proposed polarization projection algorithm. In this paper, we focus on revealing the physical significance of polarization transform between the antenna plane and the propagation plane. In practice, it makes retrieving the field patterns by electromagnetic computation possible. Meanwhile, the impact imposed by distinct antenna orientations is geometrically illustrated and consequently incorporated into the proposed algorithm. The result is analytically verified by the closed-form expression of the dipole field pattern and we find that its 2D degenerative case is aligned with that defined in 3GPP TR 25.996. The benefit is to significantly reduce the cost on generating channel coefficients in WINNER channel simulations. Lihua Pang, Yang Zhang 0013, Yanjun Ma, Bing Lan, Jiandong Li 0001 |
VTC Spring | 3 |
| 2013 | Coverage probability driven dynamic spectrum allocation in heterogeneous wireless networks
Jiandong Li 0001, Yanjun Ma |
Sci. China Inf. Sci. | 3 |
| 2012 | Mining User's Real Social Circle in MicroblogabstractAs a media and communication platform, microblog is more and more popular around the world. Users can follow anyone ranges from well-known individuals to real friends, and read their tweets without their permission. Most users follow a large number of celebrities and public media in microblog, however, these celebrities do not necessarily follow all their fans. Such one-way relationship abounds in the user network and is displayed in the forms of users' followees and followers, which make it difficult to identify users' real friends who are contained in the merged list of followees and followers. The aim of this paper is to propose a general algorithm for mining users' real friends in social media and dividing them into different social circles automatically according to the closeness of their relationships. To verify the effectiveness of the proposed algorithm, we build a microblog application which presents the social circles for users identified by the algorithm and enable users to modify the proposed results according to her/his real social circles. We demonstrate that our algorithm is superior to traditional clustering method in terms of F measure and Mean Average Precision. Hailong Qin, Yanjun Ma |
ASONAM | 3 |
| 2012 | An Evaluation of Statistical Post-Editing Systems Applied to RBMT and SMT Systems
Hannah Béchara, Raphaël Rubino, Yifan He 0007, Yanjun Ma, Josef van Genabith |
COLING | 4 |
| 2012 | What types of word alignment improve statistical machine translation?
Patrik Lambert, Simon Petit-Renaud, Yanjun Ma, Andy Way |
Mach. Transl. | 3 |
| 2011 | Consistent Translation using Discriminative Learning - A Translation Memory-inspired Approach
Yanjun Ma, Yifan He 0007, Andy Way, Josef van Genabith |
ACL | 1 |
| 2011 | Preliminary Experiments on Using Users' Post-Editions to Enhance a SMT System Oracle-based Training for Phrase-based Statistical Machine Translation
Ankit K. Srivastava, Yanjun Ma, Andy Way |
EAMT | 2 |
| 2011 | Statistical Post-Editing for a Statistical MT System
Hannah Béchara, Yanjun Ma, Josef van Genabith |
MTSummit | 2 |
| 2011 | Rich Linguistic Features for Translation Memory-Inspired Consistent Translation
Yifan He 0007, Yanjun Ma, Andy Way, Josef van Genabith |
MTSummit | 2 |
| 2010 | Bridging SMT and TM with Translation Recommendation
Yifan He 0007, Yanjun Ma, Josef van Genabith, Andy Way |
ACL | 2 |
| 2010 | Statistical Analysis of Alignment Characteristics for Phrase-based Machine Translation
Patrik Lambert, Simon Petit-Renaud, Yanjun Ma, Andy Way |
EAMT | 3 |
| 2010 | Kam-Fai Wong, Wenjie Li, Ruifeng Xu and Zheng-sheng Zhang: Introduction to Chinese natural language processing - Morgan and Claypool publishers, 2010, x+148pp
Yanjun Ma |
Mach. Transl. | 1 |
| 2009 | Bilingually Motivated Domain-Adapted Word Segmentation for Statistical Machine Translation
Yanjun Ma, Andy Way |
EACL | 1 |
| 2009 | Using Supertags as Source Language Context in SMT
Rejwanul Haque, Sudip Kumar Naskar, Yanjun Ma, Andy Way |
EAMT | 3 |
| 2009 | Tuning Syntactically Enhanced Word Alignment for Statistical Machine Translation
Yanjun Ma, Patrik Lambert, Andy Way |
EAMT | 1 |
| 2009 | Bilingually Motivated Word Segmentation for Statistical Machine TranslationabstractWe introduce a bilingually motivated word segmentation approach to languages where word boundaries are not orthographically marked, with application to Phrase-Based Statistical Machine Translation (PB-SMT). Our approach is motivated from the insight that PB-SMT systems can be improved by optimizing the input representation to reduce the predictive power of translation models. We firstly present an approach to optimize the existing segmentation of both source and target languages for PB-SMT and demonstrate the effectiveness of this approach using a Chinese--English MT task, that is, to measure the influence of the segmentation on the performance of PB-SMT systems. We report a 5.44% relative increase in Bleu score and a consistent increase according to other metrics. We then generalize this method for Chinese word segmentation without relying on any segmenters and show that using our segmentation PB-SMT can achieve more consistent state-of-the-art performance across two domains. There are two main advantages of our approach. First of all, it is adapted to the specific translation task at hand by taking the corresponding source (target) language into account. Second, this approach does not rely on manually segmented training data so that it can be automatically adapted for different domains. Yanjun Ma, Andy Way |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2007 | Bootstrapping Word Alignment via Word Packing
Yanjun Ma, Nicolas Stroppa, Andy Way |
ACL | 1 |