Mixtures of biased sentiment analysers

Mar 1, 2014·

Michael Salter-Townshend

Thomas Brendan Murphy

· 0 min read

Project

Abstract

Modelling bias is an important consideration when dealing with inexpert annotations. We are concerned with training a classifier to perform sentiment analysis on news media articles, some of which have been manually annotated by volunteers. The classifier is trained on the words in the articles and then applied to non-annotated articles. In previous work we found that a joint estimation of the annotator biases and the classifier parameters performed better than estimation of the biases followed by training of the classifier. An important question follows from this result: can the annotators be usefully clustered into either predetermined or data-driven clusters, based on their biases? If so, such a clustering could be used to select, drop or otherwise categorise the annotators in a crowdsourcing task. This paper presents work on fitting a finite mixture model to the annotators’ bias. We develop a model and an algorithm and demonstrate its properties on simulated data. We then demonstrate the clustering that exists in our motivating dataset, namely the analysis of potentially economically relevant news articles from Irish online news sources.

Type

Publication

Advances in Data Analysis and Classification

Last updated on Mar 1, 2014

← Bayesian inference for palaeoclimate with time uncertainty and stochastic volatility Jan 1, 2015

Exploring the Relationship between Membership Turnover and Productivity in Online Communities Jan 30, 2014 →