Periodical

A Common Misassumption in Online Experiments with Machine Learning Models.

Bibliographic Details
Title: A Common Misassumption in Online Experiments with Machine Learning Models.
Authors: Jeunen, Olivier
Source: SIGIR Forum; Jun2023, Vol. 57 Issue 1, p1-9, 9p
Abstract: Online experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal effect of replacing system variant "A" with variant "B", on some metric of interest. These variants can differ in many aspects. In this paper, we focus on the common use-case where they correspond to machine learning models. The online experiment then serves as the final arbiter to decide which model is superior, and should thus be shipped. The statistical literature on causal effect estimation from RCTs has a substantial history, which contributes deservedly to the level of trust researchers and practitioners have in this "gold standard" of evaluation practices. Nevertheless, in the particular case of machine learning experiments, we remark that certain critical issues remain. Specifically, the assumptions that are required to ascertain that A/B-tests yield unbiased estimates of the causal effect, are seldom met in practical applications. We argue that, because variants typically learn using pooled data, a lack of model interference cannot be guaranteed. This undermines the conclusions we can draw from online experiments with machine learning models. We discuss the implications this has for practitioners, and for the research literature. [ABSTRACT FROM AUTHOR]
Subject Terms: MACHINE learning, RANDOMIZED controlled trials, RESEARCH personnel
Copyright of SIGIR Forum is the property of Association for Computing Machinery and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
ISSN: 01635840
DOI: 10.1145/3636341.3636358
Database: Complementary Index