Thanks for the feedback Betsy - some replies inline: (In reply to Betsy Mikel [:betsymi] from comment #3) > I'm fine to do this test in this way — that is, to add a recommendation regarding email masking with a clickable CTA — but would like us to avoid pitting recommendations and various products/services against each other when we analyze the results. There are too many variables to isolate this as a test with a true control. > Users clicking on product A versus product B doesn't mean that A is a winner and B is a loser. Product A and Product B solve strive to solve entirely different privacy-related problems. Users may or may not have a level of understanding of what those privacy problems even are. Those users may or may not be interested in learning about one right now. I broadly agree - we should not think of more users clicking on CTA for "product A" as a slam dunk case that product A is more interesting than B, has better market fit, etc. No one experiment could answer these questions definitively, and if we do this test we need to follow up with other things e.g. user research, more experiments, prototypes etc regardless of results. That said, I do think we need some kind of comparison for reference and decision making. It doesn't necessarily have to be another product - It could be just an a priori number, e.g. "we will proceed with more research if 3+% of users exposed to the CTA click on it". But presumably this number would have some reason behind it. I don't want us to run the study without a point of comparison and then have to "read the entrails" wrt the results. Say 3% of users exposed click the CTA. Is that enough to warrant further work? Maybe, especially if we have (totally making up a number here) 1.5M visitors per day. But, without some point of reference, its hard to say. Maybe 3% of users would click on a banner that just says "click here" and nothing else (I've seen weirder). Its true that FPN is a different solution for a different (but overlapping) set of problems. But, we've done research there, and have reasons to believe it has at least enough baseline interest for viability. That's what we're trying to get data on here (I think?). If we're careful about wording the CTAs to be as close to each other as possible (that's where we could really need your help), and the CTR for the FPR greatly outperforms FPN (a product that has already had some amount of vetting), I would think that this tells us *something* about its intrinsic interest level. What if FPR only performs half as well as FPN? I think that's also informative, but again not definitive. Once more, we're not trying to prove anything beyond a shadow of a doubt here, just trying to get one more shred of data to shore up our priors on what to do next. I'm totally open to other ideas about what the comparison should be here, BTW. Not at all sold that it has to be FPN. But, I think we need something. > We also cannot know is this the first time a user is seeing breach recommendations (perhaps this is their first breach), the 10th (perhaps they've received many breach alerts and have seen this page without this product countless times), or any number in between. And all privacy threats are not created equal. The fact that someone selects recommendation A doesn't mean they don't care about the others. It only means that at this very isolated moment in time, they clicked that particular one. It is true that each user will be exposed to the CTA(s) in different contexts and that this will contribute to their likelihood of clicking on it, perhaps independently of their "true" interest level in the product. However, if we do a reasonable job at random branch assignments then most of these issues should mainly just contribute noise (which will average out), rather than bias (i.e. if we enroll enough users we should have roughly the same amount of users who have been exposed to 10+ breach alerts in each branch).
Bug 1606003 Comment 4 Edit History
Note: The actual edited comment in the bug view page will always show the original commenter’s name and original timestamp.
Thanks for the feedback Betsy - some replies inline: (In reply to Betsy Mikel [:betsymi] from comment #3) > I'm fine to do this test in this way — that is, to add a recommendation regarding email masking with a clickable CTA — but would like us to avoid pitting recommendations and various products/services against each other when we analyze the results. There are too many variables to isolate this as a test with a true control. > Users clicking on product A versus product B doesn't mean that A is a winner and B is a loser. Product A and Product B solve strive to solve entirely different privacy-related problems. Users may or may not have a level of understanding of what those privacy problems even are. Those users may or may not be interested in learning about one right now. I broadly agree - we should not think of more users clicking on CTA for "product A" as a slam dunk case that product A is more interesting than B, has better market fit, etc. No one experiment could answer these questions definitively, and if we do this test we need to follow up with other things e.g. user research, more experiments, prototypes etc regardless of results. That said, I do think we need some kind of comparison for reference and decision making. It doesn't necessarily have to be another product - It could be just an a priori number, e.g. "we will proceed with more research if 3+% of users exposed to the CTA click on it". But presumably this number would have some reason behind it. I don't want us to run the study without a point of comparison and then have to "read the entrails" wrt the results. Say 3% of users exposed click the CTA. Is that enough to warrant further work? Maybe, especially if we have (totally making up a number here) 1.5M visitors per day. But, without some point of reference, its hard to say. Maybe 3% of users would click on a banner that just says "click here" and nothing else (I've seen weirder). Its true that FPN is a different solution for a different (but overlapping) set of problems. But, we've done research there, and have reasons to believe it has at least enough baseline interest for viability. That's what we're trying to get data on here (I think?). If we're careful about wording the CTAs to be as close to each other as possible (that's where we would really need your help), and the CTR for the FPR greatly outperforms FPN (a product that has already had some amount of vetting), I would think that this tells us *something* about its intrinsic interest level. What if FPR only performs half as well as FPN? I think that's also informative, but again not definitive. Once more, we're not trying to prove anything beyond a shadow of a doubt here, just trying to get one more shred of data to shore up our priors on what to do next. I'm totally open to other ideas about what the comparison should be here, BTW. Not at all sold that it has to be FPN. But, I think we need something. > We also cannot know is this the first time a user is seeing breach recommendations (perhaps this is their first breach), the 10th (perhaps they've received many breach alerts and have seen this page without this product countless times), or any number in between. And all privacy threats are not created equal. The fact that someone selects recommendation A doesn't mean they don't care about the others. It only means that at this very isolated moment in time, they clicked that particular one. It is true that each user will be exposed to the CTA(s) in different contexts and that this will contribute to their likelihood of clicking on it, perhaps independently of their "true" interest level in the product. However, if we do a reasonable job at random branch assignments then most of these issues should mainly just contribute noise (which will average out), rather than bias (i.e. if we enroll enough users we should have roughly the same amount of users who have been exposed to 10+ breach alerts in each branch).