Closed Bug 1606003 Opened 6 years ago Closed 6 years ago

Review proposed Firefox Private Relay Market Test

Categories

(Data Science :: Experiment Collaboration, task)

task
Not set
normal

Tracking

(Not tracked)

RESOLVED INACTIVE

People

(Reporter: ssage, Assigned: loines, NeedInfo)

Details

Brief Description of the request (required):

In January, we plan to run an experiment on the Monitor website to determine the market interest in email masking (ie should we further explore the Firefox Private Relay project). Before we build anything here, we first aim to see if there is an identifiable user problem(s) that people resonate with.

Business purpose for this request (required):

To inform our holistic privacy & security bundle of apps & services as it relates to how/where/if email masking is included.

Requested timelines for the request or how this fits into roadmaps or critical decisions (required):

1/10/2019 - In theory, this review shouldn't involve much. If I'm misinterpreting the scope of this request, then we'll need to adjust this timeline.

As it relates to Monitor & Lockwise requests, this is probably the higher of priorities as we explore new opportunities on the Monitor website.

Links to any assets (e.g Start of a PHD, BRD; any document that helps describe the project):

Feature document: https://docs.google.com/document/d/1rJI8orhfnyW_QNAtAWtqd-PB3T-L02vZTq-Jpg4C6nM/edit#

Name of Data Scientist (If Applicable):

Leif

Please note if it is found that not enough information has been given this will delay the triage of this request.

Flags: needinfo?(loines)
Assignee: nobody → loines
Flags: needinfo?(loines)

Left some first thoughts in the doc via comments. The overall idea seems OK to me. To my mind, these are the big things to think about right now:

  1. we have not done an a/b test on the monitor site yet. typically, engineering would put in the plumbing for this, and we would verify it with an a/a test before doing a "real" experiment on which decisions hinge. do we need to do that here? if so, we can file another bug with DS to help with the design and analysis of that.

  2. the doc suggests we compare the click rate of an FPR CTA to that of similar lockwise and FPN CTAs. such a design would tell us, most directly, whether users are more, less or similarly willing to view information about FPR than they are to view information about lockwise and FPN. Let's say for the sake of argument that the FPR CTA does perform better than Lockwise or FPN. This would be consistent with any of the following (not mutually exclusive or exhaustive):
    a. Users are more interested in FPR than they are Lockwise or FPN
    b. The wording on the FPR CTA was more convincing/clear/easier to read than the other CTAs
    c. The FPR CTA is more novel than the Lockwise or FPN CTAs (especially a possibility if the user has already seen - and potentially clicked on - CTAs for those products on the monitor site or elsewhere)

My guess is that, in reality, we would want to conclude (a) while eliminating (b) and (c) to the best of our ability. There probably isn't a design that does this perfectly, but we would want to give it our best shot. We can rely on our content experts to reduce the chance of (b) as much as possible. My biggest concern is therefore (c):

We know that FPR will be totally new for all the users seeing that branch, but Lockwise and FPN may not be new to users in those branches. Therefore FPR may have an inherent advantage (or disadvantage?) due to novelty and that could be largely independent of how genuinely interesting the product is to users. I believe there is a good chance users seeing the Lockwise CTA may have already seen some promotion related to Lockwise on the monitor site. Maybe the problem is slightly less for FPN, I'm unsure if its promoted on monitor, but users may have been exposed via other channels - especially email if they are FxA (and the initial design calls for this to only be run on signed-in users, so that may increase the risk).

What could we do about this?

  1. only use FPN as the control - users are less likely to have been exposed to it already, and hey, it also starts with "Firefox Private...." in the title :P
  2. instead of calling out FPN and Lockwise by name in the CTA, make the wording more vague, e.g. "Learn how to manage your passwords better with Firefox" or "Learn how you can more securely browse the web using Firefox", you see my point.
  3. Other ideas?

I'd like for the FPR promotion to be on as even a footing with its comparison as possible.

(In reply to Leif Oines [:loines] from comment #1)

  1. we have not done an a/b test on the monitor site yet. typically, engineering would put in the plumbing for this, and we would verify it with an a/a test before doing a "real" experiment on which decisions hinge. do we need to do that here? if so, we can file another bug with DS to help with the design and analysis of that.

Filed https://github.com/mozilla/blurts-server/issues/1459 for this but don't know that we need that level of rigor for this experiment & decision?

Appreciate point #2 also. Sandy, :betsymi, :jdavidson - what should we do/build here?

Flags: needinfo?(ssage)
Flags: needinfo?(jdavidson)
Flags: needinfo?(bmikel)

I'm fine to do this test in this way — that is, to add a recommendation regarding email masking with a clickable CTA — but would like us to avoid pitting recommendations and various products/services against each other when we analyze the results. There are too many variables to isolate this as a test with a true control.

Users clicking on product A versus product B doesn't mean that A is a winner and B is a loser. Product A and Product B solve strive to solve entirely different privacy-related problems. Users may or may not have a level of understanding of what those privacy problems even are. Those users may or may not be interested in learning about one right now. We also cannot know is this the first time a user is seeing breach recommendations (perhaps this is their first breach), the 10th (perhaps they've received many breach alerts and have seen this page without this product countless times), or any number in between. And all privacy threats are not created equal. The fact that someone selects recommendation A doesn't mean they don't care about the others. It only means that at this very isolated moment in time, they clicked that particular one.

So with that context and nuance, I think seeing if people click recommendation A at all, without comparing it to recommendations B, C, D, etc., is the right way to go.

Flags: needinfo?(bmikel)

Thanks for the feedback Betsy - some replies inline:

(In reply to Betsy Mikel [:betsymi] from comment #3)

I'm fine to do this test in this way — that is, to add a recommendation regarding email masking with a clickable CTA — but would like us to avoid pitting recommendations and various products/services against each other when we analyze the results. There are too many variables to isolate this as a test with a true control.

Users clicking on product A versus product B doesn't mean that A is a winner and B is a loser. Product A and Product B solve strive to solve entirely different privacy-related problems. Users may or may not have a level of understanding of what those privacy problems even are. Those users may or may not be interested in learning about one right now.

I broadly agree - we should not think of more users clicking on CTA for "product A" as a slam dunk case that product A is more interesting than B, has better market fit, etc. No one experiment could answer these questions definitively, and if we do this test we need to follow up with other things e.g. user research, more experiments, prototypes etc regardless of results. That said, I do think we need some kind of comparison for reference and decision making. It doesn't necessarily have to be another product - It could be just an a priori number, e.g. "we will proceed with more research if 3+% of users exposed to the CTA click on it". But presumably this number would have some reason behind it. I don't want us to run the study without a point of comparison and then have to "read the entrails" wrt the results. Say 3% of users exposed click the CTA. Is that enough to warrant further work? Maybe, especially if we have (totally making up a number here) 1.5M visitors per day. But, without some point of reference, its hard to say. Maybe 3% of users would click on a banner that just says "click here" and nothing else (I've seen weirder).

Its true that FPN is a different solution for a different (but overlapping) set of problems. But, we've done research there, and have reasons to believe it has at least enough baseline interest for viability. That's what we're trying to get data on here (I think?). If we're careful about wording the CTAs to be as close to each other as possible (that's where we would really need your help), and the CTR for the FPR greatly outperforms FPN (a product that has already had some amount of vetting), I would think that this tells us something about its intrinsic interest level. What if FPR only performs half as well as FPN? I think that's also informative, but again not definitive. Once more, we're not trying to prove anything beyond a shadow of a doubt here, just trying to get one more shred of data to shore up our priors on what to do next.

I'm totally open to other ideas about what the comparison should be here, BTW. Not at all sold that it has to be FPN. But, I think we need something.

We also cannot know is this the first time a user is seeing breach recommendations (perhaps this is their first breach), the 10th (perhaps they've received many breach alerts and have seen this page without this product countless times), or any number in between. And all privacy threats are not created equal. The fact that someone selects recommendation A doesn't mean they don't care about the others. It only means that at this very isolated moment in time, they clicked that particular one.

It is true that each user will be exposed to the CTA(s) in different contexts and that this will contribute to their likelihood of clicking on it, perhaps independently of their "true" interest level in the product. However, if we do a reasonable job at random branch assignments then most of these issues should mainly just contribute noise (which will average out), rather than bias (i.e. if we enroll enough users we should have roughly the same amount of users who have been exposed to 10+ breach alerts in each branch).

Here's an interesting experiment :bmiroglio pointed me to, measuring the CTR for various firefox features.

I'm going to clear the needinfo for me. My recommendation is simply: Don't lead users to believe a product exists when it doesn't (and I've reviewed the landing page design, and I think we're ok), and make sure the landing page gets usability tested. I'm meeting with Sandy about usability testing the landing page next week.

Flags: needinfo?(jdavidson)

Work for the DS team is now tracked in Jira. You can search with the Data Science Jira project for the corresponding ticket.

Status: NEW → RESOLVED
Closed: 6 years ago
Resolution: --- → INACTIVE
You need to log in before you can comment on or make changes to this bug.