Closed Bug 1594849 Opened 6 years ago Closed 6 years ago

Use anonymized IP data to understand retention edge cases

Categories

(Data Science :: Investigation, task, P1)

x86_64
Unspecified
task
Points:
3

Tracking

(Not tracked)

RESOLVED INACTIVE

People

(Reporter: harter, Assigned: harter)

References

Details

There are a few common edge-cases related to temporary profiles that consistently come up when discussing retention. I'm interested in getting a rough understanding of how prominent these issues are using anonymized IP address data.

Initial investigations could include:

  • What proportion of clients come from IP addresses producing many clients? E.g. is there an obvious bot that's churning thousands of profiles from a single IP address?
  • How long do most clients keep an IP address? Are there any IPs that consistently send fresh client_ids? E.g. are there IPs that are obviously churning client_ids daily?

Felix has some prior art in this space here.

Priority: -- → P1
Depends on: 1595136
Status: NEW → ASSIGNED
Group: mozilla-employee-confidential

Romain has some background on retention issues in this document (in particular here).

Quick update for the record:

A quick analysis of the blocklist ping suggests that ~10% of our our ADI came from Amazon ISPs (source.

ADI does not include a client_id so it's very possible that this actually represents much fewer than 10% of DAU. We can get a more accurate read on the size of this problem once we get ISP associated with anonymized client_ids. This work to refresh this dataset is being tracked in Bug 1595136.

A quick review of the existing IP data from ~2 years ago showed very few clients come from "chatty IPs". That is, very few clients came from IP addresses associated with many clients. This suggests the library/computer lab issue is small. Before putting more effort into this analysis it makes sense to refresh the dataset (again Bug 1595136) so we're working with more contemporary data.

Work for the DS team is now tracked in Jira. You can search with the Data Science Jira project for the corresponding ticket.

Status: ASSIGNED → RESOLVED
Closed: 6 years ago
Resolution: --- → INACTIVE
You need to log in before you can comment on or make changes to this bug.