Closed Bug 720489 Opened 14 years ago Closed 11 years ago

Telemetry view of current vs last release

Categories

(Mozilla Metrics :: Frontend Reports, defect)

defect
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED
Moved to JIRA

People

(Reporter: lmandel, Unassigned)

Details

(Whiteboard: [JIRA BIV-8] [Telemetry])

Attachments

(3 files)

It would be useful to compare the current state of Firefox (via the most current Telemetry data) vs the last release. It may make sense to compare one 6 week period with the next 6 week period rather than previous Firefox release day to the current day. I'm open to suggestions here. The overall goal is to provide a view that allows for easy comparison of how the lastest beta/aurora/nightly is doing in relation to the previous release as it comes to data sets such as memory consumption, start-up time, cycle collection time, etc.
Would this be the equivalent of going to the evolution dashboard and overlaying, eg, nightly and aurora?
I don't know that simply overlaying the 6 week development periods would be useful. I'm not as interested in comparing day 1 of Nightly with day 1 of Aurora as I am in comparing the end result of our time on Nightly with the current state of Aurora or the current state of Firefox 10 beta with the Firefox 9 release. Do you have any other suggestions for how we can compare releases/builds?
Hey there. At the moment we are a bit full with the current requests, but we'll try to fit this in the our next goals. Please keep updating this bug with your ideas regarding the report and we'll contact you in a couple of weeks in order to piece this together. Would that be ok with you?
Yes. That works.
Whiteboard: [Telemetry] → [Telemetry:P2]
Assignee: nobody → paulo.pires
Group: metrics-private
Attached image Telemetry Comparison_01
Attached image Telemetry Comparison_02
Attached image Telemetry Comparison_03
Hi, attached you will find the mockup dashboard. Looking forward to your feedback. Joana
This is what I had in mind with creating a custom dashboard in bug 720485. I don't see how these mockups will help me compare releases. Can you provide some more details? TBH, evolution with the release flags may be enough at the moment to compare releases and close this bug. In the future I would like to see how we're trending release to release (% improvement). It may be useful to be able to state the Firefox E starts 5% faster than Firefox D, 7% faster than Firefox C, and 10% faster than Firefox B.
You lost me, why is this related to bug 720485? The goal of this dashboard is exactly what you described. Maybe that wasn't clear, but the selectors on the left, that refer to the platform build ids for each channen, will automagically populate with the platform build id that corresponds to the latest code. Like you said initially, and if I got it right,, you want to compare the most recent codes for each channel. Telemetry wouldn't help there. The table we see in the screenshots is to help you draw the conclusions you aim for: beta x% better than release, etc
instead of "Telemetry wouldn't help there.", i meant "Telemetry evolution wouldn't help there"
(In reply to Pedro Alves from comment #9) > You lost me, why is this related to bug 720485? > > > The goal of this dashboard is exactly what you described. Maybe that wasn't > clear, but the selectors on the left, that refer to the platform build ids > for each channen, will automagically populate with the platform build id > that corresponds to the latest code. I completely missed that as I was focused on the "duplicate" functionality in the dashboard. This is the piece that I thought was similar to bug 720485 as I could create a custom dashboard with multiple histograms. > > > Like you said initially, and if I got it right,, you want to compare the > most recent codes for each channel. Telemetry wouldn't help there. You're right. I should have read my comment 2. That is what this bug is about and this view does look to satisfy my request and answer my question. At this point we have mockups but no code changes, correct? If so, let me ping a few other people to get their views before we get into the code changes.
Ah, got it. Yes, the duplicate feature is something we have already and plan to keep using it :-) This is just a mockup, so by all means lets discuss it here as much as possible. Only when everyone's confortable with it we'll start imementing it
I mentioned this in https://bugzilla.mozilla.org/show_bug.cgi?id=733468 but I think that a much better measure of how we're doing on a given metric is the % of telemetry pings above a certain cutoff rather than the mean, median or standard deviation. Is there a way we can get that into the summary data below the graph?
Feedback from mbest: Maybe worth considering graphing the mean change or perhaps the change relative to the current version. It's really the metric you care about most at least at a glance. The main graph let's you dive into detail which is awesome but the first thing you want to know is generally where things sit and then move into the minutia. It's easier to get a sense of the degree of chance at a glance when graphed rather than via just displaying a number. Numerical table below does a good job of flagging that there is an issue and this maybe enough so I'm raising it as a possible improvement but the teams using it are the best judge of how useful this could be.
Feedback from dmandelin: I'm glad to see better comparison features coming. I totally agree with Martin. The way I put it: side-by-side comparisons on 31 x values is overwhelming to the viewer. I can see that ultimately you'd want to be able to spot patterns like one version having more top outlier, but I'm not sure that the view shown is even correct for that. - repeat, Martin's idea about a summary graph of means is good. Would be good to have confidence intervals and/or quantiles there too. - On the graph itself: - I think it would be much easier to understand if the y values were scaled. Example: right now the thing that pops out the most is that the light blue bars are highest in the middle, but this is actually meaningless because it's just a reflection of release having more users. - Lines might be easier to read than bars. - Does Tofte have some examples of showing distributions for comparison like this? Some detailed comments: - Channels should be ordered like Release, Beta, Aurora, Nightly (or the reverse) - ID is hard to read. - Splitting it out into date components would help. - Better yet, show the version number and hover for the ID - Some actual statistician can come in and correct me if I'm wrong, but I think sample stdev by itself is not a very useful statistic. Standard error is the more useful one, but hard to interpret unless you are familiar with it. I think quantiles is probably better. - Would be very nice to show t test results for significance of the difference of means. Need to tell people to take it with a grain of salt, though, because data look non-normal. - Not sure what m/m means.
Status: NEW → ASSIGNED
Migrating to JIRA to manage this work. These tickets are replaced by METRICS-775.
Status: ASSIGNED → RESOLVED
Closed: 14 years ago
Resolution: --- → INCOMPLETE
updated whiteboard field
Whiteboard: [Telemetry:P2] → Telemetry P2 migrated to JIRA
updating status and assignee
Assignee: paulo.pires → nobody
Status: RESOLVED → REOPENED
Resolution: INCOMPLETE → ---
New JIRA migration identifier - 55 BZ issues being considered in current planning for Q3
Target Milestone: Unreviewed → Targeted - JIRA
Whiteboard: Telemetry P2 migrated to JIRA → [JIRA BIV-8] Telemetry
Whiteboard: [JIRA BIV-8] Telemetry → [JIRA BIV-8] [Telemetry]
Telemetry metrics are being rethought out as part of the Executive Metrics project.
Status: ASSIGNED → RESOLVED
Closed: 14 years ago11 years ago
Resolution: --- → FIXED
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Creator:
Created:
Updated:
Size: