Page MenuHomePhabricator

Impression weirdness in 2020-12-07 en6C desktop large test
Open, Needs TriagePublic

Description

Hi @AndyRussG and @Pcoombe; Peter, we've talked a bit about this, but I'm finally getting around to creating a task. Any help understanding this issue would be appreciated!

On December 7, 2020, PCoombe put up an a/b test in the large banner buckets (A&B) of our FY2021 en6C desktop CentralNotice campaign.

We've since realized that one of the banners in that test, B2021_120716_en6C_dsk_p1_lg_dsn_squareCorners, registered about a million more impressions than the other banner, B2021_120716_en6C_dsk_p1_lg_dsn_cnt.

Here's a Turnilo chart illustrating the issue; blue line = _squareCorners.

image.png (1,417×916 px, 106 KB)

I've looked at the logs for that campaign and haven't caught any human error (e.g., someone leaving _squareCorners in a bucket when setting up the next test).

Finally and interestingly, donations did not continue to accumulate for _squareCorners, just impressions, making it look like an enormous loss from a testing perspective. Note the discrepancy in trials (impressions).

The obvious anxiety is that impressions could be unevenly distributed in more/most tests. I've done some spot-checking of other tests from that time period, and recently, but haven't caught other alarming impression imbalances.

Event Timeline

@EYener and @RMurthy is this related to your investigations about lower performing banners in December? Have you looked at this and/or do you have any data? I just want to be sure we have all available info.

I did a search for any banners from the 6C campaigns where the test banner and the control banner had a pgehres impression count difference >5%. This was the only case.

https://docs.google.com/spreadsheets/d/1NvueAtix7jBn4BI1S4KokooB45j4nhRIU4D1vG6o9N4/edit?usp=sharing

(the B2021_122914_en6C_m_p2_sm_dow_fixbtntxt* banners in the sheet are explained by the fact that we re-ran one of them against the control in a later test)

Sorry I'm late here @DStrine -- I think we had been talking about portal? I don't recall this particular issue, and it sounds like Peter found the isolated case where it happened. Thanks!

Hi all! @DStrine : The weirdness you're referring to was isolated to mobile devices only. We saw a slight increase in mobile impressions and a simultaneous slight decrease in mobile donations, leading to an overall notable decrease in mobile donation rate. However, it doesn't appear there was a technical bug....that we've identified, yet.

As @Pcoombe notes, this was an isolated incident to this test - and, actually, isolated within this test as well. Here is a Turnilo link broken down by hour and variant, where you can see that the test and control diverge at hour 19:00. Until then, the two tests are in sync.

Given the time of the divergence, is there something in setup / takedown that could be identified? Of course, people cache banners, too. Is there a way to check the IP addresses coming through post-19:00? I don't believe there is, but want to check with FR-Tech first. If there is not, I sort of doubt we would be able to determine a root cause...

I don't know about IP addresses, but I did find that the difference in impressions (turnilo) and the persistence of impressions after the banner supposedly came down (turnilo) only happened in the US and not in other countries.