Page MenuHomePhabricator

Intermittent connectivity issues in eqiad's row C
Closed, ResolvedPublic

Description

  • Spike of 503 at 2018-08-02T13:59 https://grafana.wikimedia.org/dashboard/db/varnish-http-errors?refresh=5m&panelId=7&fullscreen&orgId=1&from=1533217668620&to=1533219291066
  • connectivity issues between several hosts on asw2-b-eqiad T201039 elastic1049, elastic1038, starting 2018-07-31 15:00 (not intermittent, probably unrelated)
  • Before 2018-08-02T14:21:54 (not sure when, based on icinga monitoring), loss of connectivity from dbproxy1002 to db1065 (we believe without much proof to be a dbproxy1002 issue, because it didn't fail at the other proxy, dbproxy1008, but we are not sure)
  • Before 2018-08-03T01:53:32, (not sure when, based on icinga monitoring) loss of connectivity from dbproxy1001 and dbproxy1006 to db1063 (we believe to be a db1063 issue because both failed over).
    • At the same time (to be more precise - from 1:50:44 to 1:51:30), the etcd cluster for kubernetes suffered from high latencies. Looking at the logs, etcd1001 was unable to talk to either etcd1002 or etcd1003 at the time.

Event Timeline

Restricted Application added a subscriber: Aklapper. · View Herald Transcript

@Joe gave more timestamps from etcd logs on IRC:

  • Aug 2 13:59:52
  • Aug 3 01:19-01:20
  • Aug 3 01:28-01:29
  • Aug 3 01:50-01:51
  • Aug 3 02:06-02:07

These are potentially network partition hiccups/events. These seem to correlate well with the other events (dbproxy/db etc.) listed here.

They also seem to correlate well with DDOS_PROTOCOL_VIOLATION_SET events in cr1-eqiad:

Aug  2 13:59:49  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol NDPv6:aggregate is violated at fpc 3 for 1 times, started at 2018-08-02 13:59:48 UTC
Aug  2 13:59:56  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIMv6:aggregate is violated at fpc 3 for 2 times, started at 2018-08-02 13:59:56 UTC
Aug  2 14:00:03  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIM:aggregate is violated at fpc 3 for 1 times, started at 2018-08-02 14:00:03 UTC
Aug  2 14:05:52  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIM:aggregate has returned to normal. Violated at fpc 3 for 1 times, from 2018-08-02 14:00:03 UTC to 2018-08-02 14:00:51 UTC
Aug  2 14:05:52  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIMv6:aggregate has returned to normal. Violated at fpc 3 for 2 times, from 2018-08-02 13:59:56 UTC to 2018-08-02 14:00:51 UTC
Aug  2 14:05:52  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol NDPv6:aggregate has returned to normal. Violated at fpc 3 for 1 times, from 2018-08-02 13:59:48 UTC to 2018-08-02 14:00:51 UTC
Aug  2 14:14:57  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol Rejectv6:aggregate is violated at fpc 3 for 188 times, started at 2018-08-02 14:14:56 UTC
Aug  2 14:19:57  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol Rejectv6:aggregate has returned to normal. Violated at fpc 3 for 188 times, from 2018-08-02 14:14:56 UTC to 2018-08-02 14:14:56 UTC
Aug  3 01:18:37  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIM:aggregate is violated at fpc 3 for 2 times, started at 2018-08-03 01:18:36 UTC
Aug  3 01:18:47  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol NDPv6:aggregate is violated at fpc 3 for 2 times, started at 2018-08-03 01:18:46 UTC
Aug  3 01:23:37  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIM:aggregate has returned to normal. Violated at fpc 3 for 2 times, from 2018-08-03 01:18:36 UTC to 2018-08-03 01:18:36 UTC
Aug  3 01:25:07  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol NDPv6:aggregate has returned to normal. Violated at fpc 3 for 2 times, from 2018-08-03 01:18:46 UTC to 2018-08-03 01:20:06 UTC
Aug  3 01:25:12  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIM:aggregate is violated at fpc 3 for 3 times, started at 2018-08-03 01:25:11 UTC
Aug  3 01:25:12  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol NDPv6:aggregate is violated at fpc 3 for 3 times, started at 2018-08-03 01:25:11 UTC
Aug  3 01:31:12  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol NDPv6:aggregate has returned to normal. Violated at fpc 3 for 3 times, from 2018-08-03 01:25:11 UTC to 2018-08-03 01:26:11 UTC
Aug  3 01:34:17  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIM:aggregate has returned to normal. Violated at fpc 3 for 3 times, from 2018-08-03 01:25:11 UTC to 2018-08-03 01:29:16 UTC
Aug  3 01:51:22  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIM:aggregate is violated at fpc 3 for 4 times, started at 2018-08-03 01:51:21 UTC
Aug  3 01:51:22  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol NDPv6:aggregate is violated at fpc 3 for 4 times, started at 2018-08-03 01:51:21 UTC
Aug  3 01:56:42  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol NDPv6:aggregate has returned to normal. Violated at fpc 3 for 4 times, from 2018-08-03 01:51:21 UTC to 2018-08-03 01:51:41 UTC
Aug  3 02:00:47  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIM:aggregate has returned to normal. Violated at fpc 3 for 4 times, from 2018-08-03 01:51:21 UTC to 2018-08-03 01:55:46 UTC
Aug  3 02:04:47  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIM:aggregate is violated at fpc 3 for 5 times, started at 2018-08-03 02:04:46 UTC
Aug  3 02:04:47  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol NDPv6:aggregate is violated at fpc 3 for 5 times, started at 2018-08-03 02:04:46 UTC
Aug  3 02:06:42  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol PIMv6:aggregate is violated at fpc 3 for 3 times, started at 2018-08-03 02:06:41 UTC
Aug  3 02:09:52  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol NDPv6:aggregate has returned to normal. Violated at fpc 3 for 5 times, from 2018-08-03 02:04:46 UTC to 2018-08-03 02:04:46 UTC
Aug  3 02:12:37  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIM:aggregate has returned to normal. Violated at fpc 3 for 5 times, from 2018-08-03 02:04:46 UTC to 2018-08-03 02:07:36 UTC
Aug  3 02:12:37  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_CLEAR: Protocol PIMv6:aggregate has returned to normal. Violated at fpc 3 for 3 times, from 2018-08-03 02:06:41 UTC to 2018-08-03 02:07:36 UTC
Aug  3 02:24:46  re0.cr1-eqiad jddosd[1959]: DDOS_PROTOCOL_VIOLATION_SET: Protocol VRRP:aggregate is violated at fpc 3 for 16 times, started at 2018-08-03 02:24:46 UTC
[…]

Note that:

  • There are many more of these continuing up to ~now (Aug 3 10:17 UTC)
  • …but not (m)any before the first event logged here
  • cr2-eqiad doesn't seem to have the same events.
faidon renamed this task from Intermitent connectivity issues between eqiad servers? to Intermitent connectivity issues in eqiad's row C.Aug 3 2018, 10:35 AM
faidon triaged this task as Unbreak Now! priority.
faidon updated the task description. (Show Details)

The current guess is that those errors were side effects of the other switches failing.
"DDOS_PROTOCOL_VIOLATION" syslog seem to be read hearing.
Still monitoring, let me me know if the same issues happen so we can live troubleshot it.

New issue: there seems to be connectivity issues between es1014 (B1) and prometheus1004 (B4), not intermitent, they are unable to ping .

root@es1014:/run/mysqld$ ping prometheus1003.eqiad.wmnet
PING prometheus1003.eqiad.wmnet(prometheus1003.eqiad.wmnet (2620:0:861:101:10:64:0:123)) 56 data bytes
64 bytes from prometheus1003.eqiad.wmnet (2620:0:861:101:10:64:0:123): icmp_seq=1 ttl=63 time=0.220 ms
64 bytes from prometheus1003.eqiad.wmnet (2620:0:861:101:10:64:0:123): icmp_seq=2 ttl=63 time=0.240 ms
^C
--- prometheus1003.eqiad.wmnet ping statistics ---
2 packets transmitted, 2 received, 0% packet loss, time 1024ms
rtt min/avg/max/mdev = 0.220/0.230/0.240/0.010 ms

root@es1014:/run/mysqld$ ping prometheus1004.eqiad.wmnet
PING prometheus1004.eqiad.wmnet(prometheus1004.eqiad.wmnet (2620:0:861:102:10:64:16:38)) 56 data bytes
From 2620:0:861:102:46a8:42ff:fe35:72b8 (2620:0:861:102:46a8:42ff:fe35:72b8) icmp_seq=1 Destination unreachable: Address unreachable
From 2620:0:861:102:46a8:42ff:fe35:72b8 (2620:0:861:102:46a8:42ff:fe35:72b8) icmp_seq=2 Destination unreachable: Address unreachable
From 2620:0:861:102:46a8:42ff:fe35:72b8 (2620:0:861:102:46a8:42ff:fe35:72b8) icmp_seq=3 Destination unreachable: Address unreachable
^C
--- prometheus1004.eqiad.wmnet ping statistics ---
5 packets transmitted, 0 received, +3 errors, 100% packet loss, time 4056ms

root@prometheus1004:/srv/prometheus/ops/targets$ ping es1013.eqiad.wmnet
PING es1013.eqiad.wmnet (10.64.16.186) 56(84) bytes of data.
64 bytes from es1013.eqiad.wmnet (10.64.16.186): icmp_seq=1 ttl=64 time=0.139 ms
64 bytes from es1013.eqiad.wmnet (10.64.16.186): icmp_seq=2 ttl=64 time=0.154 ms
64 bytes from es1013.eqiad.wmnet (10.64.16.186): icmp_seq=3 ttl=64 time=0.229 ms
^C
--- es1013.eqiad.wmnet ping statistics ---
3 packets transmitted, 3 received, 0% packet loss, time 2053ms
rtt min/avg/max/mdev = 0.139/0.174/0.229/0.039 ms
root@prometheus1004:/srv/prometheus/ops/targets$ ping es1014.eqiad.wmnet
PING es1014.eqiad.wmnet (10.64.16.187) 56(84) bytes of data.
^C
--- es1014.eqiad.wmnet ping statistics ---
9 packets transmitted, 0 received, 100% packet loss, time 8187ms
ayounsi lowered the priority of this task from Unbreak Now! to Medium.Aug 23 2018, 11:34 PM

Lowering the priority, as as far I know this didn't happen again. Focusing on the row A/B issues.

Aklapper renamed this task from Intermitent connectivity issues in eqiad's row C to Intermittent connectivity issues in eqiad's row C.Aug 24 2018, 9:29 AM

Has anything happened on this? IIRC at our meetings we talked about investigating this further e.g. with the help of JTAC, and exploring whether we should disable the JunOS' DDoS protection.

I looked at it some time ago, the spike of DDOS_PROTOCOL_VIOLATION matches spikes of broadcast/multicast traffic we observed on asw2-a

jddosd-1m.png (1,315×751 px, 63 KB)

Spike of syslog messages from process jddos.

My guess is that asw2-a-eqiad failure triggered an internal loop/storm which hammed cr1-eqiad (not cr2 as asw2 was connected to cr1 and asw only).

This triggered (appropriately) the DDoS limiter of the router and dropped related packets, including legitimate packets to other switches as for most of the protocols, the rate limiting is done by FPC where several switches are connected.

Disabling DDOS_PROTOCOL_VIOLATION would put the RE at risk (eg. fully crash) if the same issue happen.

This is an issue as it means a faulty switch stack could potentially take down or impact other part of a site.

On top of my mind I can think of a few longer term action to mitigate the issue:

  • Different switch fabric design (eg. routing closer to the host)
  • Terminate each rows on a dedicated FPCs (so DDOS_PROTOCOL_VIOLATION doesn't impact other rows)
  • Enable STP (to prevent looping an access port)
  • Investigate DDOS_PROTOCOL_VIOLATION so see how much we can fine tune it

As an update to T201139#4483590 es1014 continues to show strange network patterns- I only see them at app layer, so take this with a grain of salt, but aside from continuing being unpollable from prometheus, I have relatively frequent connection errors to its master:

I leave them here in case the timestamps can be helpful to correlate with other events:

Sep 24 16:04:53 es1014 mysqld[3467]: 2018-09-24 16:04:53 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002657' at position 271877183; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38807149,180355214-180355214-24248112,180359340-180359
Sep 24 19:09:32 es1014 mysqld[3467]: 2018-09-24 19:09:32 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 24 19:09:32 es1014 mysqld[3467]: 2018-09-24 19:09:32 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002657' at position 804336515; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38818228,180355214-180355214-24367207,180359340-180359
Sep 24 20:16:15 es1014 mysqld[3467]: 2018-09-24 20:16:15 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 24 20:16:15 es1014 mysqld[3467]: 2018-09-24 20:16:15 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002657' at position 993768261; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38822231,180355214-180355214-24415099,180359340-180359
Sep 24 22:00:14 es1014 mysqld[3467]: 2018-09-24 22:00:14 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 24 22:00:14 es1014 mysqld[3467]: 2018-09-24 22:00:14 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002658' at position 228942087; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38828470,180355214-180355214-24478218,180359340-180359
Sep 24 23:22:07 es1014 mysqld[3467]: 2018-09-24 23:22:07 140234388567808 [Note] Slave: received end packet from server, apparent master shutdown:
Sep 24 23:22:07 es1014 mysqld[3467]: 2018-09-24 23:22:07 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002658' at position 428598733; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38833383,180355214-180355214-24520039,180359340-180359
Sep 25 05:20:56 es1014 mysqld[3467]: 2018-09-25  5:20:56 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 05:20:56 es1014 mysqld[3467]: 2018-09-25  5:20:56 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002659' at position 155654014; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38854912,180355214-180355214-24681677,180359340-180359
Sep 25 06:31:06 es1014 mysqld[3467]: 2018-09-25  6:31:06 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 06:31:06 es1014 mysqld[3467]: 2018-09-25  6:31:06 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002659' at position 285537002; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38859122,180355214-180355214-24712507,180359340-180359
Sep 25 10:51:38 es1014 mysqld[3467]: 2018-09-25 10:51:38 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 10:51:38 es1014 mysqld[3467]: 2018-09-25 10:51:38 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002659' at position 918297323; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38874754,180355214-180355214-24842815,180359340-180359
Sep 25 13:23:53 es1014 mysqld[3467]: 2018-09-25 13:23:53 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 13:23:53 es1014 mysqld[3467]: 2018-09-25 13:23:53 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002660' at position 238413833; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38883889,180355214-180355214-24925291,180359340-180359
Sep 25 15:08:59 es1014 mysqld[3467]: 2018-09-25 15:08:59 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 15:08:59 es1014 mysqld[3467]: 2018-09-25 15:08:59 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002660' at position 512853640; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38890195,180355214-180355214-24985878,180359340-180359
Sep 25 16:16:26 es1014 mysqld[3467]: 2018-09-25 16:16:26 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 16:16:26 es1014 mysqld[3467]: 2018-09-25 16:16:26 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002660' at position 681083513; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38894242,180355214-180355214-25023943,180359340-180359
Sep 25 18:07:19 es1014 mysqld[3467]: 2018-09-25 18:07:19 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 25 18:07:19 es1014 mysqld[3467]: 2018-09-25 18:07:19 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002660' at position 965573523; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38900895,180355214-180355214-25086989,180359340-180359
Sep 25 21:05:44 es1014 mysqld[3467]: 2018-09-25 21:05:44 140234388567808 [Note] Slave: received end packet from server, apparent master shutdown:
Sep 25 21:05:44 es1014 mysqld[3467]: 2018-09-25 21:05:44 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002661' at position 370905599; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38911600,180355214-180355214-25196666,180359340-180359
Sep 26 00:27:20 es1014 mysqld[3467]: 2018-09-26  0:27:20 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 00:27:20 es1014 mysqld[3467]: 2018-09-26  0:27:20 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002661' at position 808217748; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38923696,180355214-180355214-25312709,180359340-180359
Sep 26 01:54:14 es1014 mysqld[3467]: 2018-09-26  1:54:14 140234388567808 [Note] Slave: received end packet from server, apparent master shutdown:
Sep 26 01:54:14 es1014 mysqld[3467]: 2018-09-26  1:54:14 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002661' at position 990294225; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38928910,180355214-180355214-25360033,180359340-180359
Sep 26 03:26:05 es1014 mysqld[3467]: 2018-09-26  3:26:05 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 03:26:05 es1014 mysqld[3467]: 2018-09-26  3:26:05 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002662' at position 158370503; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38934421,180355214-180355214-25408644,180359340-180359
Sep 26 04:53:36 es1014 mysqld[3467]: 2018-09-26  4:53:36 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 04:53:36 es1014 mysqld[3467]: 2018-09-26  4:53:36 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002662' at position 359652383; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38939672,180355214-180355214-25453094,180359340-180359
Sep 26 07:16:44 es1014 mysqld[3467]: 2018-09-26  7:16:44 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 07:16:44 es1014 mysqld[3467]: 2018-09-26  7:16:44 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002662' at position 672440798; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38948260,180355214-180355214-25526798,180359340-180359
Sep 26 09:16:16 es1014 mysqld[3467]: 2018-09-26  9:16:16 140234388567808 [Note] Slave: received end packet from server, apparent master shutdown:
Sep 26 09:16:16 es1014 mysqld[3467]: 2018-09-26  9:16:16 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002662' at position 939207968; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38955432,180355214-180355214-25590971,180359340-180359
Sep 26 10:21:10 es1014 mysqld[3467]: 2018-09-26 10:21:10 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 10:21:10 es1014 mysqld[3467]: 2018-09-26 10:21:10 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002663' at position 52078423; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38959326,180355214-180355214-25626292,180359340-1803593
Sep 26 12:41:11 es1014 mysqld[3467]: 2018-09-26 12:41:11 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 12:41:11 es1014 mysqld[3467]: 2018-09-26 12:41:11 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002663' at position 412530486; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38967727,180355214-180355214-25707838,180359340-180359
Sep 26 16:47:55 es1014 mysqld[3467]: 2018-09-26 16:47:55 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 16:47:55 es1014 mysqld[3467]: 2018-09-26 16:47:55 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002664' at position 25268307; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38982531,180355214-180355214-25858367,180359340-1803593
Sep 26 20:56:11 es1014 mysqld[3467]: 2018-09-26 20:56:11 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 20:56:11 es1014 mysqld[3467]: 2018-09-26 20:56:11 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002664' at position 678268187; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-38997427,180355214-180355214-26005302,180359340-180359
Sep 26 22:37:01 es1014 mysqld[3467]: 2018-09-26 22:37:01 140234388567808 [ERROR] Error reading packet from server: Lost connection to MySQL server during query (server_errno=2013)
Sep 26 22:37:01 es1014 mysqld[3467]: 2018-09-26 22:37:01 140234388567808 [Note] Slave I/O thread: Failed reading log event, reconnecting to retry, log 'es1017-bin.002664' at position 927587744; GTID position '171970747-171970747-336240191,0-180359340-438187354,171974721-171974721-39003477,180355214-180355214-26061120,180359340-180359

Thanks for the update, note that es1014 is in row B (issues tracked in T201039)
Rough timeline is to get row B fixed next week, no ETA yet for row C (but I'm not aware of ongoing issues with row C).

There are 2 parallel issues here.

1/ IPv6 neighbor discovery randomly broken when igmp-snooping is enabled. This has been worked-around by disabling igmp-snooping yesterday T201039#4653040

2/ Non-juniper compliant Virtual Chassis Fabric cabling. Fixed for asw2-a/b-eqiad but not c yet. This has been stable until today.
The logs started to get flooded with:

Oct 10 18:12:09  asw2-c-eqiad fpc8 [EX-BCM PIC] ex_bcm_linkscan_handler: Link 56 UP
Oct 10 18:12:09  asw2-c-eqiad fpc8 [EX-BCM PIC] ex_bcm_pic_get_an_info: Failed to get the remote ability for Rear QSFP+ PIC port 2

Until:
Oct 10 18:20:13 asw2-c-eqiad chassisd[1837]: CHASSISD_IPC_CONNECTION_DROPPED: Dropped IPC connection for FPC 8

This caused at least a spike of 500
My intrepretation so far is that "link 56" which is the VC link between fpc3 and fpc8 failed in a way that it blackholed traffic (including traffic transiting through fpc8).
The switches went back to a stable state on their own.
I disabled that link to ensure that doesn't happen again.
fpc8 still have a link to fpc6 and fpc7 so it's safe to not re-enable fpc3-fpc8,

ayounsi mentioned this in Unknown Object (Task).Oct 10 2018, 7:39 PM

As far as I know this didn't reproduce since.

1/ has been solved by removing IGMP and 2/ by disabling the fpc3-fpc8 link