As mentioned in T313095 , we lost a master due to hardware failure, and then a second master during a routine reimage operation.
While there was no user impact, we believe that increasing the number of Elastic master-eligible nodes from 3 to 5 will reduce the likelihood of these types of failures without significantly affecting performance or operational complexity.
We should also choose newer chassis as our master eligibles, as many of the current master-eligibles are slated for a hardware refresh