User Details
- User Since
- Oct 3 2014, 8:06 AM (615 w, 3 d)
- Availability
- Available
- IRC Nick
- godog
- LDAP User
- Filippo Giunchedi
- MediaWiki User
- FGiunchedi (WMF) [ Global Accounts ]
Yesterday
Cluster in warning also blocks ceph reboots for {T431659}, not sure if we want to --force though ?
Good point @dcaro, ok thank you! So as far as Toolforge is concerned we can move to bind-mounting /mnt/nfs with mount propagation, I'll send out patches
Thu, Jul 16
Wed, Jul 15
Thank you @Odinaldo for reaching out and the useful information. I don't think 60GB is problematic, I'm looking forward to the size estimation for the full database. Also in terms of sizing I'd like to know if you have a sense of system load in terms of requests per second coming from users? thank you
A side effect of the toolsbeta deployment from yesterday was T432115: [jobs-emailer, toolsbeta] No emails sent in the last hour i.e. pods getting stuck while starting. toolforge volume-admission expects dumps mountpoint directories to be present, therefore we'll have to make sure to not remove them via puppet.
Tue, Jul 14
I'm going through https://wikitech.wikimedia.org/wiki/Server_Lifecycle#Move_existing_server_between_rows/racks,_changing_IPs and figuring out ownership for the various steps, cc @cmooney @ayounsi
@VRiley-WMF I thought we could start with one host, cloudvirt1048 and test-run the whole procedure there. Once we have nailed the process then we can do the rest in batch, what do you think?
As of now dumps_use_nfs_lb: true is deployed to toolsbeta, unfortunately for k8s workers that means a roll-reboot because containers keep a reference to /mnt/nfs mountpoints. That in turn means that pre-reboot the mount attempts will keep contacting the old clouddumps nfs server address. A reboot is the easiest way to ensure new mounts hit dumps-nfs lb address.
I forgot to attach a task number, however:
Mon, Jul 13
During the last reboots of clouddumps hosts we have not observed toolforge/cloudvps unavailability. At any rate, the service will be improving when T411248: Plan to make clouddumps more resilient and easier to operate is complete
This is still an issue: SRE got paged for 10 5xx req/s on swift which was doing 2k req/s at the time (esams). I very much doubt it is worth paging engineers on absolute number of errors, irrespective of service traffic
Fri, Jul 10
Agreed re: 'sshd is quite reliable enough'.
ok two paths as I see it:
This is done in the sense that maintenance aggregate is no longer a thing. I have updated https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Maintenance#cloudvirtXXXX with commands to list maint status
Thu, Jul 9
Sweet, thank you !
Wed, Jul 8
And another one for your eyes @Andrew https://gitlab.wikimedia.org/repos/sre/wikitech-static-docker/-/merge_requests/6
+1 LGTM
Tentatively resolving, feel free to reopen if sth is amiss
Tue, Jul 7
I did some more testing of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308126 (general scaffolding to support dumps-nfs.w.o) and https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308128 (cloud vps support).
Ok I spent some time debugging the script (https://script.google.com/home/projects/1k0oMpC0CtdiwLN9zCbvFM_g2z5LtdtMDfV4QMJWLgMi7gljqfJAJb0iI) and it is definitely doing something, even recently:
Mon, Jul 6
I'm not sure there's anything immediate to do here as trying again worked. I suspect the proper and more time-intensive fix is to move away from wmcs-drain-hypervisor and onto cookbooks, either using cli or openstacksdk clients.
+1 to just get rid of renaming/broken functionality
Current situation wrt aggregates ceph, network-ovs and maintenance
Fri, Jul 3
The tool availability metric will need to be a recording rule because istio_requests_total is huge. And the same per-tool availability metric can also feed the KR metrics for webservice unavailability we were talking about yesterday (cc @CCiufo-WMF @aputhin)
Indeed now with the work by @taavi on T392356: Replace ingress-nginx before upstream EOL date we do have per-tool status codes as seen by istio, i.e. the metric used in https://grafana.wmcloud.org/d/fnhp8st/tool-error-rates based on istio_requests_total
Could be yeah, I don't know if puppet agent is actually supposed to be doing its own crl management/refresh.
On balance I like option 2, which ensures truly end-to-end monitoring of container build + deploy. In other words the wikitech_static_update_timestamp_seconds metric going stale (time() - wikitech_static_update_timestamp_seconds) means something went wrong, either the container failed to build or failed to deploy. Either way we can take a look when/if that happens, @Andrew I sent the MR your way, please let me know what you think
On the specifics on why exactly this happens, speaking as a former o11y member, I'm not going to look deeper into icinga-issued pages because IMHO we shouldn't be doing that anymore in the first place.
To provide some context/history: the "re-page on acked but not resolved incidents" is a VO setting which we set to 24h and can be disabled (cfr T259465: VictorOps behavior on long-ack'd incidents). I am +1 on changing the behavior to not re-page
I looked at, and reported, logs as part of clinic duty though did not touch the db FWIW
Thu, Jul 2
Ok so I spent some time getting a better understanding of NFS and how Linux clients reacts to a server failover. Below my findings:
All steps done, I'm tentatively resolving though @Rscout please reach out and reopen if something is amiss
Wed, Jul 1
@Rsilvola FYI this is pending your approval
Tue, Jun 30
I'll be starting the trials on toolsbeta and specifically toolsbeta-test-k8s-worker-nfs-8.toolsbeta by overriding its hiera values with the following:
hosts are in service. I have updated https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Maintenance#new_cloudvirt_install with the instructions, manual for now
https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306672 is the bandaid I propose for now to get unblocked
@Rscout we need to verify your ssh key out of band, please let me know when it would be a good time for a quick google meet. feel also free to send a meeting invite my way: fgiunchedi@wikimedia.org
@Milimetric @Ahoelzl @Ottomata I'm seeking analytics-privatedata-users approval for Mona, an former WMDE intern. thank you !
Thank you all!
Mon, Jun 29
Would we have capacity (power, space) to move two hosts to their final allocation in E4/F4 ?
@KFrancis I could not find an NDA on file for Mona Thierse, would you mind arranging one? thank you so much!
Hello @Monrac5, thank you for reaching out -- just to confirm: you are not part of WMDE staff, correct ?