Page MenuHomePhabricator

fgiunchedi (Filippo Giunchedi)
/* No comment */

Today

  • No visible events.

Tomorrow

  • No visible events.

Thursday

  • No visible events.

User Details

User Since
Oct 3 2014, 8:06 AM (615 w, 3 d)
Availability
Available
IRC Nick
godog
LDAP User
Filippo Giunchedi
MediaWiki User
FGiunchedi (WMF) [ Global Accounts ]

Recent Activity

Yesterday

fgiunchedi added a comment to T429387: cloudceph HEALTH_WARN, multiple OSD(s) experiencing slow operations in BlueStore.

Cluster in warning also blocks ceph reboots for {T431659}, not sure if we want to --force though ?

Mon, Jul 20, 12:11 PM · tools-infrastructure-team, Cloud-VPS
fgiunchedi updated the task description for T432587: Migrate to dumps-nfs.w.o in Cloud VPS.
Mon, Jul 20, 11:49 AM · tools-infrastructure-team, Cloud-VPS
fgiunchedi updated the task description for T432212: Migrate to dumps-nfs.w.o in production.
Mon, Jul 20, 11:30 AM · tools-infrastructure-team, Data-Platform-SRE
fgiunchedi created T432587: Migrate to dumps-nfs.w.o in Cloud VPS.
Mon, Jul 20, 10:32 AM · tools-infrastructure-team, Cloud-VPS
fgiunchedi created T432583: Make sure dumps-nfs mounts are propagated inside PAWS containers.
Mon, Jul 20, 10:22 AM · PAWS, tools-platform-team, tools-infrastructure-team
fgiunchedi renamed T432325: Make sure dumps-nfs mount/umount is propagated inside Toolforge containers from Make sure dumps-nfs mount/umount is propagated inside containers to Make sure dumps-nfs mount/umount is propagated inside Toolforge containers.
Mon, Jul 20, 10:19 AM · tools-platform-team, tools-infrastructure-team, Cloud-VPS
fgiunchedi added a comment to T432325: Make sure dumps-nfs mount/umount is propagated inside Toolforge containers.

Good point @dcaro, ok thank you! So as far as Toolforge is concerned we can move to bind-mounting /mnt/nfs with mount propagation, I'll send out patches

Mon, Jul 20, 10:18 AM · tools-platform-team, tools-infrastructure-team, Cloud-VPS
fgiunchedi added a comment to T430933: clean up nova IDs for newer cloudvirts.

I think I'm behind on something. When you say...

should not be hardcoded in puppet anymore

Are you talking about the profile::openstack::base::nova::compute::id values in hiera, or a different nova ID somewhere else in puppet?

Mon, Jul 20, 8:57 AM · tools-infrastructure-team, Cloud-VPS

Thu, Jul 16

fgiunchedi created T432325: Make sure dumps-nfs mount/umount is propagated inside Toolforge containers.
Thu, Jul 16, 9:33 AM · tools-platform-team, tools-infrastructure-team, Cloud-VPS
fgiunchedi removed a subtask for T411248: Plan to make clouddumps more resilient and easier to operate: T430651: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020.
Thu, Jul 16, 6:47 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi removed a parent task for T430651: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020: T411248: Plan to make clouddumps more resilient and easier to operate.
Thu, Jul 16, 6:47 AM · tools-infrastructure-team, Infrastructure-Foundations, Traffic, netops

Wed, Jul 15

fgiunchedi added a comment to T400675: Page on ATS backend errors relative to traffic.

Thank you for the feedback @hnowlan @ssingh ! re: exclusion list I'm also expecting that moving to ratio-based paging will carry more signal (i.e. we can shrink the exclusion list)

Wed, Jul 15, 4:00 PM · SRE Observability, SRE-SLO, Traffic, SRE
fgiunchedi added a comment to T416590: Advice for Hybrid Architecture for AI Tool.

Thank you @Odinaldo for reaching out and the useful information. I don't think 60GB is problematic, I'm looking forward to the size estimation for the full database. Also in terms of sizing I'd like to know if you have a sense of system load in terms of requests per second coming from users? thank you

Wed, Jul 15, 3:43 PM · Cloud-VPS, tools-infrastructure-team
fgiunchedi created T432212: Migrate to dumps-nfs.w.o in production.
Wed, Jul 15, 9:41 AM · tools-infrastructure-team, Data-Platform-SRE
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

A side effect of the toolsbeta deployment from yesterday was T432115: [jobs-emailer, toolsbeta] No emails sent in the last hour i.e. pods getting stuck while starting. toolforge volume-admission expects dumps mountpoint directories to be present, therefore we'll have to make sure to not remove them via puppet.

Wed, Jul 15, 6:38 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS

Tue, Jul 14

fgiunchedi added a comment to T431682: Rebalance cloudvirts out of E4.

Thanks @fgiunchedi! I will be running through this today. I just got back from vacation and catching up on a few different things here. I will let you know when this first one is done.

Tue, Jul 14, 2:03 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi updated subscribers of T431682: Rebalance cloudvirts out of E4.

I'm going through https://wikitech.wikimedia.org/wiki/Server_Lifecycle#Move_existing_server_between_rows/racks,_changing_IPs and figuring out ownership for the various steps, cc @cmooney @ayounsi

Tue, Jul 14, 1:56 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi added a comment to T431682: Rebalance cloudvirts out of E4.

@VRiley-WMF I thought we could start with one host, cloudvirt1048 and test-run the whole procedure there. Once we have nailed the process then we can do the rest in batch, what do you think?

Tue, Jul 14, 12:24 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi created P94814 (An Untitled Masterwork).
Tue, Jul 14, 10:08 AM
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

As of now dumps_use_nfs_lb: true is deployed to toolsbeta, unfortunately for k8s workers that means a roll-reboot because containers keep a reference to /mnt/nfs mountpoints. That in turn means that pre-reboot the mount attempts will keep contacting the old clouddumps nfs server address. A reboot is the easiest way to ensure new mounts hit dumps-nfs lb address.

Tue, Jul 14, 9:05 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

I forgot to attach a task number, however:

Tue, Jul 14, 7:45 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi created T432093: Consider per-rack availability-zones in nova.
Tue, Jul 14, 7:40 AM · tools-infrastructure-team, Cloud-VPS

Mon, Jul 13

fgiunchedi closed T391369: If the inactive clouddumps host goes down, it causes a ripple effect on Cloud VPS and Toolforge, a subtask of T403154: Upgrade clouddumps hosts to bookworm/trixie, as Invalid.
Mon, Jul 13, 3:23 PM · tools-infrastructure-team, Data-Platform-SRE, Cloud-VPS, cloud-services-team
fgiunchedi closed T391369: If the inactive clouddumps host goes down, it causes a ripple effect on Cloud VPS and Toolforge as Invalid.

During the last reboots of clouddumps hosts we have not observed toolforge/cloudvps unavailability. At any rate, the service will be improving when T411248: Plan to make clouddumps more resilient and easier to operate is complete

Mon, Jul 13, 3:23 PM · tools-infrastructure-team, Toolforge, Cloud-VPS
fgiunchedi created T432012: Get rid of neutronclient deprecation warnings.
Mon, Jul 13, 1:50 PM · tools-infrastructure-team, Cloud-VPS
fgiunchedi added a comment to T400675: Page on ATS backend errors relative to traffic.

This is still an issue: SRE got paged for 10 5xx req/s on swift which was doing 2k req/s at the time (esams). I very much doubt it is worth paging engineers on absolute number of errors, irrespective of service traffic

Mon, Jul 13, 7:56 AM · SRE Observability, SRE-SLO, Traffic, SRE

Fri, Jul 10

fgiunchedi added a comment to T429013: Upgrade cloudsw1-e4-eqiad.

Thanks @fgiunchedi.

This isn't so urgent we want to cause stress for you guys. So 2 or 3 is also fine if you want to move some of those hosts first. But 1 obviously works for us too :)

Fri, Jul 10, 2:49 PM · tools-infrastructure-team, Cloud-VPS, netops, Infrastructure-Foundations, SRE
fgiunchedi added a comment to T395448: Discuss about "host down" semantics.

Agreed re: 'sshd is quite reliable enough'.

Fri, Jul 10, 2:47 PM · SRE Observability (FY2025/2026-Q1), Observability-Metrics
fgiunchedi added a comment to T429013: Upgrade cloudsw1-e4-eqiad.

ok two paths as I see it:

Fri, Jul 10, 1:40 PM · tools-infrastructure-team, Cloud-VPS, netops, Infrastructure-Foundations, SRE
fgiunchedi created T431829: dashes are allowed when creating a new action, but not when saving.
Fri, Jul 10, 1:19 PM · Hiddenparma
fgiunchedi closed T424802: Revisit cloudvirt maintenance story as Resolved.

This is done in the sense that maintenance aggregate is no longer a thing. I have updated https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Maintenance#cloudvirtXXXX with commands to list maint status

Fri, Jul 10, 8:34 AM · tools-infrastructure-team, Cloud-VPS, cloud-services-team

Thu, Jul 9

fgiunchedi updated the task description for T431682: Rebalance cloudvirts out of E4.
Thu, Jul 9, 12:25 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi raised the priority of T284747: openstack: alert for cloudvirts without aggregate or with unexpected set of them from Low to Medium.
Thu, Jul 9, 10:46 AM · tools-infrastructure-team, Cloud-VPS
fgiunchedi edited projects for T284747: openstack: alert for cloudvirts without aggregate or with unexpected set of them, added: tools-infrastructure-team; removed cloud-services-team.
Thu, Jul 9, 10:35 AM · tools-infrastructure-team, Cloud-VPS
fgiunchedi updated the task description for T284747: openstack: alert for cloudvirts without aggregate or with unexpected set of them.
Thu, Jul 9, 10:31 AM · tools-infrastructure-team, Cloud-VPS
fgiunchedi updated the task description for T431682: Rebalance cloudvirts out of E4.
Thu, Jul 9, 10:05 AM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi added a comment to T424658: Ensure cloudvirt capacity is more evenly spread out among racks.

@fgiunchedi I was curious to know if there is an estimated date for these moves? Will we be doing it one server at a time.

Thu, Jul 9, 10:03 AM · cloud-services-team (Hardware), tools-infrastructure-team, SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi created T431682: Rebalance cloudvirts out of E4.
Thu, Jul 9, 10:01 AM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi added a comment to T431259: 'cluster overview' collapsible rows panel show up with incorrect width.

Sweet, thank you !

Thu, Jul 9, 6:45 AM · SRE Observability, Grafana

Wed, Jul 8

fgiunchedi added a comment to T362397: Move wikitech-static monitoring off Icinga.

And another one for your eyes @Andrew https://gitlab.wikimedia.org/repos/sre/wikitech-static-docker/-/merge_requests/6

Wed, Jul 8, 12:18 PM · cloud-services-team, wikitech.wikimedia.org, Observability-Alerting
fgiunchedi added a comment to T431531: Request creation of wdp-gitlab-runner VPS project.

+1 LGTM

Wed, Jul 8, 10:05 AM · User-dcaro, Wikidata Platform Team, Cloud-VPS (Project-requests)
fgiunchedi closed T431303: Problems with SRE Team Vacations Calendar sync as Resolved.

Tentatively resolving, feel free to reopen if sth is amiss

Wed, Jul 8, 8:54 AM · SRE
fgiunchedi placed T430933: clean up nova IDs for newer cloudvirts up for grabs.
Wed, Jul 8, 7:36 AM · tools-infrastructure-team, Cloud-VPS

Tue, Jul 7

fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

I did some more testing of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308126 (general scaffolding to support dumps-nfs.w.o) and https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308128 (cloud vps support).

Tue, Jul 7, 3:32 PM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi updated the task description for T411248: Plan to make clouddumps more resilient and easier to operate.
Tue, Jul 7, 3:25 PM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

@fgiunchedi maybe silly question but to make sure we're preserving public IPs, why can't we use the same IPs for dumps and nfs dumps? Like we do for other services listening on different ports (eg. gerrit-http/gerrit-ssh)?

Not a silly question at all -- please double check my reasoning, service is not in production yet (more on that later) and there's still time to change with basically-zero impact.

  • The service is fundamentally different than http/rsync in my mind, it is public only because cloud/prod both access it. Unlike http/rsync which are public by design
  • Easier to match/isolate the load balanced nfs traffic from http/rsync

Having said that, I don't feel very strongly nor have stronger motivations, maybe some other folks do

Thanks, then let's be frugal about IPs and use the same one for both protocols if that's ok with you.

Using ports + IPs seems fine enough to match/isolate traffic (ACLs, analytics).

SGTM, I'll be moving the service to the same IP as http/rsync

Tue, Jul 7, 12:39 PM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi added a comment to T431303: Problems with SRE Team Vacations Calendar sync.

I've just deleted and re-created my OoO starting 21 July. Thanks for looking into this :)

Tue, Jul 7, 9:38 AM · SRE
fgiunchedi added a comment to T431303: Problems with SRE Team Vacations Calendar sync.

Ok I spent some time debugging the script (https://script.google.com/home/projects/1k0oMpC0CtdiwLN9zCbvFM_g2z5LtdtMDfV4QMJWLgMi7gljqfJAJb0iI) and it is definitely doing something, even recently:

Tue, Jul 7, 8:11 AM · SRE

Mon, Jul 6

fgiunchedi created T431307: Consider deprecating the 'suspend' VM feature.
Mon, Jul 6, 2:11 PM · tools-infrastructure-team, Cloud-VPS
fgiunchedi updated the task description for T431300: NeutronAgentDown firing during cloudvirt.safe_reboot despite silences.
Mon, Jul 6, 1:38 PM · Cloud-VPS, tools-infrastructure-team
fgiunchedi added projects to T431300: NeutronAgentDown firing during cloudvirt.safe_reboot despite silences: Cloud-VPS, cloud-services-team.
Mon, Jul 6, 1:34 PM · Cloud-VPS, tools-infrastructure-team
fgiunchedi created T431300: NeutronAgentDown firing during cloudvirt.safe_reboot despite silences.
Mon, Jul 6, 1:34 PM · Cloud-VPS, tools-infrastructure-team
fgiunchedi added a comment to T431282: nova.exception.GroupAffinityViolation while live-migrating tools-redis-6.

I'm not sure there's anything immediate to do here as trying again worked. I suspect the proper and more time-intensive fix is to move away from wmcs-drain-hypervisor and onto cookbooks, either using cli or openstacksdk clients.

Mon, Jul 6, 12:36 PM · Toolforge, Cloud-VPS, tools-infrastructure-team, cloud-services-team
fgiunchedi updated the task description for T431282: nova.exception.GroupAffinityViolation while live-migrating tools-redis-6.
Mon, Jul 6, 12:16 PM · Toolforge, Cloud-VPS, tools-infrastructure-team, cloud-services-team
fgiunchedi created T431282: nova.exception.GroupAffinityViolation while live-migrating tools-redis-6.
Mon, Jul 6, 11:58 AM · Toolforge, Cloud-VPS, tools-infrastructure-team, cloud-services-team
fgiunchedi created T431259: 'cluster overview' collapsible rows panel show up with incorrect width.
Mon, Jul 6, 9:41 AM · SRE Observability, Grafana
fgiunchedi added a comment to T429738: [infra,o11y] ToolforgeWebHighErrorRate should not page if a single tool is down.

Moving the alerting to be based on Istio-level metrics means we lose visibility into any potential connectivity issues between HAProxy and Istio, right?

Mon, Jul 6, 8:36 AM · tools-platform-team, Toolforge
fgiunchedi added a comment to T431092: Web proxy service does not check domain policies when changing the domain of an existing proxy.

+1 to just get rid of renaming/broken functionality

Mon, Jul 6, 8:24 AM · tools-infrastructure-team, Cloud-VPS, Security
fgiunchedi renamed T424802: Revisit cloudvirt maintenance story from cloudvirt1075 in 'maintenance' aggregate to Revisit cloudvirt maintenance story.
Mon, Jul 6, 8:10 AM · tools-infrastructure-team, Cloud-VPS, cloud-services-team
fgiunchedi added a comment to T424802: Revisit cloudvirt maintenance story.

Current situation wrt aggregates ceph, network-ovs and maintenance

Mon, Jul 6, 8:03 AM · tools-infrastructure-team, Cloud-VPS, cloud-services-team
fgiunchedi added a comment to T424802: Revisit cloudvirt maintenance story.

I'm not sure why this is happening. Most likely it's due to a different cookbook (maybe wmcs.openstack.cloudvirt.safe_reboot) failing and leaving things in an inconsistent state.

As far as I know the 'maintenance' aggregate doesn't do anything specific, it was just created as in indicator to human eyeballs that a host is getting worked on. The real action is in the 'ceph' aggregate which allows new VMs to be scheduled. So we could potentially do away with the maintenance aggregate entirely.

Ok now I'm even more confused, the set_maintenance and unset_maintenance cookbooks operate on maintenance aggregate to supposedly stop/start VMs being scheduled on the hypervisor.

Mon, Jul 6, 7:48 AM · tools-infrastructure-team, Cloud-VPS, cloud-services-team

Fri, Jul 3

fgiunchedi updated subscribers of T429738: [infra,o11y] ToolforgeWebHighErrorRate should not page if a single tool is down.

The tool availability metric will need to be a recording rule because istio_requests_total is huge. And the same per-tool availability metric can also feed the KR metrics for webservice unavailability we were talking about yesterday (cc @CCiufo-WMF @aputhin)

Fri, Jul 3, 9:56 AM · tools-platform-team, Toolforge
fgiunchedi updated subscribers of T429738: [infra,o11y] ToolforgeWebHighErrorRate should not page if a single tool is down.

Indeed now with the work by @taavi on T392356: Replace ingress-nginx before upstream EOL date we do have per-tool status codes as seen by istio, i.e. the metric used in https://grafana.wmcloud.org/d/fnhp8st/tool-error-rates based on istio_requests_total

Fri, Jul 3, 9:51 AM · tools-platform-team, Toolforge
fgiunchedi added a comment to T429298: puppet failing on metricsinfra-prometheus-2.metricsinfra due to expired crl.

Could be yeah, I don't know if puppet agent is actually supposed to be doing its own crl management/refresh.

Fri, Jul 3, 9:04 AM · tools-infrastructure-team, cloud-services-team, Cloud-VPS
fgiunchedi updated subscribers of T362397: Move wikitech-static monitoring off Icinga.

On balance I like option 2, which ensures truly end-to-end monitoring of container build + deploy. In other words the wikitech_static_update_timestamp_seconds metric going stale (time() - wikitech_static_update_timestamp_seconds) means something went wrong, either the container failed to build or failed to deploy. Either way we can take a look when/if that happens, @Andrew I sent the MR your way, please let me know what you think

Fri, Jul 3, 8:28 AM · cloud-services-team, wikitech.wikimedia.org, Observability-Alerting
fgiunchedi added a comment to T431030: Splunk Oncall paged at 5:30am for nothing.

On the specifics on why exactly this happens, speaking as a former o11y member, I'm not going to look deeper into icinga-issued pages because IMHO we shouldn't be doing that anymore in the first place.

Fri, Jul 3, 7:09 AM · Observability-Alerting
fgiunchedi added a comment to T431030: Splunk Oncall paged at 5:30am for nothing.

To provide some context/history: the "re-page on acked but not resolved incidents" is a VO setting which we set to 24h and can be disabled (cfr T259465: VictorOps behavior on long-ack'd incidents). I am +1 on changing the behavior to not re-page

Fri, Jul 3, 7:05 AM · Observability-Alerting
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

@fgiunchedi maybe silly question but to make sure we're preserving public IPs, why can't we use the same IPs for dumps and nfs dumps? Like we do for other services listening on different ports (eg. gerrit-http/gerrit-ssh)?

Not a silly question at all -- please double check my reasoning, service is not in production yet (more on that later) and there's still time to change with basically-zero impact.

  • The service is fundamentally different than http/rsync in my mind, it is public only because cloud/prod both access it. Unlike http/rsync which are public by design
  • Easier to match/isolate the load balanced nfs traffic from http/rsync

Having said that, I don't feel very strongly nor have stronger motivations, maybe some other folks do

Thanks, then let's be frugal about IPs and use the same one for both protocols if that's ok with you.

Using ports + IPs seems fine enough to match/isolate traffic (ACLs, analytics).

Fri, Jul 3, 6:50 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi added a comment to T429759: Cannot join Toolforge: Striker throws IntegrityError while saving user.

I looked at, and reported, logs as part of clinic duty though did not touch the db FWIW

Fri, Jul 3, 6:40 AM · User-bd808, tools-platform-team, Striker, cloud-services-team

Thu, Jul 2

fgiunchedi created T430933: clean up nova IDs for newer cloudvirts.
Thu, Jul 2, 11:11 AM · tools-infrastructure-team, Cloud-VPS
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

Ok so I spent some time getting a better understanding of NFS and how Linux clients reacts to a server failover. Below my findings:

Thu, Jul 2, 10:18 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi closed T430594: Requesting access to deployment for rscout as Resolved.

All steps done, I'm tentatively resolving though @Rscout please reach out and reopen if something is amiss

Thu, Jul 2, 8:12 AM · SRE, SRE-Access-Requests
fgiunchedi updated the task description for T430594: Requesting access to deployment for rscout.
Thu, Jul 2, 8:11 AM · SRE, SRE-Access-Requests

Wed, Jul 1

fgiunchedi added a comment to T430594: Requesting access to deployment for rscout.

@Rsilvola FYI this is pending your approval

Wed, Jul 1, 3:07 PM · SRE, SRE-Access-Requests
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

@fgiunchedi maybe silly question but to make sure we're preserving public IPs, why can't we use the same IPs for dumps and nfs dumps? Like we do for other services listening on different ports (eg. gerrit-http/gerrit-ssh)?

Wed, Jul 1, 1:38 PM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi closed T430304: Requesting access to "analytics-privatedata" for mona_thierse as Resolved.

Hi all, the NDA is complete. Thanks!

Wed, Jul 1, 12:05 PM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi updated the task description for T430304: Requesting access to "analytics-privatedata" for mona_thierse.
Wed, Jul 1, 12:03 PM · Data-Engineering, SRE, SRE-Access-Requests

Tue, Jun 30

fgiunchedi updated the task description for T430594: Requesting access to deployment for rscout.
Tue, Jun 30, 2:40 PM · SRE, SRE-Access-Requests
fgiunchedi added a comment to T430304: Requesting access to "analytics-privatedata" for mona_thierse.

Thank you all!

@Monrac5 we'd need to verify your ssh public key out of band. please let me know when it would be a good time for doing so synchronously (e.g. google meet)

@fgiunchedi would tomorrow at 11:30 am CEST work for you? i could also meet anytime before that or after 1 pm CEST

Tue, Jun 30, 2:39 PM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi added a comment to T411248: Plan to make clouddumps more resilient and easier to operate.

I'll be starting the trials on toolsbeta and specifically toolsbeta-test-k8s-worker-nfs-8.toolsbeta by overriding its hiera values with the following:

Tue, Jun 30, 1:09 PM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi updated the task description for T411248: Plan to make clouddumps more resilient and easier to operate.
Tue, Jun 30, 12:59 PM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi closed T429563: Put cloudvirt10[77-80] in service, a subtask of T424658: Ensure cloudvirt capacity is more evenly spread out among racks, as Resolved.
Tue, Jun 30, 12:52 PM · tools-infrastructure-team, cloud-services-team (Hardware), SRE, DC-Ops, ops-eqiad, Cloud-VPS
fgiunchedi closed T429563: Put cloudvirt10[77-80] in service as Resolved.

hosts are in service. I have updated https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Maintenance#new_cloudvirt_install with the instructions, manual for now

Tue, Jun 30, 12:52 PM · tools-infrastructure-team, Cloud-VPS, cloud-services-team (FY2025/2026-Q3-Q4)
fgiunchedi updated the task description for T430304: Requesting access to "analytics-privatedata" for mona_thierse.
Tue, Jun 30, 12:05 PM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi added a comment to T430304: Requesting access to "analytics-privatedata" for mona_thierse.

Approved (NOTE: I initially thought there was no NDA on file, but see from the comments it's now available, so you can update that in the task description)

Tue, Jun 30, 12:03 PM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi added a comment to T430651: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020.

https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306672 is the bandaid I propose for now to get unblocked

Tue, Jun 30, 11:57 AM · tools-infrastructure-team, Infrastructure-Foundations, Traffic, netops
fgiunchedi created T430651: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020.
Tue, Jun 30, 11:56 AM · tools-infrastructure-team, Infrastructure-Foundations, Traffic, netops
fgiunchedi updated the task description for T411248: Plan to make clouddumps more resilient and easier to operate.
Tue, Jun 30, 9:12 AM · Patch-For-Review, tools-platform-team, tools-infrastructure-team, Infrastructure-Foundations, Data-Platform-SRE, Traffic, netops, Data-Services, Cloud-VPS
fgiunchedi added a comment to T430594: Requesting access to deployment for rscout.

@Rscout we need to verify your ssh key out of band, please let me know when it would be a good time for a quick google meet. feel also free to send a meeting invite my way: fgiunchedi@wikimedia.org

Tue, Jun 30, 8:03 AM · SRE, SRE-Access-Requests
fgiunchedi updated the task description for T430594: Requesting access to deployment for rscout.
Tue, Jun 30, 7:59 AM · SRE, SRE-Access-Requests
fgiunchedi added a project to T430304: Requesting access to "analytics-privatedata" for mona_thierse: Data-Engineering.

@Milimetric @Ahoelzl @Ottomata I'm seeking analytics-privatedata-users approval for Mona, an former WMDE intern. thank you !

Tue, Jun 30, 7:43 AM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi added a comment to T430304: Requesting access to "analytics-privatedata" for mona_thierse.

Thank you all!

Tue, Jun 30, 7:42 AM · Data-Engineering, SRE, SRE-Access-Requests

Mon, Jun 29

fgiunchedi updated the task description for T430310: Provision cloudvirt with firewall.
Mon, Jun 29, 1:14 PM · Cloud-VPS, tools-infrastructure-team, cloud-services-team
fgiunchedi added a comment to T419892: Q3:rack/setup/install cloudcephosd105[3456].

Would we have capacity (power, space) to move two hosts to their final allocation in E4/F4 ?

Mon, Jun 29, 10:15 AM · SRE, cloud-services-team (Hardware), ops-eqiad, DC-Ops
fgiunchedi updated the task description for T429563: Put cloudvirt10[77-80] in service.
Mon, Jun 29, 7:44 AM · tools-infrastructure-team, Cloud-VPS, cloud-services-team (FY2025/2026-Q3-Q4)
fgiunchedi updated subscribers of T430304: Requesting access to "analytics-privatedata" for mona_thierse.

@KFrancis I could not find an NDA on file for Mona Thierse, would you mind arranging one? thank you so much!

Mon, Jun 29, 7:19 AM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi added a comment to T430304: Requesting access to "analytics-privatedata" for mona_thierse.

Hello @Monrac5, thank you for reaching out -- just to confirm: you are not part of WMDE staff, correct ?

Mon, Jun 29, 7:17 AM · Data-Engineering, SRE, SRE-Access-Requests
fgiunchedi updated the task description for T430304: Requesting access to "analytics-privatedata" for mona_thierse.
Mon, Jun 29, 7:14 AM · Data-Engineering, SRE, SRE-Access-Requests

Fri, Jun 26

fgiunchedi updated the task description for T429451: Establish a blackbox network probe vantage point into cloud realm.
Fri, Jun 26, 1:51 PM · Infrastructure-Foundations, netops, tools-infrastructure-team, Cloud-VPS
fgiunchedi created T430310: Provision cloudvirt with firewall.
Fri, Jun 26, 1:46 PM · Cloud-VPS, tools-infrastructure-team, cloud-services-team