| -@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/irc-meetings] 1000378: Update opendev meeting location https://review.opendev.org/c/opendev/irc-meetings/+/1000378 | 07:26 | |
| @yaron.bar:matrix.org | Hi — my Gerrit account yaronbar (ID 39274, yaron.bar@weka.io) is still bound to a legacy Ubuntu One OpenID, https://login.ubuntu.com/+id/YBX3pXN. No contributor agreement is offered under Settings → Agreements, so I can't sign the ICLA or push. I have an OpenInfra Foundation account under <email>. Could you re-link the account to my OpenInfraID? I haven't logged in with the new identity, to avoid creating a duplicate. | 12:44 |
|---|---|---|
| @yaron.bar:matrix.org | Hi — my Gerrit account yaronbar (ID 39274, yaron.bar@weka.io) is still bound to a legacy Ubuntu One OpenID, https://login.ubuntu.com/+id/YBX3pXN. No contributor agreement is offered under Settings → Agreements, so I can't sign the ICLA or push. I have an OpenInfra Foundation account under yaron.bar@weka.io Could you re-link the account to my OpenInfraID? I haven't logged in with the new identity, to avoid creating a duplicate. | 12:45 |
| @fungicide:matrix.org | #status log Yanked PBR 7.1.1 and 7.1.2 from PyPI due to incompatibilities with old setuptools in easy-install invocations | 14:17 |
| @status:opendev.org | @fungicide:matrix.org: finished logging | 14:17 |
| @fungicide:matrix.org | Yaron Bar: who told you to sign a contributor agreement? if there's old documentation out there mentioning it, we want to know so we can try to get that cleaned up | 14:18 |
| @harbott.osism.tech:regio.chat | Yaron Bar: oh, I missed the "can't push" part earlier, can you explain what you want to push? are you aware of our documentation? https://docs.openstack.org/contributors/code-and-documentation/using-gerrit.html (that's the OpenStack variant but it should be similar for other opendev projects, too) | 14:22 |
| @harbott.osism.tech:regio.chat | mirror.dfw3.raxflex.opendev.org might be having issue, I also cannot ssh to it. | 17:50 |
| @clarkb:matrix.org | yup was just looking at that based on some pbr failures | 17:54 |
| @clarkb:matrix.org | I can't ssh to it nor can I list it | 17:54 |
| @clarkb:matrix.org | `HttpException: 401: Client Error for url: ... The request you have made requires authentication.` | 17:54 |
| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/zuul-providers] 1002365: Temporarily disable raxflex-dfw3 https://review.opendev.org/c/opendev/zuul-providers/+/1002365 | 17:55 | |
| @jim:acmegating.com | Clark: does that mean you can list nothing? | 17:55 |
| @clarkb:matrix.org | corvus: ya I have tried images and networks too and get the same error. catalog list works | 17:55 |
| @fungicide:matrix.org | so in theory zuul should stop booting instances there anyway | 17:55 |
| @clarkb:matrix.org | fungi: no we use two projects/accounts | 17:55 |
| @fungicide:matrix.org | aha, so it's just the one project and not the other? | 17:56 |
| @clarkb:matrix.org | its almost like they deleted our account (or at the very least changed our password under us) | 17:56 |
| @clarkb:matrix.org | fungi: yes | 17:56 |
| @fungicide:matrix.org | ouch | 17:56 |
| -@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [opendev/zuul-providers] 1002365: Temporarily disable raxflex-dfw3 https://review.opendev.org/c/opendev/zuul-providers/+/1002365 | 17:57 | |
| @clarkb:matrix.org | I didn't see any obvious email warnings | 17:58 |
| @clarkb:matrix.org | but ya either our account is no longer valid or our account password is not valid. Its possible we have to login via the horizon setup to get things to tie back to old keystone again? | 17:59 |
| @clarkb:matrix.org | fungi: I think you have done that in the past and there isn't anything special just log in right? | 17:59 |
| @clarkb:matrix.org | I'm going to test other regions next | 17:59 |
| @jim:acmegating.com | i think the urls are all in the file | 18:00 |
| @fungicide:matrix.org | looks like it's affecting sjc3 too, merely non-impacting since we're still not using it | 17:59 |
| @clarkb:matrix.org | the IAD3 mirror is up | 18:00 |
| @clarkb:matrix.org | so they probably didn't delete the entire project | 18:00 |
| @clarkb:matrix.org | but I cannot list servers in IAD3 | 18:01 |
| @fungicide:matrix.org | yeah, i get a 401 in all 3 regions for that account | 18:01 |
| @clarkb:matrix.org | so maybe we need to login to not horizon and see if that fixes things | 18:01 |
| @clarkb:matrix.org | should I do that or did someone else want to? | 18:02 |
| @fungicide:matrix.org | i'm trying | 18:02 |
| @fungicide:matrix.org | just taking me a few minutes | 18:02 |
| @clarkb:matrix.org | thanks. I need to prep for our meeting soon too so that is helpful | 18:03 |
| @fungicide:matrix.org | the account seems to be gone even from skyline. i can log in as `openstackjenkins@rackspace_cloud_domain` but not `openstackci@rackspace_cloud_domain` | 18:08 |
| @fungicide:matrix.org | Doug Goldstein: ^ since you seemed to be following along in irc | 18:08 |
| @fungicide:matrix.org | so it definitely seems like one of our accounts got clobbered and not the other | 18:09 |
| @fungicide:matrix.org | openstackci can still log into the manage.rackspace.com site at least, but we have no open tickets there yet | 18:12 |
| @fungicide:matrix.org | okay, once i logged into the rackspace portal, there was a link i could click there to log into skyline via sso rather than entering credentials directly into skyline, and that seems to have worked | 18:16 |
| @clarkb:matrix.org | fungi: can you see the state of the dfw3 mirror node that way? | 18:16 |
| @fungicide:matrix.org | yeah, i'm trying to work that out now. it definitely lists the instance | 18:17 |
| @fungicide:matrix.org | ah, status error | 18:17 |
| @fungicide:matrix.org | so the problem with our account credentials may be a longer-term issue and not something that just started, since we don't log into this one often | 18:18 |
| @fungicide:matrix.org | "libvirtError" | 18:19 |
| @fungicide:matrix.org | that's not good | 18:19 |
| @fungicide:matrix.org | so i think the immediate problem is that some catastrophic failure has befallen the instance | 18:19 |
| @fungicide:matrix.org | that we can't work out how to authenticate to the api may be a completely unrelated issue | 18:19 |
| @clarkb:matrix.org | ya could be two orthogonal issues. | 18:21 |
| @fungicide:matrix.org | anyway, it very well may be a down host or something, and they just haven't opened a ticket about it yet | 18:24 |
| @fungicide:matrix.org | for now, i've opened ticket #260825-ord-0001652 with rackspace to notify them of the outage | 18:40 |
| @clarkb:matrix.org | fungi: did you mention the inability to login to skyline directly or use the API? | 18:46 |
| @clarkb:matrix.org | Wondering if we need a second ticket for that | 18:46 |
| @fungicide:matrix.org | i did not yet, want to see what they come back with there first, and then do some more troubleshooting with our login info | 18:47 |
| @fungicide:matrix.org | since i can get into skyline another way i should be able to export some credentials for comparison on our own | 18:48 |
| @fungicide:matrix.org | it's possible we'll be able to identify what changed without involving them right away (or at all) | 18:48 |
| @fungicide:matrix.org | in fact, i wonder if something wasn't recorded correctly when we rotated the password for this account in january, and we simply haven't used it since then? | 18:52 |
| @fungicide:matrix.org | because comparing the openrc.sh file everything matches what we put in our clouds.yaml, with the only thing i can't verify being the password | 18:53 |
| @fungicide:matrix.org | (since the exported openrc.sh omits that intentionally) | 18:53 |
| @clarkb:matrix.org | looking at bridge that was the last time any of the credentials updated | 18:55 |
| @clarkb:matrix.org | so nothing has chagned since then. Which amkes that theory somewhat reasonable. I thought we were good about testing as we went. I will check my own personally history on bridge to see if I ran commands more recentlythan that | 18:55 |
| @clarkb:matrix.org | fungi: on April 13th I ran a listing against SJC3 servers and the mirror node there. I assume that is when we remoevd that region because the mirror had a sad. | 18:57 |
| @clarkb:matrix.org | fungi: https://meetings.opendev.org/irclogs/%23opendev/%23opendev.2026-04-13.log.html#opendev.2026-04-13.log.html#t2026-04-13T16:58:23 this implies I had working API access on that date | 18:57 |
| @fungicide:matrix.org | so in theory it was working as recently as 5 months ago, at least | 18:58 |
| @clarkb:matrix.org | yes I think this was working in April | 18:58 |
| @fungicide:matrix.org | i can also "update user password" through skyline, though it wants to know the original password | 18:58 |
| @fungicide:matrix.org | oh, though we're using the cloud api keys from our legacy accounts as passwords, according to our notes | 19:00 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/irc-meetings] 1002385: Remove extra # from matrix meeting locations https://review.opendev.org/c/opendev/irc-meetings/+/1002385 | 19:49 | |
| @jim:acmegating.com | oops ^ :) | 19:49 |
| @clarkb:matrix.org | Eric Ball: I left a note on the python3.14 change you pushed asking if we can split things up. Not sure if you saw that. Mostly I want to make reverts easy if one specific container doesn't like upgrading | 19:49 |
| @clarkb:matrix.org | Anil Belur: I responded to your review on https://review.opendev.org/c/opendev/system-config/+/1002179 (thank you for looking by the way) and you are correct that is a deficiency in that ansible role. If you'd like to work on fixing it I noted another file that should be checked for updates in my response. Otherwise let me know and I can work on an update | 19:50 |
| @clarkb:matrix.org | Anil Belur: the gerrit replication will work as long as ssh is up but if the gitea web service isn't also running and ready then gitea never updates the database and its like we never replicated according to the web ui. But the content is in git so gerrit never pushes it again automatically. That is why we have that complicated restart procedure to ensure gitea is ready for the ssh container to accept replication pushes | 19:51 |
| @fungicide:matrix.org | from skyline i can tell that the flex sjc3 mirror is also in error state, though the error detail is different | 19:52 |
| @fungicide:matrix.org | "Guest does not have a console available." | 19:52 |
| @clarkb:matrix.org | ya based on the irc logs I pulled up it was in an error state before too but with no reason given in the show output over the api | 19:53 |
| @fungicide:matrix.org | i've opened ticket #260825-ord-0001702 about the sjc3 mirror | 19:59 |
| @fungicide:matrix.org | rebooting the mirror01.dfw3.raxflex instance to see if apache will start normally or get more debug logging | 20:14 |
| @fungicide:matrix.org | it's possible it was waiting for openafs | 20:15 |
| @fungicide:matrix.org | apache still hasn't started | 20:16 |
| @fungicide:matrix.org | i can browse `/afs/openstack.org/` on it though | 20:16 |
| @fungicide:matrix.org | the cinder volume we're using for caches is also attached and logical volumes from it mounted | 20:17 |
| @fungicide:matrix.org | `systemctl status apache2` claims the service is disabled | 20:18 |
| @clarkb:matrix.org | That's weird. Our config management shouldn't do that (and may explicitly enable it) | 20:20 |
| @clarkb:matrix.org | Will systemd disable units that fill in a loop automatically? | 20:20 |
| @clarkb:matrix.org | * Will systemd disable units that fail in a loop automatically? | 20:20 |
| @fungicide:matrix.org | possible | 20:21 |
| @fungicide:matrix.org | `apache2ctl configtest` reports "Syntax OK" | 20:22 |
| @fungicide:matrix.org | i enabled it and then started it | 20:22 |
| @clarkb:matrix.org | I wonder if afs broke and put it into a restart loop and systemd gave up on it | 20:23 |
| @fungicide:matrix.org | seems to be working, i'm going to reboot again and see what happens | 20:23 |
| @fungicide:matrix.org | maybe we have a startup race or something | 20:23 |
| @clarkb:matrix.org | Google says systemd does not disable failing units | 20:23 |
| @fungicide:matrix.org | came back up fine with apache running after a reboot now | 20:27 |
| @fungicide:matrix.org | it looks like root ran `sudo systemctl disable apache2` (note the superfluous use of `sudo` there) according to its shell history, so i'll see if i can find a record in the auth log of when that happened | 20:30 |
| @fungicide:matrix.org | it was prior to 2026-07-26 which is as far back as auth.log history goes | 20:31 |
| @jim:acmegating.com | that totally sounds like me but i can't remember working on a mirror for some time... maybe it was when we were under some ddos | 20:31 |
| @fungicide:matrix.org | so i think someone was troubleshooting something, disabled the apache2 unit manually but didn't stop the service, then didn't start again automatically later at reboot | 20:32 |
| @fungicide:matrix.org | anyway, mystery (more or less) solved | 20:33 |
| @jim:acmegating.com | was that the server where i was tuning apache in real time for the ddos? | 20:33 |
| @fungicide:matrix.org | might have been, yes | 20:33 |
| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/zuul-providers] 1002390: Revert "Temporarily disable raxflex-dfw3" https://review.opendev.org/c/opendev/zuul-providers/+/1002390 | 20:36 | |
| @jim:acmegating.com | fungi: yep i think it was from the emergency work that lead to this: https://review.opendev.org/c/opendev/system-config/+/995978 | 20:38 |
| @jim:acmegating.com | sorry! | 20:38 |
| @fungicide:matrix.org | no worries, it didn't take long at all to figure out. could easily have been any of us | 20:39 |
| @clarkb:matrix.org | I just approved the revert change | 20:51 |
| @clarkb:matrix.org | (after checking the web server on that mirror works for me) | 20:51 |
| -@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [opendev/zuul-providers] 1002390: Revert "Temporarily disable raxflex-dfw3" https://review.opendev.org/c/opendev/zuul-providers/+/1002390 | 20:51 | |
| @abelur:matrix.org | Clark: I can fix up the ansible role for your change in https://review.opendev.org/c/opendev/system-config/+/1002179, i'll prep up the change and ping you once its ready or if I need some help / info on this. thanks | 22:58 |
| @abelur:matrix.org | Eric Ball: on 996553 (python3.14 images) - Clark has asked for it to be split per service group, and I've already got that done: 5 commits matching his suggested grouping (hound / haproxy-statsd+zk-statsd / bots / gerrit / base images). All 7 images build locally on 3.14. | 23:00 |
| No point stacking mine on yours - I checked, and 4 of my 5 commits cherry-pick empty onto 996553. Same change written twice, 15 of 16 | ||
| files identical. So it's just a question of which tree the split lands from. | ||
| @abelur:matrix.org | * Eric Ball: on 996553 (python3.14 images) - Clark has asked for it to be split per service group, and I've already got that done: 5 commits matching his suggested grouping (hound / haproxy-statsd+zk-statsd / bots / gerrit / base images). All 7 images build locally on 3.14. | 23:00 |
| No point stacking mine on yours - I checked, and 4 of my 5 commits cherry-pick empty onto 996553. Same change written twice, 15 of 16 files identical. So it's just a question of which tree the split lands from. | ||
| @abelur:matrix.org | 23:00 | |
| Eric Ball Ok if I push mine, with a commit carrying your base-image ARG bump so nothing of yours is lost, and credit 996553 in the messages? Happy to hand you the commits instead if you'd rather push it yourself. | ||
| One thing mine adds either way: asyncio.run() in the eavesdrop bot - get_event_loop() raises on 3.14 and the container won't start without it. | ||
| @clarkb:matrix.org | > <@abelur:matrix.org> Clark: I can fix up the ansible role for your change in https://review.opendev.org/c/opendev/system-config/+/1002179, i'll prep up the change and ping you once its ready or if I need some help / info on this. thanks | 23:03 |
| I think we should have two changes. One to improve restarts for gitea when files change and the other that updates the anubis config (this second one is what I started before getting side tracked by PBR) | ||
| @clarkb:matrix.org | So feel free to push a new change for that and I can rebase | 23:03 |
| @abelur:matrix.org | perfect, I was about to ask on this! :) | 23:03 |
| @eball-lf:matrix.org | Go ahead and push what you've got. | 23:09 |
| @abelur:matrix.org | thanks - I'll update here once its ready | 23:16 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!