| @priteau:matrix.org | Hello. Are there known connectivity issues in VEXXHOST? | 11:05 |
|---|---|---|
| @priteau:matrix.org | WARNING: Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ConnectTimeoutError(<HTTPSConnection(host='mirror.ca-ymq-1.vexxhost.opendev.org', port=443) at 0x7991b7ee3230>, 'Connection to mirror.ca-ymq-1.vexxhost.opendev.org timed out. (connect timeout=60.0)')': /pypi/simple/python-openstackclient/ | 11:05 |
| @priteau:matrix.org | https://zuul.opendev.org/t/openstack/build/6d196476377b4520904856f97aa1845a | 11:08 |
| @priteau:matrix.org | Similar issues on other jobs | 12:05 |
| @priteau:matrix.org | Err:1 https://mirror.ca-ymq-1.vexxhost.opendev.org/ubuntu noble InRelease | 12:05 |
| Could not connect to mirror.ca-ymq-1.vexxhost.opendev.org:443 (2604:e100:1:0:f816:3eff:fef9:1d9c), connection timed out Could not connect to mirror.ca-ymq-1.vexxhost.opendev.org:443 (162.253.55.232), connection timed out | ||
| @fungicide:matrix.org | Pierre Riteau: we saw reports of possible internal routing problems in that region for resources we could reach there from outside | 13:59 |
| @fungicide:matrix.org | if it seems to be hitting a lot of builds, we can temporarily stop booting nodes there, though i think some specific node labels may only be satisfied in that region so we could be trading for node_failure results on those | 14:01 |
| @jim:acmegating.com | #status log started new ze11 | 14:10 |
| @status:opendev.org | @jim:acmegating.com: finished logging | 14:10 |
| @jim:acmegating.com | is there a job that typically failes with the connection error? we could put in an autohold | 14:11 |
| @gthiemonge:matrix.org | looks like most of the octavia jobs are affected, for instance: https://review.opendev.org/c/openstack/octavia/+/919846?tab=change-view-tab-header-zuul-results-summary | 14:13 |
| @fungicide:matrix.org | i wonder if they're all using a high-memory label only available there | 14:15 |
| @jim:acmegating.com | gthiemonge: is one of those really reliable otherwise? i'm looking for a case where a failure *probably* means we're hitting the network issue | 14:15 |
| @gthiemonge:matrix.org | yes they are, for instance octavia-v2-dsvm-scenario-traffic-ops | 14:16 |
| @gthiemonge:matrix.org | https://zuul.opendev.org/t/openstack/builds?job_name=octavia-v2-dsvm-scenario-traffic-ops&project=openstack/octavia | 14:17 |
| @gthiemonge:matrix.org | the 6 most recent failures are due to issues with the ubuntu mirrors | 14:17 |
| @jim:acmegating.com | gthiemonge: https://zuul.opendev.org/t/openstack/autohold/0000000303 is in place -- if you want to trigger a change that will run that job, we'll hold the node from the next 2 failures | 14:18 |
| @gthiemonge:matrix.org | ack, I've just rechecked the change that I linked | 14:19 |
| @gthiemonge:matrix.org | this job uses nested-virt-ubuntu-noble, it looks like it's almost only vexxhost-ca-ymq-1 | 14:21 |
| @jim:acmegating.com | gthiemonge: https://zuul.opendev.org/t/openstack/build/d48512286a844a729cdf0844083a51ae is a held failure; did that hit the problem? | 14:49 |
| @gthiemonge:matrix.org | corvus: not this time :/ it failed due to an issue with pip (we've had a few issues with the pip mirrors too) | 14:50 |
| @jim:acmegating.com | it looks like it was able to communicate with the mirror... so i guess it's intermittent. | 14:52 |
| @gthiemonge:matrix.org | the first one I rechecked is still running, so it didn't hit the issue | 14:52 |
| @jim:acmegating.com | i'm going to leave that hold in place, so that if we hit the issue later, we'll have a control node handy and we can check to see if this node changes behavior | 14:52 |
| @jim:acmegating.com | the autohold is set up to catch one more error node | 14:53 |
| @priteau:matrix.org | Should we stop rechecking for now? | 15:05 |
| @fungicide:matrix.org | my suspicion is that there's a hypervisor host in that region that has broken routing or similar, so when job nodes get scheduled to that host the build can't reach the mirror in that region | 15:28 |
| @clarkb:matrix.org | ya I brought this theory up yesterday but at the time we thought the issue was a 100% consistent failure. This new data would give more credence to this theory though | 15:40 |
| @jim:acmegating.com | oh whoops, it looks like we typically move /opt to /var/lib/zuul on executors | 16:07 |
| @jim:acmegating.com | i'll stop ze11 and do that | 16:07 |
| @jim:acmegating.com | #status log moved data disk to /var/lib/zuul on ze11 and restarted | 16:22 |
| @status:opendev.org | @jim:acmegating.com: finished logging | 16:22 |
| -@gerrit:opendev.org- Antoine Musso proposed: | 21:48 | |
| - [opendev/git-review] 1004139: Modernize str.format() for Python 3.1+ https://review.opendev.org/c/opendev/git-review/+/1004139 | ||
| - [opendev/git-review] 1004140: Fix color shifting when listing topics https://review.opendev.org/c/opendev/git-review/+/1004140 | ||
| - [opendev/git-review] 1004141: Insert column for when wip changes are present https://review.opendev.org/c/opendev/git-review/+/1004141 | ||
| @amusso:matrix.org | I decided I could slightly improve git-review to display changes that are under wip :) | 21:57 |
| @amusso:matrix.org | last one needs tests though | 21:58 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!