Skip to content

asyncio: ProactorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF #156333

Description

@discovery-ukraine

Bug report

Bug description:

On Windows, BaseProactorEventLoop wakes itself through a self-pipe, which is a loopback TCP socketpair created by socket.socketpair(). If that connection reaches a clean EOF while the loop is running -- for example the OS tears the idle loopback connection down across a power/session state change -- the loop spins at 100% of one core forever, with no exception raised and nothing logged.

Lib/asyncio/proactor_events.py:

def _loop_self_reading(self, f=None):
    try:
        if f is not None:
            f.result()  # may raise
        if self._self_reading_future is not f:
            return
        f = self._proactor.recv(self._ssock, 4096)
    except exceptions.CancelledError:
        return
    except (SystemExit, KeyboardInterrupt):
        raise
    except BaseException as exc:
        self.call_exception_handler({
            'message': 'Error on reading from the event loop self pipe',
            'exception': exc,
            'loop': self,
        })
    else:
        self._self_reading_future = f
        f.add_done_callback(self._loop_self_reading)

At EOF, f.result() returns b''. That is not an exception, so no except branch runs and control falls through to else, which re-arms recv() on a socket that is at EOF. That recv completes immediately, whose callback is _loop_self_reading, which re-arms again -- a tight infinite loop.

The loop is otherwise completely idle. py-spy shows MainThread as active+gil with:

_loop_self_reading (asyncio/proactor_events.py:804)
  recv             (asyncio/windows_events.py:493)
    _register      (asyncio/windows_events.py:720)

and _run_once locals sched_count: 0, ntodo: 1, timeout: None -- nothing is scheduled, one handle re-runs forever.

Related precedent

IocpProactor.recv() in Lib/asyncio/windows_events.py returns b'' in this situation, while recv_into() was changed to return 0 in bpo-41467 -- "asyncio: recv_into() must not return b'' if the socket/pipe is closed". The same reasoning applies to recv(). Independently of that, _loop_self_reading has no EOF handling at all, so a dead self-pipe can never be recovered from.

Reproducer

Deterministic -- 12/12 runs on both interpreters tested.

import asyncio
import socket
import time


async def main():
    loop = asyncio.get_running_loop()

    # Let the loop arm its self-pipe read first.
    await asyncio.sleep(0.1)

    # Graceful half-close: the read half sees a clean EOF, which is what an OS
    # teardown of the loopback connection looks like.
    loop._csock.shutdown(socket.SHUT_WR)

    start = time.process_time()
    await asyncio.sleep(3)
    print(f"CPU consumed while sleeping 3s: {time.process_time() - start:.2f}s")


asyncio.run(main())  # Windows default is ProactorEventLoop

Expected: roughly 0.00s -- the process is asleep.
Actual: roughly 2.90s -- a full core burned for a three second sleep.

Important: it must be a graceful half-close. An abortive close() surfaces as an exception, which the existing except BaseException branch handles, so it does not reproduce reliably (~80% of runs, and only when the recv is issued after the socket is already gone).

How this was hit in production

Three unrelated long-running Python programs on the same Windows 11 machine -- two independent MCP servers plus a minimal control program written purely to isolate this -- all began pinning a core at the same instant, roughly 69 minutes into their lifetime. They had accumulated near-identical CPU time (5465s, 5465s, 5485.7s, 5482.7s over ~9638s of uptime), which is what pointed at a shared external trigger rather than three independent bugs.

The trigger turned out to be waking the screen after the machine had been locked -- not going idle. With the display off the machine was silent; moving the mouse lit the screen and every asyncio process immediately pinned a core.

Socket state confirms the mechanism:

sockets CPU
healthy process 127.0.0.1:A->B ESTABLISHED + 127.0.0.1:B->A ESTABLISHED (plus vestigial 0.0.0.0:A BOUND) 0%
spinning process both ESTABLISHED halves gone, only BOUND left 97-99%

Nothing is logged, no exception is raised, and the affected processes never recover. From a user's point of view the machine simply starts running hot after every unlock, with several cores pinned, and the only remedy is killing the processes.

The machine has no Modern Standby (powercfg /a reports S3 only), so this is an ordinary session/display transition on a fully awake system, not a sleep/resume cycle.

Environment

  • Windows 11
  • Python 3.12.10 (tags/v3.12.10:0cc8128, MSC v.1943 64 bit) -- affected
  • Python 3.13.12 (main, MSC v.1944 64 bit) -- affected
  • The code path is unchanged on main as of this writing.

Suggested fix

Treat a zero-length result as the self-pipe being gone, and rebuild it instead of re-arming a read that can never block again:

def _loop_self_reading(self, f=None):
    try:
        if f is not None:
            data = f.result()  # may raise
            if not data:
                # The self-pipe reached EOF: the socketpair is gone (this can
                # happen when the OS tears down the loopback connection across a
                # power or session state change). Re-arming here would spin the
                # CPU forever, so rebuild the pipe instead.
                self._self_reading_future = None
                self._close_self_pipe()
                self._make_self_pipe()
                return
        ...

A more conservative variant would be to leave the rebuild out and only stop re-arming, which at least turns an invisible 100% CPU spin into a loop that can no longer be woken -- but rebuilding keeps the loop functional, which seems strictly better.

Making IocpProactor.recv() return a length rather than b'', matching what bpo-41467 did for recv_into(), would additionally make this class of bug harder to reintroduce.

CPython versions tested on:

3.12, 3.13

Operating systems tested on:

Windows

Linked PRs

Activity

  1. added
    stdlibStandard Library Python modules in the Lib/ directory
    on Aug 25, 2026
  2. aidaodedjl commented on Aug 25, 2026

    @aidaodedjl

    I can reproduce this deterministically on a self-built main (3.16.0a0): 346,702 _loop_self_reading re-arms during a 3s idle sleep after the SHUT_WR half-close, ~1.4s of CPU burned. The empty-result re-arm chain is exactly as analyzed — f.result() returns b'', no except branch runs, recv on the dead socket completes immediately, forever.

    One extra data point while digging: the selector twin has the same hole. SelectorEventLoop._read_from_self breaks on the empty read but leaves the reader registered, and a readable-forever closed socket spins the select loop the same way — same reproducer shape against asyncio.SelectorEventLoop gives 582,692 _read_from_self calls over 3s idle. I'll file that separately.

    Fix for the proactor side is up in #156343: rebuild the socketpair on the empty result (new pair allocated before touching the old sockets, wakeup fd re-registered before the old sockets close, next read armed on the new socket). Reproducer drops to 0 re-arms / 0.00s CPU with it, full test_asyncio suite clean.

  3. pangi commented on Sep 1, 2026

    @pangi

    Sent here from modelcontextprotocol/python-sdk, where this was diagnosed and
    closed as upstream. The mechanism in both issues matches what I measured, so I
    will not restate it. Four things from the field instead.

    1. A trigger that needs no suspend, and takes ten seconds

    The reports so far reach the EOF through a suspend/resume cycle. On this machine
    the cause turned out to be an in-path WFP filter driver re-applying its
    filters
    — a consumer ad-blocker with a network shield
    (adgnetworkwfpdrv.sys, loaded at System start).

    Toggling that product's protection off and back on, with nothing else happening:

    19:11:41  monitoring 2 idle asyncio processes, self-pipe pairs Established
              (protection off, ~20s, protection on)
    19:16:57  process A - loopback pair 49209<->49210 no longer Established
    19:17:00  process A - 99% of one core
    19:17:00  process B - loopback pair 49233<->49234 no longer Established
    19:17:04  process B - 100% of one core
    

    Both processes, three seconds apart. No sleep involved; system uptime
    continuous. Stacks are the ones already in this issue —
    _loop_self_reading (proactor_events.py:802) re-arming, _poll returning
    immediately.

    Repeated later the same evening with the same result, so two for two on this
    machine. What I have not isolated is which of the two WFP transitions does it,
    disabling or re-enabling: both deaths land within three seconds of each other,
    which means one transition kills both processes, but the test covered both.

    It is a clean EOF, not a reset. In the second run the loop was carrying an
    instrumented _loop_self_reading that reports why it fired, and it said EOF
    both times — never ConnectionResetError. So this really is the graceful-close
    path, which is the one that busy-loops silently rather than raising.

    So a class of ordinary consumer software reproduces this on demand.

    2. It leaves no trace in any Windows log

    This is the part I would flag for severity. At the moment the pair died, nothing
    was recorded in System, Application, NetworkProfile, Kernel-PnP or Defender:

    19:16:57   no events in any watched log in the last 90s
    

    Two consequences:

    • A user cannot diagnose this. There is no error, no log line, no exception —
      only a pinned core and, if several asyncio processes are running, several
      pinned cores at once. Every idle asyncio process on the box goes together,
      because they all lose their socketpair in the same instant.
    • It is probably underreported for the same reason. The suspend/resume trigger
      is memorable; "my CPU is at 100% and nothing says why" is not something people
      connect to asyncio.

    3. What the socket table shows, in case it helps the fix

    The connected rows disappear entirely — the pair is not left half-closed:

    healthy   127.0.0.1:55306 -> 127.0.0.1:55307  Established
              0.0.0.0:55307   -> 0.0.0.0:0        Bound
              127.0.0.1:55307 -> 127.0.0.1:55306  Established
    
    spinning  0.0.0.0:55307   -> 0.0.0.0:0        Bound
    

    Same process, same thread, twelve minutes apart. The Bound remnant keeps the
    identical port and the handle count is unchanged, so the interpreter never
    closed either fd — the connection was removed underneath it. A synthetic pair
    behaves the same way: half-closed states last a couple of seconds and then both
    rows are gone while both fds stay open.

    If a fix rebuilds the pair on EOF, that is exactly the state it has to recover
    from: two live fds whose connection no longer exists in the TCP table.

    4. Rebuild-on-EOF holds up outside a test harness

    Small vote of confidence for the approach in the open PRs. I am running a
    monkeypatched _loop_self_reading / _read_from_self that rebuilds the pair on
    EOF or error, on both loop types, inside two long-lived stdio servers. With the
    trigger above applied:

    20:10:14  self-pipe pair 49603<->49604 no longer Established
    20:10:17  same process measured at 0% of one core
              server log: "self-pipe was dead (EOF); rebuilt"
    

    New pair, new ports, no spin, and the servers kept answering. Before the patch
    the identical trigger put both of them at 99-100% of a core within three
    seconds.

    I also checked the thing a rebuild could plausibly break: set_wakeup_fd has to
    follow the new _csock, or cross-thread wakeups go quiet in a way nothing
    reports. call_soon_threadsafe still works after the rebuild on both loop types.
    Worth an assertion in the PR tests if there is not one already.

    Environment

    Windows 10.0.19045, uv-managed CPython 3.13.14, mcp 1.29.x, two idle stdio
    servers spawned by the same host application. Full py-spy dumps (four samples
    per process, --native), socket snapshots at detection time, and a healthy
    baseline of both are available for every episode since 2026-08-29 — say the word
    and I will attach them.

  4. added a commit that references this issue on Sep 4, 2026
  5. aidaodedjl commented on Sep 4, 2026

    @aidaodedjl

    All four points land, and the WFP trigger is the most useful thing this issue has gotten — a reproducer that needs no suspend cycle is the difference between "we believe the reporter" and "anyone can confirm in ten seconds."

    On point 4: the wakeup-fd concern is already covered in both PRs. #156343 re-registers signal.set_wakeup_fd(csock.fileno()) before the old sockets are closed, and the test asserts the new fd; #156345 is stricter — it only migrates when the wakeup fd actually points at the loop's own _csock (a foreign registration is restored untouched), and keeps the old write end alive when the rebuild runs on a worker thread. call_soon_threadsafe after rebuild is a regression test on both loop types, so nothing goes quietly.

    The clean-EOF detail in point 1 matters more than it might look: it confirms this is the graceful-close path, which is exactly the one that spins silently instead of raising. A ConnectionResetError would at least have surfaced somewhere.

    Yes — please attach the py-spy dumps and socket snapshots. Three episodes of field data beat a synthetic reproducer if a core dev asks how this behaves on real machines.

  6. added a commit that references this issue on Sep 4, 2026
  7. pangi commented on Sep 4, 2026

    @pangi

    Attached: cpython-156333-field-data.zip (32 kB) — py-spy dumps, socket snapshots and monitor traces, grouped into eight directories with a README that says what each one demonstrates.

    Short version:

    • 01, 02 — the first two episodes: proactor, then the same machine with WindowsSelectorEventLoopPolicy forced. 01 includes a healthy baseline of the same server, which is the useful part — identical frames, MainThread (idle) instead of (active+gil).
    • 03 — no suspend at all: continuous uptime, no Kernel-Power 42/107/131, zero wakes in powercfg /lastwake.
    • 04 — same pids and same threads, twelve minutes apart, healthy then spinning. Clearest look at what happens to the socket table.
    • 05 — a three-second S3 was enough, with the surrounding event log.
    • 06 — the WFP filter trigger, deliberately applied.
    • 07 — negative control: a Hyper-V vSwitch NIC torn down and rebuilt on purpose. The socketpairs survived on the same ports, so a network-adapter rebuild alone does not do it.
    • 08 — the same trigger with rebuild-on-EOF in place.

    Three days on from 08, that monkeypatch has rebuilt the pair 13 times, every one of them with cause EOF and not a single error case. Over the same three days the watchdog — sampling every five minutes, and with 14 runaways logged before — has recorded none.

    One caveat on the 13: I cannot give you individual timestamps. Python block-buffers stderr to a pipe, so those lines only reached the host's log when each process exited, which parks every pair of them at a session boundary regardless of when the rebuild happened. The zero-runaway figure comes from the watchdog's
    own timestamped log, so that half is solid. The patch now timestamps and flushes its own output, so anything from here on will be exact.

    Notes on the files, so nothing looks odd: the Windows user name is replaced with USER and nothing else was altered; one 2026-08-29 capture is deliberately excluded because it is a synthetic while True: pass decoy I used to test the watchdog, not field data; and 05-...-pid27176 has three of four samples because
    the process exited while py-spy was sampling it.

    Good to hear the wakeup-fd migration is already handled in both PRs, and that call_soon_threadsafe after rebuild is a regression test on both loop types — that was the one thing I could see going quiet without anyone noticing.

    The capture runs automatically now, so if you want a specific measurement taken at the moment the pair dies — socket options, an SO_ERROR read on _ssock, a handle table snapshot — name it and it will be in the next episode.

    cpython-156333-field-data.zip

  8. aidaodedjl commented on Sep 4, 2026

    @aidaodedjl

    Downloaded and read the whole archive — this is now the evidence set I'd point a core dev at, in this order: 04 (same pid, twelve minutes apart, the TCP table going from one loopback pair Established both ways plus the Bound remnant to only the Bound remnant) is the clearest single exhibit for what rebuild-on-EOF has to recover from, and 01's healthy baseline makes the spin legible at a glance: identical frames, MainThread (idle) becomes (active+gil).

    07 is more valuable than a control usually is: the pairs survived a deliberate NIC teardown/rebuild on the same ports, so this is not generic network-stack churn — it is something about the WFP filter re-application pass specifically. That plus 03 (continuous uptime, zero power events) closes off the "it's really just suspend/resume being misdiagnosed" explanation completely.

    The stderr block-buffering caveat is good data hygiene to state explicitly, and the watchdog's own timestamped zero-runaway figure is the one that matters anyway — 14 runaways before the patch, zero in the three days after, on the same machine doing the same work.

    For the capture rig, two things would be worth having at the moment of death, if the next episode allows:

    1. An SO_ERROR getsockopt on both ends at detection time. If it returns 0 while the TCP rows are gone, that documents why the re-arm loop spins silently — nothing on the socket itself reports an error, only the read can reveal the EOF. That is the strongest argument that a recv-side fix is the only possible interception point.
    2. Whichever WFP transition is the killer (disable or re-enable), caught with netsh wfp show state captured immediately before and after the toggle. That would name the layer, and "filter re-application removes loopback TCP connections" is the kind of sentence that gets this fixed at the right level — or at least gets the issue a footnote in WFP-facing release notes.

    The 13-rebuilds-all-EOF stat is exactly what the PRs needed: it says the graceful-close path is not an edge case triggered by an unusual test, it is the only path this bug uses in the wild.

  9. pangi commented on Sep 4, 2026

    @pangi

    Correction to my earlier comment: the WFP-toggle reproducer does not hold up, and
    I withdraw it.
    Please don't spend time trying to confirm it as described.

    What I tested today

    Same machine, same two servers, now with the guard timestamping and flushing its own
    output so every self-pipe death has an exact time. Five attempts:

    attempt result
    disable protection only no death
    enable protection only no death
    disable only, again, servers 6 min old no death
    enable only, again no death
    disable, wait ~20 s, enable — the exact original sequence no death

    Pairs unchanged throughout, same ports, CPU flat, and the watchdog logged no runaway.

    Why I got it wrong

    Both original hits landed within seconds of a toggle, and I took that for causation.
    What I had not noticed at the time: over the three days since the guard went in, it
    has rebuilt the pair 13 times with nobody touching the ad-blocker at all. So the
    trigger fires on its own a few times a day. Two coincidences in a row with something
    that frequent is unremarkable — and I only ever measured while we were toggling.

    So the trigger is not identified. It is something that happens spontaneously a
    few times a day on this machine, silently, and I do not yet know what.

    What does still stand

    None of the following depends on the trigger:

    • The mechanism, with the same-process before/after in directory 04 of the archive:
      connected loopback rows gone, Bound remnant on the same port, handle count
      unchanged, _loop_self_reading re-arming.
    • It needs no suspend: directory 03, continuous uptime, zero wakes in
      powercfg /lastwake.
    • A three-second S3 is enough when suspend is involved: directory 05.
    • Rebuild-on-EOF works: 13 rebuilds, all with cause EOF, and zero runaways in three
      days against 14 in the days before, on the same machine doing the same work.

    The SO_ERROR reading you asked for

    Captured inside the process at the moment of detection, before the old sockets are
    closed:

    2026-09-04 13:37:22 self-pipe was dead (EOF); rebuilt
      _ssock local=127.0.0.1:61320 peer=127.0.0.1:61321 SO_ERROR=0
      _csock local=127.0.0.1:61321 peer=127.0.0.1:61320 SO_ERROR=0
    

    Zero on both ends, both fds still open, and getpeername() still resolving on
    both — while the connected rows are already gone from the TCP table. Nothing on the
    socket reports anything wrong; only the read reveals it. That is your argument for a
    recv-side interception point, in the strongest form I can give it.

    The WFP diff, and a caveat on the wording

    netsh wfp show state before the toggle, after disabling, and after re-enabling
    (attached as wfp-summaries.zip — the provider's filters and callouts per layer,
    reduced from three 28 MB dumps).

    The three extracts are byte-identical, same SHA-256, apart from the source
    filename in the header:

    baseline protection off protection on again
    provider's filters 176 176 176
    provider's callouts 32 32 32
    filters on the box 13972 13972 13972
    FWPM_LAYER_STREAM_V4 / _V6 2 / 2 2 / 2 2 / 2

    The provider keeps every filter and callout registered whether protection is on or
    off. Where it sits, for what it is worth: 2 of the 4 filters at
    FWPM_LAYER_STREAM_V4, the same at _V6, 2 of 6 at
    FWPM_LAYER_ALE_FLOW_ESTABLISHED_V4/V6, both filters at
    FWPM_LAYER_ALE_ENDPOINT_CLOSURE_V4/V6, and 56 Stream Callout plus 72
    Flow Established Callout entries.

    So "filter re-application removes loopback TCP connections" is not what this
    measures — the filter set does not change across the transition. If an in-path
    filter is involved at all, it would have to be a callout dropping flows it was
    inspecting while its registrations stay put. I would rather hand you the empty diff
    than a sentence that reads well and is not supported.

    Next

    The guard now timestamps and flushes, and captures SO_ERROR on both ends
    automatically. The next spontaneous event — a few times a day, going by the count —
    will arrive with an exact second and the surrounding event log, and I will post it.
    Given the trigger leaves no trace in any Windows log, that may be the only way to
    name it.

    wfp-summaries.zip

  10. aidaodedjl commented on Sep 4, 2026

    @aidaodedjl

    Withdrawing your own reproducer after two public threads credited it is not the usual move — noted and respected. Correcting my own comment above: disregard the "narrowed to the WFP filter re-application pass" reading of 06/07. The toggles were coincidences; the archive's mechanism exhibits (03/04/05, the negative control in 07 as a negative control) all stand on their own.

    The empty WFP diff is worth more than a matching one would have been. Same SHA-256 across the transition, every filter and callout still registered: there is no static configuration change to point at, and a callout dropping a flow it inspects would leave exactly this kind of nothing behind. Combined with the 13 unattended rebuilds, the honest summary is: the trigger is spontaneous, silent, unlogged, and possibly never going to be nameable.

    Which flips the framing for this issue. Until today the story was "reproduces via WFP toggling, suspend/resume, three-second S3" — multiple causes, and a reviewer could reasonably ask which one is the real target. The corrected story is stronger: the cause does not matter and cannot be relied on. SO_ERROR=0 on both ends with getpeername() still resolving, while the connected rows are already gone from the TCP table, is the whole argument in one capture — the kernel has torn the connection down, the socket objects have no idea, and the only API surface that can possibly discover the state is the read. There is no earlier interception point and no diagnostic that beats it. Any loop that treats "read returned EOF" as a reason to immediately re-arm the same read is wrong on its face; what the EOF means (peer half-closed? connection destroyed underneath us?) is exactly what the loop cannot know, which is why rebuilding is the only bounded response available.

    So the pull requests are not waiting on a trigger attribution, and I would argue against deferring them until one arrives — an unnameable trigger is precisely the case where the runtime has to be defensive. If a future capture does name the cause, that becomes a nice README footnote, not a prerequisite.

  11. added a commit that references this issue on Sep 4, 2026
  12. msullivan commented on Oct 3, 2026

    @msullivan
    Contributor

    I somewhat lost track of the thread of robots talking to each other in this issue, but is there a real repro procedure (that doesn't involve manually sabotaging event loop internals)?

  13. xksk-doer commented on Oct 4, 2026

    @xksk-doer

    Thanks for the analysis in this issue — we hit the sibling failure mode of the same code line on Windows in production, and it is not covered by the current patch in #156343.

    What we see

    _loop_self_reading() can also lose the pipe through the exception path, not just EOF. When the pending recv() on self._ssock is aborted, f.result() raises instead of returning b'':

    ERROR asyncio: Error on reading from the event loop self pipe
    loop: <ProactorEventLoop running=True closed=False>
    Traceback (most recent call last):
      File ".../Lib/asyncio/proactor_events.py", line 789, in _loop_self_reading
        f.result()  # may raise
      File ".../Lib/asyncio/windows_events.py", line 807, in _poll
        value = callback(transferred, key, ov)
      File ".../Lib/asyncio/windows_events.py", line 467, in finish_socket_func
        raise ConnectionResetError(*exc.args)
    ConnectionResetError: [WinError 995] The I/O operation has been aborted because of either a thread exit or an application request
    

    finish_socket_func() maps ERROR_OPERATION_ABORTED (995) and ERROR_NETNAME_DELETED (1236) to ConnectionResetError, so this reads like "the peer reset us", but the winerror says it was aborted locally.

    On main, that path goes to except BaseException → call_exception_handler(...) → return with no re-arm, so the loop keeps running=True while its self-pipe is dead. Every subsequent call_soon_threadsafe() / run_coroutine_threadsafe() from another thread enqueues a callback that is never run, because nothing wakes the loop any more.

    Note this is strictly worse than the EOF case in this issue: EOF (and #156343's fix) pins a core and is visible; the abort path is silent — one log line, then wakeups are lost forever. It also affects any exception raised there, not only 995.

    Reproducer (deterministic, Windows-only, 3.11.15 and 3.14.7)

    import asyncio, threading, time, sys, ctypes, os
    from ctypes import wintypes
    
    k32 = ctypes.WinDLL("kernel32", use_last_error=True)
    k32.CancelIoEx.argtypes = [wintypes.HANDLE, ctypes.c_void_p]
    
    errors = []
    loop = asyncio.ProactorEventLoop()
    loop.set_exception_handler(lambda l, ctx: errors.append(ctx))
    threading.Thread(target=loop.run_forever, daemon=True).start()
    time.sleep(0.5)
    
    def wakes():
        box = {"hit": False}
        loop.call_soon_threadsafe(lambda: box.__setitem__("hit", True))
        time.sleep(1.0)
        return box["hit"]
    
    before = wakes()                                            # True
    k32.CancelIoEx(wintypes.HANDLE(loop._ssock.fileno()), None) # abort the pending read
    time.sleep(1.2)
    after = wakes()                                             # False  <-- wakeups gone for good
    err = next((c for c in errors if "self pipe" in (c.get("message") or "")), None)
    print(f"{sys.version.split()[0]} before={before} err={err is not None} "
          f"winerror={getattr(err.get('exception'), 'winerror', None) if err else None} after={after}")
    os._exit(0)

    Output on both interpreters (3/3 runs each):

    3.14.7  before=True err=True winerror=995 after=False
    3.11.15 before=True err=True winerror=995 after=False
    

    CancelIoEx is only used to inject the abort deterministically; closing loop._ssock from another thread produces the same permanent loss (with winerror=1236). Note shutdown(SHUT_RD)/shutdown(SHUT_RDWR) on the pipe sockets does not trigger it — it is specific to a cancelled/interrupted overlapped operation.

    Why #156343 as written doesn't fix it

    The patch only rebuilds on the EOF branch:

    if f is not None and not data:
        self._rebuild_self_pipe()

    The except BaseException branch is unchanged, so an aborted read still leaves the loop un-wakeable. Would you be open to also rebuilding (or at least re-arming) there? The rebuild helper looks like the right recovery for both cases; if rebuilding on the abort path is considered risky, re-arming at minimum makes the loop recover on the next wakeup instead of staying broken forever.

    Production evidence (why we chased this)

    Long-lived Windows processes that create a private event loop per streaming HTTP request (several hundred loops/hour). In one 4-hour window:

    • 8 of 8 "stream stalled, no chunks received" episodes began at exactly the timestamp of a self-pipe error — measured as stale_report_time − stale_threshold (i.e. request start) vs the logged error timestamp: delta 0.0 s in all 8 cases.
    • The affected loops were running=True, closed=False, and died 2–20 s after creation.
    • WinError 10054 / WSAECONNRESET / RemoteDisconnected: 0 occurrences — the transport was never reset by the remote side.
    • The event is provider-agnostic, and it always affects whichever loop is in use at that moment.

    Happy to test a patch on Windows (3.14.7 + 3.11.15) if that helps.

  14. added a commit that references this issue on Oct 5, 2026
  15. aidaodedjl commented on Oct 5, 2026

    @aidaodedjl

    Short answer: no clean one today.

    • S3 suspend/resume kills idle proactor loops reliably (directory 05 of the archive), but it isn't on-demand or CI-able.
    • The WFP-toggle trigger posted earlier didn't survive retesting and was withdrawn. On that machine the underlying event is spontaneous: a few times a day, silent, absent from every Windows event log, so nobody can invoke it at will.
    • The closest thing to a real procedure without monkeypatching is injecting at the operation layer instead of the loop layer: CancelIoEx on the pending self-pipe read. It delivers the same failure the kernel reports for a cancelled or reset operation (WinError 995, surfacing as ConnectionResetError through finish_socket_func), it's four lines of ctypes against a stock interpreter (no internals replaced, one handle passed to one API), and it lands the loop in the running=True, wakeups-dead state deterministically. Still injection, but it exercises the exception path the suspend/EOF repros never reach; I used it to verify the exception-path recovery on stock 3.13.15 and 3.10.11 (last commit in gh-156333: rebuild the proactor self-pipe on EOF instead of busy-looping #156343).

    With SO_ERROR reading 0 on both ends at the moment of death, I wouldn't bet on a nameable real-world trigger. The read is the only surface that observes this state. If a CI-able repro matters for the PRs, CancelIoEx is the honest best available.

  16. msullivan commented on Oct 5, 2026

    @msullivan
    Contributor
  17. aidaodedjl commented on Oct 5, 2026

    @aidaodedjl

    Tested it in a Windows 11 guest on Parallels (26200, stock 3.13.13): a 30-second VM suspend/resume with the loop idling left wakeups working, no self-pipe error after resume. The guest's TCP state is part of the snapshot, so the loopback pair comes back exactly as it went in.

    Also tried the sharper variant, disconnecting and reconnecting the guest's vNIC from the host while the loop idled: wakeups still fine. Loopback traffic never touches the NIC, and that matches the dir 07 negative control in the archive (a deliberate Hyper-V vSwitch rebuild didn't kill the pairs either).

    So suspending a VM reproduces the freeze without the death. Hardware S3 kills (dir 05) because the host reinitializes the network stack under the running process; a VM snapshot just pauses and restores it. A hypervisor setup that reattaches network hardware on resume instead of restoring it might behave differently, but plain suspend on Parallels does not.

  18. msullivan commented on Oct 5, 2026

    @msullivan
    Contributor
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    stdlibStandard Library Python modules in the Lib/ directorytopic-asynciotype-bugAn unexpected behavior, bug, or error

    Projects

    • Status
      Todo

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions