Python: stop the MCP lifecycle owner after a failed connect - #8755
Merged
Eduard van Valkenburg (eavanvalkenburg) merged 1 commit intoSep 28, 2026
Conversation
When a connect action failed, _run_lifecycle_owner set the exception on the caller's future and looped back to await queue.get(). The task then owned nothing and never exited, so asyncio reported "Task was destroyed but it is pending!" once the tool was garbage collected. Reaching this needs only a server that rejects the credentials: Agent._prepare_run_context enters every MCP tool through its AsyncExitStack, __aenter__ re-raises ToolException, and __aexit__ never runs because __aenter__ raised. The owner already retires in three situations, all guarded by having nothing connected and nothing queued: a cancelled connect, a connect whose caller never accepted the result, and an explicit close. A failed connect was the missing fourth case, so this adds it with the same guard. The connected check is load-bearing rather than defensive. is_connected is set before tools and prompts are loaded, so a failure while loading leaves a live session that still needs this owner to close it later. Only a connect that left no session behind may retire the owner. Verified on all three transports. Stdio leaks the same way with a command that cannot start, so the fix is placed in the shared base rather than in MCPStreamableHTTPTool, and callers see an unchanged ToolException. The existing coverage for this scenario, test_connect_cleanup_on_initialization _failure, asserts only that the exit stack is closed, which is why the leaked task went unnoticed.
Manjunath Janardhan (manjunathshiva)
deployed
to
github-app-auth
September 25, 2026 15:37 — with
GitHub Actions
Active
Manjunath Janardhan (manjunathshiva)
deployed
to
github-app-auth
September 25, 2026 15:37 — with
GitHub Actions
Active
Copilot started reviewing on behalf of
Manjunath Janardhan (manjunathshiva)
September 25, 2026 15:37
View session
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The targeted lifecycle fix is correctly guarded and comprehensively tested.
Review effort: Balanced
Findings: None
What changed in this PR
Fixes #8752 by retiring idle MCP lifecycle-owner tasks after failed connections while preserving owners for live sessions.
Changes:
- Stops owner tasks when failed connects leave no session or queued work.
- Adds coverage for cleanup, retry, and live-session behavior.
| File | Description |
|---|---|
python/packages/core/agent_framework/_mcp.py |
Retires unused lifecycle owners after connection failure. |
python/packages/core/tests/core/test_mcp.py |
Tests cleanup, retries, and live-session retention. |
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
Manjunath Janardhan (manjunathshiva)
marked this pull request as ready for review
September 25, 2026 15:40
Manjunath Janardhan (manjunathshiva)
requested review from
SergeyMenshykh,
Tao Chen (TaoChenOSU),
Eduard van Valkenburg (eavanvalkenburg),
Giles Odigwe (giles17),
Jose Alvarez (jpalvarezl),
Evan Mattson (moonbox3),
Roger Barreto (rogerbarreto) and
westey (westey-m)
as code owners
September 25, 2026 15:40
Manjunath Janardhan (manjunathshiva)
deployed
to
github-app-auth
September 25, 2026 15:40 — with
GitHub Actions
Active
Eduard van Valkenburg (eavanvalkenburg)
approved these changes
Sep 28, 2026
Eduard van Valkenburg (eavanvalkenburg)
left a comment
Member
There was a problem hiding this comment.
Manjunath Janardhan (@manjunathshiva) Looks good to me.
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Sep 28, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation & Context
When a connect action fails,
_run_lifecycle_ownersets the exception on the caller's future and loops back toawait queue.get(). The task then owns nothing and never exits, so asyncio reportsTask was destroyed but it is pending!once the tool is garbage collected, as an unrelated-looking error with no link to the failed run.Reaching this needs only a server that rejects the credentials.
Agent._prepare_run_contextenters every MCP tool through itsAsyncExitStack,MCPTool.__aenter__re-raisesToolException, and__aexit__never runs because__aenter__raised.Description & Review Guide
ToolExceptionand its message are identical. Repeated and concurrent failures no longer accumulate tasks, and a defensiveclose()after a failed connect remains safe. Since the fix is in the shared base, it covers all three transports: stdio leaks the same way with a command that cannot start, and is fixed by the same change.not self.is_connectedcondition. It is load-bearing rather than defensive, becauseis_connectedis set before tools and prompts are loaded, so a failure while loading leaves a live session that still needs this owner to close it later. Only a connect that left no session behind may retire the owner. I teeth-checked this by dropping the condition and confirming the third test fails.Two notes that may be useful. The existing coverage for this scenario,
test_connect_cleanup_on_initialization_failure, asserts only that the exit stack is closed, which is why the leaked task went unnoticed. Separately, theCould not cleanly close MCP exit stack due to cleanup error groupwarning on a failed connect is pre-existing, appears with and without this change, and is left alone here.The issue also suggests calling
close()from__aenter__before re-raising. I left that out: everyToolExceptionpath runs inside_connect_on_owner, so the owner already sees the failure, and the owner-side fix additionally covers callers who useconnect()directly rather thanasync with. Happy to add it as well if you would prefer the symmetry with the genericexcept Exceptionbranch.Related Issue
Fixes #8752
Contribution Checklist
breaking changelabel (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.