Make the default tool-rejection message state that the decision is final - #7784
Open
manjunathshiva wants to merge 1 commit into
Open
manjunathshiva wants to merge 1 commit into
manjunathshiva wants to merge 1 commit into
Conversation
When an approval is rejected without a Reason, GenerateRejectedFunctionResults returned "Tool call invocation rejected." That sentence says neither who rejected the call nor that the decision is settled, so models read it as a transient failure and ask for approval of the same tool call again. Each approval round-trip is a separate GetResponseAsync call, so MaximumIterationsPerRequest does not bound the retries and the end user is re-prompted for a tool they already denied. Measured against Azure OpenAI, three runs per model, rejecting with no reason: gpt-4o took 1 round every time, gpt-5.5 took 1, 1 and 2, gpt-5-mini took 2, 3 and 2, and gpt-5.6-terra took 5, 2 and 5, where 5 was the harness cap rather than termination. With this change all four take exactly 1 round in every run, including gpt-4o, which was already correct and does not regress. The message names the approver rather than the user because ToolApprovalRequestContent documents approval as possibly coming from "a user prompt, a policy decision, or any other approver". The finality clause is load-bearing: with attribution alone, gpt-5-mini still retried in one run of three.
Author
|
@dotnet-policy-service agree company="Accenture" |
Collaborator
🎉 Good job! The coverage increased 🎉
Full code coverage report: https://dev.azure.com/dnceng-public/public/_build/results?buildId=1611801&view=codecoverage-tab |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #7783
Problem
When an approval is rejected without a
Reason,GenerateRejectedFunctionResultsreturns"Tool call invocation rejected."to the model. That sentence says neither who rejected the call nor that the decision is settled, so models treat it as a transient failure and request approval for the same tool call again. The end user is re-prompted for a tool they already denied.MaximumIterationsPerRequestdoes not help here: it bounds iterations inside oneGetResponseAsynccall, while each approval round-trip is a separate call.Change
One string. The default rejection message now states who rejected the call and that it is final. The 21 assertions that pin the old text are updated to match.
Evidence
Approval rounds against Azure OpenAI, three runs per model, rejecting with no reason:
For gpt-5.6-terra, 5 was my harness cap rather than termination. Supplying a
Reasongives 1 round in every configuration, before and after. The rejected function is never executed in any run, so this is about wasted round-trips and repeated user prompts, not incorrect execution. gpt-4o was already correct and does not regress.Two wording decisions, both measured
ToolApprovalRequestContentdocuments approval as possibly coming from "a user prompt, a policy decision, or any other approver", so attributing the rejection to a user would be wrong whenever it comes from policy. The pre-Add Reason property to FunctionApprovalResponseContent for custom rejection messages #7140 wording said "by user"; restoring it verbatim would reintroduce that inaccuracy."...was rejected by the approver.") gpt-5-mini still retried in 1 of 3 runs. It is not decoration.Compatibility
This deliberately changes an observable string. No public API moves and the value is not a documented contract, but an application that displays, logs or asserts on the rejected-result text will see different output. Flagging it explicitly given the behavioral-compatibility guidance in CONTRIBUTING.
Happy to take whatever wording you prefer, or to close this if you would rather own the phrasing — the measurements should be useful either way. Context: reported downstream as microsoft/agent-framework#8503, and the attribution was dropped incidentally by #7140 while resolving #7139, which had only asked for custom rejection messages.
Validation
Microsoft.Extensions.AI.Tests: 761 passed, 0 failed.Microsoft Reviewers: Open in CodeFlow