Bump GitHub.Copilot.SDK - #1202
dependabot[bot] wants to merge 3 commits into
Conversation
Bumps GitHub.Copilot.SDK from 1.0.11 to 1.0.14 Bumps xunit.v3.mtp-v2 from 4.0.0 to 4.0.1 --- updated-dependencies: - dependency-name: GitHub.Copilot.SDK dependency-version: 1.0.14 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-other-nuget - dependency-name: xunit.v3.mtp-v2 dependency-version: 4.0.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: all-other-nuget ... Signed-off-by: dependabot[bot] <support@github.com>
|
❌ Evaluation did not complete successfully (the evaluate job reported 46 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
…-other-nuget-17ae3c7d6c
|
❌ Evaluation did not complete successfully (the evaluate job reported 16 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
|
👋 @dependabot[bot] — this PR has merge conflict. When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
…-other-nuget-17ae3c7d6c
|
❌ Evaluation did not complete successfully (the evaluate job reported |
|
❌ Evaluation did not complete successfully (the evaluate job reported |
|
❌ Evaluation did not complete successfully (the evaluate job reported |
|
❌ Evaluation did not complete successfully (the evaluate job reported |
|
❌ Evaluation did not complete successfully (the evaluate job reported |
|
✅ Approved by @Evangelink @JanKrivanek. cc @dotnet/skills-merge-approvers — ready to merge. |
📊 Skill and Agent Evaluation Results60 model/target results across 30 targets and 2 models — ✅ 24 improved, ➖ 29 not proven improved, Measurement identity: evaluated commit Measurement health: 60 expected / 60 observed / 60 written; 0 missing, 0 unexpected, 6 invalid; 1 recovered comparison error slot and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
|
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Avoid legacy Xamarin.Forms lifecycle methods | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Avoid legacy Xamarin.Forms lifecycle methods:Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — maui-app-lifecycle (gpt-5.6-luna)
Why: Net win +75.0% (3W/1T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +45.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 3W/1T/0L; d=3; p=0.125; net +75.0%
Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs
Overfit: High (score 0.59)
Repeated-run reliability (not used by the gate): 4 paired runs (3W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Avoid legacy Xamarin.Forms lifecycle methods | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▲ Window lifecycle event subscription | Eligible | +100.0% | +40.0% | 1/0/0 |
Illustrative judge evidence:
Avoid legacy Xamarin.Forms lifecycle methods:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — maui-data-binding (claude-sonnet-5)
Why: Net win +25.0% (2W/1T/1L over 4 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 2W/1T/1L; d=3; p=0.500; net +25.0%
Overfit: High (score 0.55)
Repeated-run reliability (not used by the gate): 4 paired runs (2W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Create and use an IValueConverter | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Implement MVVM ViewModel with ObservableObject | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Create and use an IValueConverter:Both answers fully satisfy the requested converter, XAML registration, binding, and fallback behavior. A is marginally more complete and directly usable due to its full XAML context and namespace/accessibility guidance; B is concise and correct but has less complete setup cont...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — maui-data-binding (gpt-5.6-luna)
Why: Net win +100.0% (4W/0T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +55.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 4W/0T/0L; d=4; p=0.063; net +100.0%
Overfit: Moderate (score 0.48)
Repeated-run reliability (not used by the gate): 4 paired runs (4W/0T/0L).
⚠️ Underpowered — maui-safe-area (claude-sonnet-5)
Why: Net win +100.0% (4W/0T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +85.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 4W/0T/0L; d=4; p=0.063; net +100.0%
Overfit: Moderate (score 0.47)
Repeated-run reliability (not used by the gate): 4 paired runs (4W/0T/0L).
⚠️ Underpowered — maui-safe-area (gpt-5.6-luna)
Why: Net win +50.0% (3W/0T/1L over 4 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +65.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=4; 3W/0T/1L; d=4; p=0.312; net +50.0%
Warnings: Activation-only stop: isolated 1 failed run
Overfit: Moderate (score 0.37)
Repeated-run reliability (not used by the gate): 4 paired runs (3W/0T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Handle notch and status bar safe areas on iOS | Eligible | +100.0% | +100.0% | 1/0/0 |
| ▼ Keyboard avoidance with safe area for chat UI | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Keyboard avoidance with safe area for chat UI:Response A wins because it provides a more complete and practical answer by explicitly warning about a common mistake—applying SoftInput directly to ScrollView/CollectionView—which is a critical caveat that developers would encounter. While Response B is more concise and uses ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⛔ Activation contract failed — mtp-hot-reload (gpt-5.6-luna)
Why: Net win +40.0% (7W/0T/3L over 10 preference-eligible stimulus vote(s), sign test p=0.172), mean preference +29.1% across 11 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)
Next action: Narrow skill routing so the listed off-target scenarios stay dormant.
State: VALID_NO_CHANGE (activation_contract_failed)
Gate evidence: n=10; 7W/0T/3L; d=10; p=0.172; net +40.0%; 1 dormancy excluded
Warnings: Dormancy contract: 1 unexpected activation(s); Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs
Overfit: Moderate (score 0.38)
Repeated-run reliability (not used by the gate): 11 paired runs (8W/0T/3L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Decline Test Explorer hot reload integration | Excluded (activation contract) | +100.0% | +40.0% | 1/0/0 |
| ▼ Enable a configured host that does not react to edits | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Filter an xUnit v3 hot-reload host to one test | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▼ Suggest hot reload for failing test in MTP project (SDK 10) | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Suggest launchSettings.json configuration for hot reload | Eligible | -100.0% | -100.0% | 0/0/1 |
Illustrative judge evidence:
Enable a configured host that does not react to edits:While both responses correctly identify the missing TESTINGPLATFORM_HOTRELOAD_ENABLED=1 setting and provide valid relaunch commands, Response A provides a more complete answer by offering both bash and PowerShell examples. This cross-platform guidance is more helpful for users...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.code-testing-generator (claude-sonnet-5)
Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible (sign test p=0.500 > 0.05) — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (target_agent_not_activated)
Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%
Warnings: Activation: isolated 2/5; plugin 0/5
Repeated-run reliability (not used by the gate): 5 paired runs (3W/0T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Generate a project-wide pytest suite across modules | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Generate collaborating Go package tests | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Generate layered Vitest coverage for an async cart | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▼ Generate project-wide xUnit tests for a .NET library | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Preserve a classic MSTest project while adding broad coverage | Eligible | +100.0% | +40.0% | 1/0/0 |
Illustrative judge evidence:
Generate a project-wide pytest suite across modules:Both are strong, passing, comprehensive test-suite submissions with useful evidence maps. A has a marginally more thorough reported behavioral test set, especially around edge-case boundaries and rollover effects, while B presents the evidence more cleanly.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.code-testing-generator (gpt-5.6-luna)
Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 5 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/0T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Generate collaborating Go package tests | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Preserve a classic MSTest project while adding broad coverage | Eligible | -100.0% | -100.0% | 0/0/1 |
Illustrative judge evidence:
Preserve a classic MSTest project while adding broad coverage:While Response B claims more test methods (20 vs 16), Response A delivers significantly higher quality and reliability. Response A provides concrete, verifiable evidence with specific line numbers and actual source code verification shown in the session output. It preserves al...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.test-quality-auditor (claude-sonnet-5)
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +10.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (target_agent_not_activated)
Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 1 dormancy excluded
Warnings: Activation: isolated 1/5; plugin 0/5
Repeated-run reliability (not used by the gate): 6 paired runs (4W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Assertion quality analysis | Eligible | -100.0% | -100.0% | 0/0/1 |
| ▲ Comprehensive test quality audit of weak test suite | Eligible | +100.0% | +40.0% | 1/0/0 |
| = Decline request to generate new tests | Excluded (activation contract) | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose test smells and propose a repair order | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Identify behavior gaps that existing tests would miss | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Targeted anti-pattern review | Eligible | +100.0% | +40.0% | 1/0/0 |
Illustrative judge evidence:
Assertion quality analysis:A answers the user's question with concrete, test-by-test assessment and actionable assertion-variety guidance. B fails to recover from a file-viewing issue and gives no substantive analysis.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.test-quality-auditor (gpt-5.6-luna)
Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +13.3% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (target_agent_not_activated)
Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%; 1 dormancy excluded
Warnings: Activation: isolated 1/5; plugin 1/5
Repeated-run reliability (not used by the gate): 6 paired runs (4W/0T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Assertion quality analysis | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▼ Diagnose test smells and propose a repair order | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Identify behavior gaps that existing tests would miss | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Targeted anti-pattern review | Eligible | +100.0% | +40.0% | 1/0/0 |
Illustrative judge evidence:
Diagnose test smells and propose a repair order:Response A delivers a more comprehensive and practical assessment. It identifies additional coverage gaps beyond the core test smells (RemoveItem, duplicate-ID merging, discount edge cases), explains risks with more behavioral detail, and provides a significantly more granular...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.testability-migration (claude-sonnet-5)
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +36.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05) — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (target_agent_not_activated)
Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
Warnings: Activation: isolated 0/5; plugin 0/5
Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Full pipeline: detect statics and recommend migration plan | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Inventory static dependencies without modifying the project | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▼ Migrate time dependencies and add deterministic tests | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Replace filesystem statics without touching unrelated dependencies | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Targeted request: just migrate DateTime to TimeProvider | Eligible | +100.0% | +100.0% | 1/0/0 |
Illustrative judge evidence:
Migrate time dependencies and add deterministic tests:Both implementations satisfy the task well. A is marginally stronger overall because its test suite adds an explicit after-expiry case in addition to the requested exact-boundary coverage. B has a slightly cleaner required-constructor DI shape, but the difference is minor.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.testability-migration (gpt-5.6-luna)
Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +44.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
State: VALID_NO_CHANGE (target_agent_not_activated)
Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%
Warnings: Activation: isolated 0/5; plugin 0/5
Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Full pipeline: detect statics and recommend migration plan | Eligible | +100.0% | +40.0% | 1/0/0 |
| = Inventory static dependencies without modifying the project | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▲ Migrate time dependencies and add deterministic tests | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Replace filesystem statics without touching unrelated dependencies | Eligible | +100.0% | +40.0% | 1/0/0 |
| ▲ Targeted request: just migrate DateTime to TimeProvider | Eligible | +100.0% | +100.0% | 1/0/0 |
Illustrative judge evidence:
Inventory static dependencies without modifying the project:Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — assertion-quality (claude-sonnet-5)
Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +32.5% across 16 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%
Warnings: Activation: isolated 7/8; plugin 8/8; Activation-only stop: isolated 1 failed run
Overfit: Moderate (score 0.37)
Repeated-run reliability (not used by the gate): 16 paired runs (12W/2T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Judge assertion strength in a shallow Jest suite | Eligible | +0.0% | +0.0% | 1/0/1 |
| ▼ Recognize structural and interaction checks in Jest while flagging vacuous tests | Eligible | -50.0% | -20.0% | 0/1/1 |
Illustrative judge evidence:
Judge assertion strength in a shallow Jest suite:A inspected the relevant code and delivered a precise, criterion-complete diagnosis plus actionable replacement assertions. B merely requested file paths despite the files being discoverable and provided no analysis.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — assertion-quality (gpt-5.6-luna)
Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.3% across 16 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%
Overfit: Moderate (score 0.33)
Repeated-run reliability (not used by the gate): 16 paired runs (9W/6T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Distinguish weak and meaningful assertions in a pytest suite | Eligible | +0.0% | +0.0% | 0/2/0 |
| ▼ Recognize structural and interaction checks in Jest while flagging vacuous tests | Eligible | -50.0% | -20.0% | 0/1/1 |
Illustrative judge evidence:
Distinguish weak and meaningful assertions in a pytest suite:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — code-testing-agent (claude-sonnet-5)
Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.2% across 18 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%
Warnings: Activation: isolated 6/9; plugin 3/9
Overfit: High (score 0.51)
Repeated-run reliability (not used by the gate): 18 paired runs (14W/0T/4L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▲ Add focused xUnit tests for one reservation class | Eligible | +100.0% | +40.0% | 2/0/0 |
| ▼ Expand a healthy existing pytest suite to every ledger boundary | Eligible | -100.0% | -40.0% | 0/0/2 |
| ▲ Generate a layered Vitest suite for an async shopping cart | Eligible | +100.0% | +40.0% | 2/0/0 |
| = Generate a project-wide Go suite across collaborating packages | Eligible | +0.0% | +0.0% | 1/0/1 |
| ▲ Generate a project-wide pytest suite across multiple modules | Eligible | +100.0% | +40.0% | 2/0/0 |
| = Generate project-wide tests for an SDK-style xUnit library | Eligible | +0.0% | +0.0% | 1/0/1 |
Illustrative judge evidence:
Expand a healthy existing pytest suite to every ledger boundary:Both are strong, passing suite extensions with Decimal-based boundary assertions. A is marginally better on substantive test completeness, particularly its explicit positive-balance coverage with and without a limit, while B's superior baseline verification does not outweigh t...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — coverage-analysis (claude-sonnet-5)
Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +20.0% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 4 dormancy excluded
Overfit: Moderate (score 0.35)
Repeated-run reliability (not used by the gate): 24 paired runs (13W/7T/4L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Reconcile a coverage target spread across several members | Eligible | +0.0% | +0.0% | 0/2/0 |
| ▼ Refactoring safety assessment from coverage data | Eligible | -100.0% | -40.0% | 0/0/2 |
| = Stay dormant for behavioral gap analysis | Excluded (activation contract) | +0.0% | +0.0% | 1/0/1 |
| ▼ Stay dormant for one-member CRAP analysis | Excluded (activation contract) | -50.0% | -20.0% | 0/1/1 |
| = Stay dormant for static source-to-test pairing | Excluded (activation contract) | +0.0% | +0.0% | 0/2/0 |
Illustrative judge evidence:
Reconcile a coverage target spread across several members:Both responses are correct, complete, self-contained, and directly address the coverage arithmetic, insufficiency of Apply alone, and practical next steps. B adds a correct reconciliation of the one uncovered line outside the listed methods, while A gives a similarly sound ris...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — coverage-analysis (gpt-5.6-luna)
Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.5% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded
Warnings: Activation: isolated 7/8; plugin 7/8
Overfit: Moderate (score 0.23)
Repeated-run reliability (not used by the gate): 24 paired runs (14W/8T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Coverage plateau diagnosis | Eligible | -50.0% | -20.0% | 0/1/1 |
| = Distinguish partially covered branches from covered lines | Eligible | +0.0% | +0.0% | 1/0/1 |
| = Project-wide coverage analysis with existing Cobertura data | Eligible | +0.0% | +0.0% | 0/2/0 |
Illustrative judge evidence:
Coverage plateau diagnosis:Response A delivers more actionable guidance with specific test recommendations (Criterion 4: much-better), while Response B provides clearer diagnostic structure (Criterion 2: slightly-better). Both fail to calculate the impact of fixing CalculateGpa (Criterion 3: tie) and bo...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — crap-score (claude-sonnet-5)
Why: Net win +44.4% (5W/3T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +17.8% across 9 paired run(s) — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=9; 5W/3T/1L; d=6; p=0.109; net +44.4%
Overfit: Moderate (score 0.27)
Repeated-run reliability (not used by the gate): 9 paired runs (5W/3T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Generate coverage then compute CRAP score | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Recognize when complexity alone blocks the CRAP threshold | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Recompute complexity instead of trusting a stale source comment | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Report a fully covered method at its complexity floor | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Generate coverage then compute CRAP score:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — detect-static-dependencies (gpt-5.6-luna)
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded
Overfit: Moderate (score 0.41)
Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Avoid false positives when ambient resources are already abstracted | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Detect statics inside lambda expressions and LINQ queries | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Detect time-related statics and recommend TimeProvider | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Stay dormant for Python timezone review | Excluded (activation contract) | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Avoid false positives when ambient resources are already abstracted:Both responses reached the correct conclusion (0 dependencies requiring new seams) and properly identified all injected dependencies and deterministic code. However, Response A provides accurate line numbers (9, 10, 11) while Response B reports incorrect line numbers (18, 21, ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — generate-testability-wrappers (claude-sonnet-5)
Why: Net win -28.6% (2W/1T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -0.8% across 24 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=7; 2W/1T/4L; d=6; p=0.344; net -28.6%; 1 dormancy excluded
Warnings: Activation-only stop: plugin 1 failed run
Overfit: Moderate (score 0.46)
Repeated-run reliability (not used by the gate): 24 paired runs (11W/3T/10L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Generate TimeProvider adoption for DateTime.UtcNow | Eligible | +0.0% | +0.0% | 1/1/1 |
| ▼ Generate a minimal process runner wrapper | Eligible | -33.3% | -53.3% | 1/0/2 |
| ▼ Generate custom Environment wrapper | Eligible | -33.3% | -13.3% | 1/0/2 |
| ▼ Generate only the console members a prompt uses | Eligible | -100.0% | -80.0% | 0/0/3 |
| ▼ Make time controllable in a library that has no DI container | Eligible | -33.3% | -13.3% | 1/0/2 |
| ▲ Recommend System.IO.Abstractions for file system calls | Eligible | +100.0% | +80.0% | 3/0/0 |
Illustrative judge evidence:
Generate TimeProvider adoption for DateTime.UtcNow:Both deliver the right .NET 10 abstraction, compile, and include working deterministic tests. A is stronger because it demonstrates advancing FakeTimeProvider and verifying changed time-dependent behavior, directly satisfying the more demanding test criterion. B's explicit pro...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — maui-collectionview (claude-sonnet-5)
Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%
Warnings: Activation-only stop: isolated 1 failed run
Overfit: Moderate (score 0.31)
Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Basic CollectionView with data binding and DataTemplate | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Selection and pull-to-refresh with CollectionView | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Avoid ListView and ViewCell mistakes:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — maui-collectionview (gpt-5.6-luna)
Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%
Overfit: Moderate (score 0.23)
Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Grid layout with CollectionView | Eligible | -100.0% | -40.0% | 0/0/1 |
| = ItemSizingStrategy placement for uniform items | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Avoid ListView and ViewCell mistakes:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — maui-dependency-injection (claude-sonnet-5)
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Warnings: Activation-only stop: isolated 2 failed runs
Overfit: High (score 0.53)
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Avoid AddScoped pitfall in MAUI | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose a page whose injected dependencies are missing | Eligible | +100.0% | +40.0% | 1/0/0 |
| = Register services with correct lifetimes in MauiProgram.cs | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Avoid AddScoped pitfall in MAUI:Both give useful, concise EF Core guidance including a context factory, transient lifetime, and explicit scopes. A is better because its explanation of MAUI scoped lifetime is accurate, whereas B incorrectly asserts automatic per-window IServiceScopes.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — maui-shell-navigation (claude-sonnet-5)
Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +32.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%
Overfit: High (score 0.62)
Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Set up Shell navigation with tabs and flyout | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Set up Shell navigation with tabs and flyout:Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — maui-shell-navigation (gpt-5.6-luna)
Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%
Overfit: High (score 0.62)
Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Diagnose common Shell navigation mistakes | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Set up Shell navigation with tabs and flyout | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Stable routes for deep linking into tabs | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Diagnose common Shell navigation mistakes:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — maui-theming (gpt-5.6-luna)
Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +44.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%
Overfit: High (score 0.56)
Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Detect and respond to system theme changes | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Detect and respond to system theme changes:Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — migrate-static-to-wrapper (gpt-5.6-luna)
Why: Net win +50.0% (6W/3T/1L over 10 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +26.0% across 10 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=10; 6W/3T/1L; d=7; p=0.063; net +50.0%
Repeated-run reliability (not used by the gate): 10 paired runs (6W/3T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Migrate a static helper class without breaking its callers | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Migrate only in scoped files, leaving others untouched | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Preserve DateTimeOffset values during TimeProvider migration | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Preserve local calendar semantics when migrating DateTime.Now | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Migrate a static helper class without breaking its callers:Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — run-tests (gpt-5.6-luna)
Why: Net win +13.6% (6W/13T/3L over 22 preference-eligible stimulus vote(s), sign test p=0.254), mean preference +8.2% across 22 paired run(s) — not credible (sign test p=0.254 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=22; 6W/13T/3L; d=9; p=0.254; net +13.6%
Warnings: Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 1 failed run
Repeated-run reliability (not used by the gate): 22 paired runs (6W/13T/3L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Collect a crash dump on an MTP project (SDK 9) | Eligible | -100.0% | -100.0% | 0/0/1 |
| = Collect coverage on an SDK 9 MTP bridge | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Collect coverage with VSTest | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Enable a diagnostic log for a VSTest project | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Enable diagnostic logs for a native MTP project | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Filter NUnit tests on an SDK 9 MTP bridge | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Filter TUnit tests by class using treenode-filter | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Filter one NUnit class on VSTest | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Filter xUnit v3 tests by class on MTP | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Filter xUnit v3 tests by class pattern and trait using query filter language | Eligible | -100.0% | -100.0% | 0/0/1 |
| = Generate TRX from a VSTest project | Eligible | +0.0% | +0.0% | 0/1/0 |
| = MTP project on SDK 10 passes args directly | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Run one VSTest invocation without rebuilding | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Run tests in a VSTest MSTest project | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Select one target framework in a multi-targeted project | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Use the repository unit-test entry point | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Collect a crash dump on an MTP project (SDK 9):Response A provides the correct, documented command using standard dotnet CLI options (--blame-crash --blame-crash-dump-type full) for collecting crash dumps from test runs. Response B provides an incorrect command using a non-existent--crashdumpparameter with inaccurate...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Routine passing details for 5 results are in Full Results.
Details for 26 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.
🔍 Full Results - all metrics and investigation details
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1202 in dotnet/skills, download eval artifacts with
gh run download 36708764237 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://github.com/raw/dotnet/skills/b4fba51282c6a0f28029015f1f094256da2e4c31/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
Rebasing might not happen immediately, so don't worry if this takes some time.
Note: if you make any changes to this PR yourself, they will take precedence over the rebase.
Updated GitHub.Copilot.SDK from 1.0.11 to 1.0.14.
Release notes
Sourced from GitHub.Copilot.SDK's releases.
1.0.14
Feature: typed message provenance for user, system, and agent sources
Messages sent through the SDK can now carry typed source provenance, distinguishing human
userinput, internalsysteminjections, and identifiedagent-senders, so recipients can reliably tell agent input from human authorization. (#2573)Feature: Auto model routing Fast tier
Sessions using
automodel routing can now select thefasttier alongside the existingefficiency,balance, andintelligencetiers, giving integrators a latency-focused routing preset across all six SDKs. (#2669)Feature: force-refresh managed settings cache
The new
managedSettings.clearCacheRPC method wipes the persistent server-policy cache and drops the runtime's in-memory retained policy, giving hosts a primitive for a "force refresh account policy" action. (#2438)Feature: Rust SDK model allowlists
SessionConfigandResumeSessionConfigin the Rust SDK now accept an optionalallowed_modelslist, letting hosts restrict which model IDs a session may use without duplicating runtime validation. (#2512)Other changes
max_output_tokensto the model capabilities override, previously unreachable without an unsafe cast (#2569)PackAsToolpackages sodotnet pack --no-buildproduces a working tool (#2557)connection_closecallback-quiescence contract consistently across all six in-process C ABI adapters, preventing races with freed callback state during disposal (#2610, #2622)CatalogTrustEligibilityunknownvalue and re-exporting shared session-event types (#2631)anyOf/oneOfhandling (#2656)... (truncated)
1.0.14-preview.1
Feature: pause and resume durable factory runs at checkpoints
Agent Factories can now pause deliberately instead of only stopping at hard limits. Call
ctx.pause(key)inside a factory body to register a durable, one-shot checkpoint that ends the current attempt; resuming replays the journal and returns from that checkpoint instead of redoing prior work. Callers can also pause a running attempt from outside the factory body. (#2537)1.0.14-preview.0
Feature: typed message provenance across all SDKs
Sending a message can now declare its source as
user,system, or an identified agent (serialized asagent-<id>), so recipients can reliably distinguish human input from system injections and forwarded agent output. Ordinary sends remain unaffected: source stays omitted unless the caller opts in. (#2573)Source = MessageSource.Agent("reviewer").setSource(MessageSource.agent("reviewer")).with_source(MessageSource::Agent("reviewer".into()))Feature: force-refresh enterprise managed settings
A new
managedSettings.clearCacheRPC wipes the persistent server-policy cache and drops the runtime's in-memory retained policy, so hosts can wire up a "sync account policy" action (for example VS Code'sDeveloper: Sync Account Policycommand). It's available in TypeScript, C#, Python, Go, and Rust; Java support follows once the underlying CLI release is pinned. (#2438)Other changes
max_output_tokenson model capability overrides (#2569)PackAsToolpublish output so packed tools install correctly (#2557)... (truncated)
1.0.13
Feature: cancellation for host-owned external tools
Host-owned external tool callbacks are now cancelled when their runtime request completes or their SDK session terminates. The cancellation primitive is idiomatic per SDK: .NET passes a request token to
AIFunction, Node.js exposesToolInvocation.signal, Go cancelsToolInvocation.TraceContext, Java cancels the returnedCompletableFuture, Python cancels the handler task, and Rust drops the handler future. Go handlers that retainTraceContextfor background work must derive a separate lifetime because the invocation context is cancelled when the request ends.Feature: declare application identity with client info
Client options now accept optional client info (application name and version, integration name and version) across all six SDKs, exposed idiomatically per language (
clientInfoin Node.js,client_infoin Python and Rust,ClientInfoin Go and .NET,setClientInfoin Java). When set, the SDK forwards it on theserver.connecthandshake so the telemetry the runtime emits on the connection is attributed to the application and its Copilot integration instead of the runtime's own build. All fields are optional, and leaving client info unset keeps the runtime's default attribution. See Client info.Feature: Node Agent Factories pagination and run notifications
The experimental Node.js Agent Factories convenience API now supports paginated run history. Existing
session.factory.listRuns()calls still return the runs array, while calls withafterSeq,beforeSeq, orlimitreturn the full page with cursor and truncation metadata.Factory
runandresumeoptions now acceptnotifyOnCompleteandlogPhaseNames. The SDK forwards these options to the Copilot CLI for new and resumed runs.Feature: selectable
ask_usersession behaviorSession create and cold resume now accept a language-specific
askUserVariantoption withlegacyandelicitationvalues. SDK sessions retain the legacy question-and-answer tool by default. Selectelicitationand provide an elicitation handler to expose the structured form-basedask_usertool.Feature: rotating session-scoped GitHub credentials
All six SDKs can now acquire short-lived GitHub credentials through a session-scoped callback. The SDK registers the callback before session create or resume, maps
initialandrefreshrequests to the owning session, and removes registrations on rollback, replacement, session close, and client close. Static per-sessiongitHubTokencredentials remain supported and are mutually exclusive with the callback.Token responses use the shared tagged token/cancelled shape and require
expiresIn, expressed as the positive number of seconds remaining when the callback completes. See github/copilot-agent-runtime#16381 for the runtime credential-authority implementation.Initial acquisition occurs during create or resume; cancellation, callback errors, and invalid credentials reject that operation instead of falling back to ambient authentication. Idle sessions refresh only before their next credential-consuming operation.
Feature: extensions can request sensitive environment variables
Copilot CLI extensions can now ask for named sensitive environment variables when they join a session.
joinSession()accepts arequestedEnvironmentVariablesoption listing the variable names the extension needs. The CLI shows a permission prompt naming the extension and the exact variables requested. On approval, only those variables reach that extension and their values are written into the extension process'sprocess.envbeforejoinSession()resolves. On denial,joinSession()rejects, the extension does not load, and its tools never reach the model.An approval is remembered against the exact set of names the user saw, so an extension that later asks for one more variable prompts again. Names that are unset, or that the CLI does not filter from extensions, are not prompted for. This is the client half of the feature; it requires a Copilot CLI that supports extension environment access, and older CLIs ignore the request and grant nothing.
Feature: early session-event subscription (Rust)
The Rust SDK can now observe every event routed to a session, starting with that session's very first routed event.
Client::prepare_sessionandClient::prepare_resume_sessionreturn an inertPreparedSessionthat owns the session's event channel, so a subscription can be installed before any protocol activity begins:Feature: session-scoped GitHub token providers
Sessions now support expiry-aware GitHub token callbacks in addition to static tokens. The SDK handles refresh requests from the runtime, so extensions always receive fresh credentials. (#2412)
Feature: Java in-process native runtime on all platforms
... (truncated)
1.0.13-preview.3
Feature: rewind support across all SDKs
Sessions can now opt into file-change tracking and conversation rewind. When
enableFileChangeTrackingis enabled, the session records which files were changed during a conversation turn. You can then list pending rewind points, preview changes, and rewind the conversation history together with any tracked file modifications. (#2321)Feature: session-scoped GitHub token providers
Sessions now support a dynamic, expiry-aware GitHub token callback as an alternative to a static
gitHubToken. The SDK maps each host request (with host, session, and reason context) to your callback, handling concurrent-session isolation automatically. (#2412)Feature: built-in plugin directory support
... (truncated)
1.0.13-preview.2
Feature: rewind support across all SDKs
Sessions can now opt in to file-change tracking so that rewinding restores both conversation history and the files that were modified. Enable it with the new
enableFileChangeTrackingsession option. (#2321)Feature: session-scoped GitHub token providers
Applications can now supply a dynamic GitHub token callback instead of a static
gitHubTokenstring. The runtime calls the callback before each token use, so short-lived tokens stay fresh across long-running sessions. (#2412)1.0.13-preview.0
Feature: rewind support across all SDKs
Sessions can now opt into file-change tracking and rewind conversation history along with tracked file changes. Enable the new
enableFileChangeTrackingsession option to allow calling rewind later. (#2321)Feature: Java in-process runtime (experimental)
The Java SDK now ships platform-native classifier JARs that load the Copilot runtime directly in-process via JNA — no separate CLI child process required. Currently available for linux-x64, Windows x64, and Apple Silicon macOS. (#2301, #2393, #2402)
Feature: permission decision context forwarding
Permission handlers can now attach
decisionContextso the runtime can attribute whether a decision came from a person, host policy, or an automated recommendation. This is additive for Node, Python, Go, .NET, and Java. Rust clients that construct or matchPermissionResult::Decisiondirectly must migrate to the new struct variant. (#2294)createAttributedPermissionResult(result, context)copilot.create_attributed_permission_result(result, context)copilot.NewAttributedPermissionResult(result, context)DecisionContexton the permission decisionPermissionRequestResult.approveOnce().setDecisionContext(context)PermissionResult::approve_once().with_context(context)Feature: built-in plugin directory support
Applications can now register a set of host-bundled plugin directories that are trusted unconditionally and loaded before any user session begins. (#2330)
Feature: extensions can request sensitive environment variables (Node)
... (truncated)
1.0.12-preview.0
Feature: rewind support across all SDKs
Sessions now support rewinding conversation history and tracked file changes. Enable file-change tracking when creating a session, then rewind to a previous checkpoint to discard later turns and restore file state. (#2321)
Feature: Java in-process Copilot CLI (linux-x64)
The Java SDK now supports an in-process connection mode on linux-x64 that loads the Copilot runtime as a native library via JNA — no separate CLI child process required. Add the
copilot-sdk-java-runtimeclassifier JAR for your platform alongside the core SDK JAR. (#2301)