Skip to content

Bump GitHub.Copilot.SDK - #1202

Open
dependabot[bot] wants to merge 3 commits into
mainfrom
dependabot/nuget/eng/skill-validator/src/all-other-nuget-17ae3c7d6c
Open

dependabot[bot] wants to merge 3 commits into
mainfrom
dependabot/nuget/eng/skill-validator/src/all-other-nuget-17ae3c7d6c

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ Dependabot is rebasing this PR ⚠️

Rebasing might not happen immediately, so don't worry if this takes some time.

Note: if you make any changes to this PR yourself, they will take precedence over the rebase.


Updated GitHub.Copilot.SDK from 1.0.11 to 1.0.14.

Release notes

Sourced from GitHub.Copilot.SDK's releases.

1.0.14

Feature: typed message provenance for user, system, and agent sources

Messages sent through the SDK can now carry typed source provenance, distinguishing human user input, internal system injections, and identified agent- senders, so recipients can reliably tell agent input from human authorization. (#​2573)

await session.send("Looks good to me.", { source: "agent-reviewer" });
await session.send("Looks good to me.", source=AgentMessageSource("reviewer"))

Feature: Auto model routing Fast tier

Sessions using auto model routing can now select the fast tier alongside the existing efficiency, balance, and intelligence tiers, giving integrators a latency-focused routing preset across all six SDKs. (#​2669)

await session.setAutoTier("fast");

Feature: force-refresh managed settings cache

The new managedSettings.clearCache RPC method wipes the persistent server-policy cache and drops the runtime's in-memory retained policy, giving hosts a primitive for a "force refresh account policy" action. (#​2438)

await client.rpc.managedSettings.clearCache();
await client.Rpc.ManagedSettings.ClearCacheAsync();

Feature: Rust SDK model allowlists

SessionConfig and ResumeSessionConfig in the Rust SDK now accept an optional allowed_models list, letting hosts restrict which model IDs a session may use without duplicating runtime validation. (#​2512)

let config = SessionConfig::default().with_allowed_models(["gpt-4o", "claude-3.7-sonnet"]);

Other changes

  • feature: [Core] add factory pause checkpoints for the Node.js Agent Factories API, letting a paused run resume without losing invocation limits or execution identity (#​2537)
  • feature: forward the optional host OAuth client metadata URL across all six SDKs on session create and resume (#​2258)
  • feature: [TypeScript] add max_output_tokens to the model capabilities override, previously unreachable without an unsafe cast (#​2569)
  • bugfix: [.NET] include Copilot CLI runtime assets in PackAsTool packages so dotnet pack --no-build produces a working tool (#​2557)
  • bugfix: apply the runtime's connection_close callback-quiescence contract consistently across all six in-process C ABI adapters, preventing races with freed callback state during disposal (#​2610, #​2622)
  • bugfix: [Rust] fix codegen for CLI 1.0.84 schemas, correctly mapping the CatalogTrustEligibility unknown value and re-exporting shared session-event types (#​2631)
  • bugfix: [C#] fix codegen for runtime schema unions, unblocking single-variant anyOf/oneOf handling (#​2656)
  • bugfix: [Rust] isolate the hostless in-process runtime cache from the bundled CLI cache to prevent cross-deletion of a shared executable (#​2659)
    ... (truncated)

1.0.14-preview.1

Feature: pause and resume durable factory runs at checkpoints

Agent Factories can now pause deliberately instead of only stopping at hard limits. Call ctx.pause(key) inside a factory body to register a durable, one-shot checkpoint that ends the current attempt; resuming replays the journal and returns from that checkpoint instead of redoing prior work. Callers can also pause a running attempt from outside the factory body. (#​2537)

await ctx.step("prepare", prepareInput);
await ctx.pause("review-ready");
await ctx.agent("Review the prepared input");
const paused = await session.factory.pause(runId);

[!WARNING]

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • github.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "github.com"

See Network Configuration for more information.

Generated by Release Changelog Generator · copilot · auto · 52.5 AIC · ⌖ 5.57 AIC · ⊞ 10.3K

1.0.14-preview.0

Feature: typed message provenance across all SDKs

Sending a message can now declare its source as user, system, or an identified agent (serialized as agent-<id>), so recipients can reliably distinguish human input from system injections and forwarded agent output. Ordinary sends remain unaffected: source stays omitted unless the caller opts in. (#​2573)

await session.send({ prompt: "Reviewed and approved.", source: "agent-reviewer" });
await session.send(prompt="Reviewed and approved.", source=AgentMessageSource("reviewer"))
session.Send(ctx, copilot.SendOptions{Source: copilot.MessageSourceAgent("reviewer")})
  • C#: Source = MessageSource.Agent("reviewer")
  • Java: .setSource(MessageSource.agent("reviewer"))
  • Rust: .with_source(MessageSource::Agent("reviewer".into()))

Feature: force-refresh enterprise managed settings

A new managedSettings.clearCache RPC wipes the persistent server-policy cache and drops the runtime's in-memory retained policy, so hosts can wire up a "sync account policy" action (for example VS Code's Developer: Sync Account Policy command). It's available in TypeScript, C#, Python, Go, and Rust; Java support follows once the underlying CLI release is pinned. (#​2438)

await client.rpc.managedSettings.clearCache();

Other changes

  • feature: [TypeScript] allow setting max_output_tokens on model capability overrides (#​2569)
  • feature: forward an optional host OAuth client metadata document URL across all six SDKs (#​2258)
  • bugfix: [.NET] include Copilot CLI runtime assets in PackAsTool publish output so packed tools install correctly (#​2557)
  • bugfix: [Rust] isolate GitHub token callbacks from the request-routing loop so a callback panic can no longer stall session requests (#​2567)
  • bugfix: [Python, Node.js] serialize concurrent client startup so racing callers reuse the same connection instead of launching duplicate runtimes (#​2570)
  • bugfix: [.NET] fix a vulnerable transitive SourceLink dependency (#​2587)
  • bugfix: fix in-process callback reclamation and FFI close responsiveness across all six SDKs to prevent use-after-free races and event-loop stalls during disposal (#​2610, #​2622)
  • bugfix: [Rust] fix codegen for CLI 1.0.84 schemas (#​2631)

[!WARNING]

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • github.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

... (truncated)

1.0.13

Feature: cancellation for host-owned external tools

Host-owned external tool callbacks are now cancelled when their runtime request completes or their SDK session terminates. The cancellation primitive is idiomatic per SDK: .NET passes a request token to AIFunction, Node.js exposes ToolInvocation.signal, Go cancels ToolInvocation.TraceContext, Java cancels the returned CompletableFuture, Python cancels the handler task, and Rust drops the handler future. Go handlers that retain TraceContext for background work must derive a separate lifetime because the invocation context is cancelled when the request ends.

Feature: declare application identity with client info

Client options now accept optional client info (application name and version, integration name and version) across all six SDKs, exposed idiomatically per language (clientInfo in Node.js, client_info in Python and Rust, ClientInfo in Go and .NET, setClientInfo in Java). When set, the SDK forwards it on the server.connect handshake so the telemetry the runtime emits on the connection is attributed to the application and its Copilot integration instead of the runtime's own build. All fields are optional, and leaving client info unset keeps the runtime's default attribution. See Client info.

Feature: Node Agent Factories pagination and run notifications

The experimental Node.js Agent Factories convenience API now supports paginated run history. Existing session.factory.listRuns() calls still return the runs array, while calls with afterSeq, beforeSeq, or limit return the full page with cursor and truncation metadata.

Factory run and resume options now accept notifyOnComplete and logPhaseNames. The SDK forwards these options to the Copilot CLI for new and resumed runs.

Feature: selectable ask_user session behavior

Session create and cold resume now accept a language-specific askUserVariant option with legacy and elicitation values. SDK sessions retain the legacy question-and-answer tool by default. Select elicitation and provide an elicitation handler to expose the structured form-based ask_user tool.

Feature: rotating session-scoped GitHub credentials

All six SDKs can now acquire short-lived GitHub credentials through a session-scoped callback. The SDK registers the callback before session create or resume, maps initial and refresh requests to the owning session, and removes registrations on rollback, replacement, session close, and client close. Static per-session gitHubToken credentials remain supported and are mutually exclusive with the callback.

Token responses use the shared tagged token/cancelled shape and require expiresIn, expressed as the positive number of seconds remaining when the callback completes. See github/copilot-agent-runtime#​16381 for the runtime credential-authority implementation.

Initial acquisition occurs during create or resume; cancellation, callback errors, and invalid credentials reject that operation instead of falling back to ambient authentication. Idle sessions refresh only before their next credential-consuming operation.

Feature: extensions can request sensitive environment variables

Copilot CLI extensions can now ask for named sensitive environment variables when they join a session. joinSession() accepts a requestedEnvironmentVariables option listing the variable names the extension needs. The CLI shows a permission prompt naming the extension and the exact variables requested. On approval, only those variables reach that extension and their values are written into the extension process's process.env before joinSession() resolves. On denial, joinSession() rejects, the extension does not load, and its tools never reach the model.

An approval is remembered against the exact set of names the user saw, so an extension that later asks for one more variable prompts again. Names that are unset, or that the CLI does not filter from extensions, are not prompted for. This is the client half of the feature; it requires a Copilot CLI that supports extension environment access, and older CLIs ignore the request and grant nothing.

import { joinSession } from "@​github/copilot-sdk/extension";

const session = await joinSession({
    requestedEnvironmentVariables: ["GITHUB_TOKEN"],
});
const token = process.env.GITHUB_TOKEN;

Feature: early session-event subscription (Rust)

The Rust SDK can now observe every event routed to a session, starting with that session's very first routed event. Client::prepare_session and Client::prepare_resume_session return an inert PreparedSession that owns the session's event channel, so a subscription can be installed before any protocol activity begins:

let prepared = client.prepare_session(
    SessionConfig::default().with_event_buffer_capacity(2048),
)?;
let mut events = prepared.subscribe();
 ... (truncated)

## 1.0.13-preview.4

### Feature: rewind support across all SDKs

Sessions can now opt into file-change tracking and rewind conversation history and tracked file changes to any prior checkpoint. Enable file tracking when creating a session, then use `rewind` to roll back. ([#​2321](https://github.com/github/copilot-sdk/pull/2321))

```ts
const session = await client.createSession({ enableFileChangeTracking: true });
// ...later
const points = await session.rpc.rewind.list();
await session.rpc.rewind.rewind({ rewindTarget: points[0].id });
var session = await client.CreateSessionAsync(new SessionOptions { EnableFileChangeTracking = true });
var points = await session.Rpc.Rewind.ListAsync();
await session.Rpc.Rewind.RewindAsync(new RewindRequest { RewindTarget = points[0].Id });
session = await client.create_session(enable_file_change_tracking=True)
points = await session.rpc.rewind.list()
await session.rpc.rewind.rewind(rewind_target=points[0].id)

Feature: session-scoped GitHub token providers

Sessions now support expiry-aware GitHub token callbacks in addition to static tokens. The SDK handles refresh requests from the runtime, so extensions always receive fresh credentials. (#​2412)

const session = await client.createSession({
  gitHubTokenProvider: async ({ host, reason }) => ({ token: await fetchToken(host) })
});
var session = await client.CreateSessionAsync(new SessionOptions {
    GitHubTokenProvider = async (req, ct) =>
        new GitHubTokenResult { Token = await FetchTokenAsync(req.Host) }
});
session, _ := client.CreateSession(ctx, copilot.SessionOptions{
    GitHubTokenProvider: func(ctx context.Context, req copilot.TokenProviderRequest) (copilot.TokenProviderResult, error) {
        return copilot.TokenProviderResult{Token: fetchToken(req.Host)}, nil
    },
})

Feature: Java in-process native runtime on all platforms

... (truncated)

1.0.13-preview.3

Feature: rewind support across all SDKs

Sessions can now opt into file-change tracking and conversation rewind. When enableFileChangeTracking is enabled, the session records which files were changed during a conversation turn. You can then list pending rewind points, preview changes, and rewind the conversation history together with any tracked file modifications. (#​2321)

const session = await client.createSession({ enableFileChangeTracking: true });
const points = await session.rpc.rewind.listPendingRewindPoints();
await session.rpc.rewind.rewind({ id: points[0].id });
var session = await client.CreateSessionAsync(new SessionOptions { EnableFileChangeTracking = true });
var points = await session.Rpc.Rewind.ListPendingRewindPointsAsync();
await session.Rpc.Rewind.RewindAsync(new RewindRequest { Id = points[0].Id });
session = await client.create_session(enable_file_change_tracking=True)
points = await session.rpc.rewind.list_pending_rewind_points()
await session.rpc.rewind.rewind(id=points[0].id)

Feature: session-scoped GitHub token providers

Sessions now support a dynamic, expiry-aware GitHub token callback as an alternative to a static gitHubToken. The SDK maps each host request (with host, session, and reason context) to your callback, handling concurrent-session isolation automatically. (#​2412)

const session = await client.createSession({
  gitHubTokenProvider: async ({ host }) => ({ token: await getToken(host), expiresIn: 3600 }),
});
var session = await client.CreateSessionAsync(new SessionOptions
{
    GitHubTokenProvider = async (req, ct) =>
        new GitHubToken { Token = await GetTokenAsync(req.Host, ct), ExpiresIn = TimeSpan.FromHours(1) }
});
session, err := client.CreateSession(ctx, copilot.SessionOptions{
    GitHubTokenProvider: func(ctx context.Context, req copilot.GitHubTokenRequest) (copilot.GitHubToken, error) {
        return copilot.GitHubToken{Token: getToken(req.Host), ExpiresIn: 3600}, nil
    },
})

Feature: built-in plugin directory support

... (truncated)

1.0.13-preview.2

Feature: rewind support across all SDKs

Sessions can now opt in to file-change tracking so that rewinding restores both conversation history and the files that were modified. Enable it with the new enableFileChangeTracking session option. (#​2321)

const session = await client.startSession({ enableFileChangeTracking: true });
var session = await client.StartSessionAsync(new SessionOptions { EnableFileChangeTracking = true });
session = await client.start_session(enable_file_change_tracking=True)
session, _ := client.StartSession(ctx, &copilot.SessionOptions{EnableFileChangeTracking: true})
Session session = client.startSession(new SessionOptions().setEnableFileChangeTracking(true)).get();
let session = client.start_session(SessionOptions { enable_file_change_tracking: Some(true), ..Default::default() }).await?;

Feature: session-scoped GitHub token providers

Applications can now supply a dynamic GitHub token callback instead of a static gitHubToken string. The runtime calls the callback before each token use, so short-lived tokens stay fresh across long-running sessions. (#​2412)

const session = await client.startSession({
  gitHubTokenProvider: async ({ host, reason }) => ({ token: await fetchToken(host) })
});
var session = await client.StartSessionAsync(new SessionOptions
{
    GitHubTokenProvider = async (request, ct) => new GitHubTokenResult(await FetchTokenAsync(request.Host))
});
async def token_provider(request):
    return GitHubTokenResult(token=await fetch_token(request.host))

session = await client.start_session(github_token_provider=token_provider)
 ... (truncated)

## 1.0.13-preview.1

### Feature: `ClientMode::Empty` now disables built-in skills by default

`ClientMode::Empty` now applies deny-by-default isolation to runtime-bundled skills in addition to other built-in capabilities. `includedBuiltinSkills` defaults to `[]` in Empty mode; pass an explicit allowlist to re-enable specific skills. This behavior is consistent across all six SDKs. ([#​2410](https://github.com/github/copilot-sdk/pull/2410))

```ts
// Node — empty mode: built-in skills excluded by default
const session = await client.createSession({ mode: ClientMode.Empty });
// opt back in:
const session = await client.createSession({ mode: ClientMode.Empty, includedBuiltinSkills: ["edit"] });
// C#
var session = await client.CreateSessionAsync(new SessionOptions { Mode = ClientMode.Empty });
// opt back in:
var session = await client.CreateSessionAsync(new SessionOptions { Mode = ClientMode.Empty, IncludedBuiltinSkills = ["edit"] });
# Python
session = await client.create_session(mode=ClientMode.EMPTY)
# opt back in:
session = await client.create_session(mode=ClientMode.EMPTY, included_builtin_skills=["edit"])
// Go
session, err := client.CreateSession(ctx, copilot.SessionOptions{Mode: copilot.ClientModeEmpty})
// opt back in:
session, err := client.CreateSession(ctx, copilot.SessionOptions{Mode: copilot.ClientModeEmpty, IncludedBuiltinSkills: []string{"edit"}})

Generated by Release Changelog Generator · sonnet46 28.6 AIC · ⌖ 4.12 AIC · ⊞ 8.1K

1.0.13-preview.0

Feature: rewind support across all SDKs

Sessions can now opt into file-change tracking and rewind conversation history along with tracked file changes. Enable the new enableFileChangeTracking session option to allow calling rewind later. (#​2321)

const session = await client.createSession({ enableFileChangeTracking: true });
// later:
await session.rpc.conversation.rewind({ ...rewindPoint });
var session = await client.CreateSessionAsync(new SessionOptions { EnableFileChangeTracking = true });
session = await client.create_session(enable_file_change_tracking=True)
session, err := client.CreateSession(ctx, copilot.SessionOptions{EnableFileChangeTracking: true})

Feature: Java in-process runtime (experimental)

The Java SDK now ships platform-native classifier JARs that load the Copilot runtime directly in-process via JNA — no separate CLI child process required. Currently available for linux-x64, Windows x64, and Apple Silicon macOS. (#​2301, #​2393, #​2402)

CopilotClientOptions options = new CopilotClientOptions()
    .setConnection(RuntimeConnection.forInProcess());
CopilotClient client = new CopilotClient(options);
client.start().get();

Feature: permission decision context forwarding

Permission handlers can now attach decisionContext so the runtime can attribute whether a decision came from a person, host policy, or an automated recommendation. This is additive for Node, Python, Go, .NET, and Java. Rust clients that construct or match PermissionResult::Decision directly must migrate to the new struct variant. (#​2294)

  • TypeScript: createAttributedPermissionResult(result, context)
  • Python: copilot.create_attributed_permission_result(result, context)
  • Go: copilot.NewAttributedPermissionResult(result, context)
  • C#: set DecisionContext on the permission decision
  • Java: PermissionRequestResult.approveOnce().setDecisionContext(context)
  • Rust: PermissionResult::approve_once().with_context(context)

Feature: built-in plugin directory support

Applications can now register a set of host-bundled plugin directories that are trusted unconditionally and loaded before any user session begins. (#​2330)

Feature: extensions can request sensitive environment variables (Node)

... (truncated)

1.0.12-preview.0

Feature: rewind support across all SDKs

Sessions now support rewinding conversation history and tracked file changes. Enable file-change tracking when creating a session, then rewind to a previous checkpoint to discard later turns and restore file state. (#​2321)

const session = await client.createSession({ enableFileChangeTracking: true });
const rewindPoints = await session.rpc.rewind.listRewindPoints();
await session.rpc.rewind.rewind({ rewindPointId: rewindPoints[0].rewindPointId });
session = await client.create_session(enable_file_change_tracking=True)
rewind_points = await session.rpc.rewind.list_rewind_points()
await session.rpc.rewind.rewind(rewind_point_id=rewind_points[0].rewind_point_id)
session, _ := client.CreateSession(ctx, &copilot.SessionOptions{EnableFileChangeTracking: true})
points, _ := session.RPC.Rewind.ListRewindPoints(ctx)
_ = session.RPC.Rewind.Rewind(ctx, &copilot.RewindRequest{RewindPointId: points[0].RewindPointId})
var session = await client.CreateSessionAsync(new SessionOptions { EnableFileChangeTracking = true });
var points = await session.Rpc.Rewind.ListRewindPointsAsync();
await session.Rpc.Rewind.RewindAsync(new RewindRequest { RewindPointId = points[0].RewindPointId });
SessionOptions options = new SessionOptions().setEnableFileChangeTracking(true);
var session = client.createSession(options).get();
var points = session.getRpc().getRewind().listRewindPoints().get();
session.getRpc().getRewind().rewind(new RewindRequest().setRewindPointId(points.get(0).getRewindPointId())).get();
let session = client.create_session(SessionOptions { enable_file_change_tracking: Some(true), ..Default::default() }).await?;
let points = session.rpc.rewind.list_rewind_points().await?;
session.rpc.rewind.rewind(RewindRequest { rewind_point_id: points[0].rewind_point_id.clone() }).await?;

Feature: Java in-process Copilot CLI (linux-x64)

The Java SDK now supports an in-process connection mode on linux-x64 that loads the Copilot runtime as a native library via JNA — no separate CLI child process required. Add the copilot-sdk-java-runtime classifier JAR for your platform alongside the core SDK JAR. (#​2301)

CopilotClientOptions options = new CopilotClientOptions()
    .setConnection(RuntimeConnection.forInProcess());
CopilotClient client = new CopilotClient(options);
client.start().get();
 ... (truncated)

Commits viewable in [compare view](https:https://github.com/github/copilot-sdk/compare/v1.0.11...v1.0.14).
</details>

Updated [xunit.v3.mtp-v2](https:https://github.com/xunit/xunit) from 4.0.0 to 4.0.1.

<details>
<summary>Release notes</summary>

_Sourced from [xunit.v3.mtp-v2's releases](https://github.com/xunit/xunit/releases)._

No release notes found for this version range.

Commits viewable in [compare view](https:https://github.com/xunit/xunit/commits).
</details>

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency
- `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself)
- `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself)
- `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself)
- `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency
- `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions


</details>

Bumps GitHub.Copilot.SDK from 1.0.11 to 1.0.14
Bumps xunit.v3.mtp-v2 from 4.0.0 to 4.0.1

---
updated-dependencies:
- dependency-name: GitHub.Copilot.SDK
  dependency-version: 1.0.14
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-other-nuget
- dependency-name: xunit.v3.mtp-v2
  dependency-version: 4.0.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: all-other-nuget
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added .NET Pull requests that update .NET code dependencies Pull requests that update a dependency file labels Sep 23, 2026
@dependabot dependabot Bot added dependencies Pull requests that update a dependency file .NET Pull requests that update .NET code labels Sep 23, 2026
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 23, 2026
github-actions Bot added a commit that referenced this pull request Sep 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate ace1ca09cd7c2547a4726f3d3a84846bd19dcdd1 to retry this exact commit.

46 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed pr-state/evals-in-progress PR evaluations are in progress labels Sep 23, 2026
@Evangelink
Evangelink enabled auto-merge (squash) September 24, 2026 07:49
@github-actions github-actions Bot added pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 24, 2026
github-actions Bot added a commit that referenced this pull request Sep 24, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate cbc3e2e03b97a2a91a7cba5b606ead1892ecf077 to retry this exact commit.

16 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

@github-actions github-actions Bot added waiting-on-author PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Sep 24, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 @dependabot[bot] — this PR has merge conflict. When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

@Evangelink Evangelink changed the title Bump GitHub.Copilot.SDK and xunit.v3.mtp-v2 Bump GitHub.Copilot.SDK Sep 24, 2026
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed waiting-on-author PR state label pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 24, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b4fba51282c6a0f28029015f1f094256da2e4c31 to retry this exact commit.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 28, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b4fba51282c6a0f28029015f1f094256da2e4c31 to retry this exact commit.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b4fba51282c6a0f28029015f1f094256da2e4c31 to retry this exact commit.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b4fba51282c6a0f28029015f1f094256da2e4c31 to retry this exact commit.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b4fba51282c6a0f28029015f1f094256da2e4c31 to retry this exact commit.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress ready-to-merge PR state label and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor

✅ Approved by @Evangelink @JanKrivanek. cc @dotnet/skills-merge-approvers — ready to merge.

@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

60 model/target results across 30 targets and 2 models — ✅ 24 improved, ➖ 29 not proven improved, ⚠️ 6 invalid or underpowered, ⛔ 1 activation contract failure, 📉 0 preference losses (report only).

Measurement identity: evaluated commit b4fba51282c6a0f28029015f1f094256da2e4c31; 2 judge models.

Measurement health: 60 expected / 60 observed / 60 written; 0 missing, 0 unexpected, 6 invalid; 1 recovered comparison error slot and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.code-testing-generator claude-sonnet-5 ➖ Not proven improved n=5; 3W/0T/2L; d=5; p=0.500; net +20.0% — Activation: isolated 2/5; plugin 0/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.code-testing-generator gpt-5.6-luna ➖ Not proven improved n=5; 3W/0T/2L; d=5; p=0.500; net +20.0% — — Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.test-quality-auditor claude-sonnet-5 ➖ Not proven improved n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 1 dormancy excluded — Activation: isolated 1/5; plugin 0/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.test-quality-auditor gpt-5.6-luna ➖ Not proven improved n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%; 1 dormancy excluded — Activation: isolated 1/5; plugin 1/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.testability-migration claude-sonnet-5 ➖ Not proven improved n=5; 4W/0T/1L; d=5; p=0.188; net +60.0% — Activation: isolated 0/5; plugin 0/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.testability-migration gpt-5.6-luna ➖ Not proven improved n=5; 4W/1T/0L; d=4; p=0.063; net +80.0% — Activation: isolated 0/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
assertion-quality claude-sonnet-5 ➖ Not proven improved n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% 🟡 0.37 Activation: isolated 7/8; plugin 8/8; Activation-only stop: isolated 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
assertion-quality gpt-5.6-luna ➖ Not proven improved n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% 🟡 0.33 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
code-testing-agent claude-sonnet-5 ➖ Not proven improved n=9; 6W/2T/1L; d=7; p=0.063; net +55.6% 🔴 0.51 Activation: isolated 6/9; plugin 3/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
code-testing-agent gpt-5.6-luna ✅ Improved n=9; 6W/3T/0L; d=6; p=0.016; net +66.7% — — None.
coverage-analysis claude-sonnet-5 ➖ Not proven improved n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 4 dormancy excluded 🟡 0.35 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
coverage-analysis gpt-5.6-luna ➖ Not proven improved n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded 🟡 0.23 Activation: isolated 7/8; plugin 7/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
crap-score claude-sonnet-5 ➖ Not proven improved n=9; 5W/3T/1L; d=6; p=0.109; net +44.4% 🟡 0.27 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
crap-score gpt-5.6-luna ✅ Improved n=9; 8W/1T/0L; d=8; p=0.004; net +88.9% 🟡 0.37 — Review overfit evidence.
detect-static-dependencies claude-sonnet-5 ✅ Improved n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded 🟡 0.21 — Review overfit evidence.
detect-static-dependencies gpt-5.6-luna ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.41 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
dotnet-maui-doctor claude-sonnet-5 ✅ Improved n=8; 8W/0T/0L; d=8; p=0.004; net +100.0% 🔴 0.54 — Review overfit evidence.
dotnet-maui-doctor gpt-5.6-luna ✅ Improved n=8; 6W/2T/0L; d=6; p=0.016; net +75.0% 🔴 0.64 — Review overfit evidence.
find-untested-sources claude-sonnet-5 ✅ Improved n=9; 8W/1T/0L; d=8; p=0.004; net +88.9% 🔴 0.56 Activation: isolated 9/9; plugin 8/9 Fix activation gaps; Review overfit evidence.
find-untested-sources gpt-5.6-luna ✅ Improved n=9; 8W/1T/0L; d=8; p=0.004; net +88.9% — — None.
generate-testability-wrappers claude-sonnet-5 ➖ Not proven improved n=7; 2W/1T/4L; d=6; p=0.344; net -28.6%; 1 dormancy excluded 🟡 0.46 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
generate-testability-wrappers gpt-5.6-luna ✅ Improved n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded 🟡 0.28 1 judge slot recovered Review overfit evidence; Review recovered judge slots.
grade-tests claude-sonnet-5 ✅ Improved n=8; 8W/0T/0L; d=8; p=0.004; net +100.0% 🟡 0.48 — Review overfit evidence.
grade-tests gpt-5.6-luna ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5% 🟡 0.35 — Review overfit evidence.
maui-app-lifecycle claude-sonnet-5 ⚠️ Underpowered n=4; 3W/1T/0L; d=3; p=0.125; net +75.0% 🔴 0.62 Activation-only stop: isolated 1 failed run Predeclare more independent, discriminating stimuli; repeated runs do not add power.
maui-app-lifecycle gpt-5.6-luna ⚠️ Underpowered n=4; 3W/1T/0L; d=3; p=0.125; net +75.0% 🔴 0.59 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs Predeclare more independent, discriminating stimuli; repeated runs do not add power.
maui-collectionview claude-sonnet-5 ➖ Not proven improved n=5; 2W/2T/1L; d=3; p=0.500; net +20.0% 🟡 0.31 Activation-only stop: isolated 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
maui-collectionview gpt-5.6-luna ➖ Not proven improved n=5; 2W/2T/1L; d=3; p=0.500; net +20.0% 🟡 0.23 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
maui-data-binding claude-sonnet-5 ⚠️ Underpowered n=4; 2W/1T/1L; d=3; p=0.500; net +25.0% 🔴 0.55 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
maui-data-binding gpt-5.6-luna ⚠️ Underpowered n=4; 4W/0T/0L; d=4; p=0.063; net +100.0% 🟡 0.48 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
maui-dependency-injection claude-sonnet-5 ➖ Not proven improved n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% 🔴 0.53 Activation-only stop: isolated 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
maui-dependency-injection gpt-5.6-luna ✅ Improved n=5; 5W/0T/0L; d=5; p=0.031; net +100.0% 🟡 0.46 Activation-only stop: isolated 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
maui-safe-area claude-sonnet-5 ⚠️ Underpowered n=4; 4W/0T/0L; d=4; p=0.063; net +100.0% 🟡 0.47 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
maui-safe-area gpt-5.6-luna ⚠️ Underpowered n=4; 3W/0T/1L; d=4; p=0.312; net +50.0% 🟡 0.37 Activation-only stop: isolated 1 failed run Predeclare more independent, discriminating stimuli; repeated runs do not add power.
maui-shell-navigation claude-sonnet-5 ➖ Not proven improved n=5; 4W/1T/0L; d=4; p=0.063; net +80.0% 🔴 0.62 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
maui-shell-navigation gpt-5.6-luna ➖ Not proven improved n=5; 2W/2T/1L; d=3; p=0.500; net +20.0% 🔴 0.62 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
maui-theming claude-sonnet-5 ✅ Improved n=5; 5W/0T/0L; d=5; p=0.031; net +100.0% 🔴 0.75 — Review overfit evidence.
maui-theming gpt-5.6-luna ➖ Not proven improved n=5; 4W/1T/0L; d=4; p=0.063; net +80.0% 🔴 0.56 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
migrate-static-to-wrapper claude-sonnet-5 ✅ Improved n=10; 8W/2T/0L; d=8; p=0.004; net +80.0% 🟡 0.31 — Review overfit evidence.
migrate-static-to-wrapper gpt-5.6-luna ➖ Not proven improved n=10; 6W/3T/1L; d=7; p=0.063; net +50.0% — — Inspect tied or lost stimuli and fix inconsistent skill behavior.
mtp-hot-reload claude-sonnet-5 ✅ Improved n=10; 9W/1T/0L; d=9; p=0.002; net +90.0%; 1 dormancy excluded 🟡 0.37 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
mtp-hot-reload gpt-5.6-luna ⛔ Activation contract failed n=10; 7W/0T/3L; d=10; p=0.172; net +40.0%; 1 dormancy excluded 🟡 0.38 Dormancy contract: 1 unexpected activation(s); Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs Narrow skill routing so the listed off-target scenarios stay dormant.
platform-detection claude-sonnet-5 ✅ Improved n=15; 12W/2T/1L; d=13; p=0.002; net +73.3% 🟡 0.27 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
platform-detection gpt-5.6-luna ✅ Improved n=15; 8W/7T/0L; d=8; p=0.004; net +53.3% — — None.
run-tests claude-sonnet-5 ✅ Improved n=22; 13W/5T/4L; d=17; p=0.025; net +40.9% 🟡 0.37 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
run-tests gpt-5.6-luna ➖ Not proven improved n=22; 6W/13T/3L; d=9; p=0.254; net +13.6% — Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
scaffold-dotnet-test-project claude-sonnet-5 ➖ Not proven improved n=10; 6W/2T/2L; d=8; p=0.145; net +40.0% 🟡 0.21 Activation: isolated 9/10; plugin 9/10 Inspect tied or lost stimuli and fix inconsistent skill behavior.
scaffold-dotnet-test-project gpt-5.6-luna ➖ Not proven improved n=10; 6W/3T/1L; d=7; p=0.063; net +50.0% 🟡 0.30 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
test-anti-patterns claude-sonnet-5 ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5%; 1 dormancy excluded 🟡 0.49 — Review overfit evidence.
test-anti-patterns gpt-5.6-luna ✅ Improved n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%; 1 dormancy excluded — — None.
test-gap-analysis claude-sonnet-5 ➖ Not proven improved n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 2 dormancy excluded 🟡 0.36 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
test-gap-analysis gpt-5.6-luna ✅ Improved n=8; 6W/2T/0L; d=6; p=0.016; net +75.0%; 2 dormancy excluded — — None.
test-smell-detection claude-sonnet-5 ➖ Not proven improved n=10; 5W/3T/2L; d=7; p=0.227; net +30.0% 🟡 0.42 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
test-smell-detection gpt-5.6-luna ➖ Not proven improved n=10; 4W/4T/2L; d=6; p=0.344; net +20.0% — — Inspect tied or lost stimuli and fix inconsistent skill behavior.
test-tagging claude-sonnet-5 ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5%; 4 dormancy excluded 🟡 0.39 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
test-tagging gpt-5.6-luna ✅ Improved n=8; 8W/0T/0L; d=8; p=0.004; net +100.0%; 4 dormancy excluded 🟡 0.32 — Review overfit evidence.
testability-obstacle claude-sonnet-5 ✅ Improved n=9; 8W/0T/1L; d=9; p=0.020; net +77.8% 🔴 0.51 — Review overfit evidence.
testability-obstacle gpt-5.6-luna ➖ Not proven improved n=9; 6W/2T/1L; d=7; p=0.063; net +55.6% 🟡 0.40 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
writing-mstest-tests claude-sonnet-5 ✅ Improved n=15; 8W/6T/1L; d=9; p=0.020; net +46.7% 🟡 0.46 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
writing-mstest-tests gpt-5.6-luna ➖ Not proven improved n=15; 8W/2T/5L; d=13; p=0.291; net +20.0% — Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 4 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⚠️ Underpowered — maui-app-lifecycle (claude-sonnet-5)

Why: Net win +75.0% (3W/1T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +60.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 3W/1T/0L; d=3; p=0.125; net +75.0%

Warnings: Activation-only stop: isolated 1 failed run

Overfit: High (score 0.62)

Repeated-run reliability (not used by the gate): 4 paired runs (3W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Avoid legacy Xamarin.Forms lifecycle methods Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Avoid legacy Xamarin.Forms lifecycle methods: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — maui-app-lifecycle (gpt-5.6-luna)

Why: Net win +75.0% (3W/1T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +45.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 3W/1T/0L; d=3; p=0.125; net +75.0%

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs

Overfit: High (score 0.59)

Repeated-run reliability (not used by the gate): 4 paired runs (3W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Avoid legacy Xamarin.Forms lifecycle methods Eligible +0.0% +0.0% 0/1/0
▲ Window lifecycle event subscription Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Avoid legacy Xamarin.Forms lifecycle methods: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — maui-data-binding (claude-sonnet-5)

Why: Net win +25.0% (2W/1T/1L over 4 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 2W/1T/1L; d=3; p=0.500; net +25.0%

Overfit: High (score 0.55)

Repeated-run reliability (not used by the gate): 4 paired runs (2W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Create and use an IValueConverter Eligible -100.0% -40.0% 0/0/1
= Implement MVVM ViewModel with ObservableObject Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Create and use an IValueConverter: Both answers fully satisfy the requested converter, XAML registration, binding, and fallback behavior. A is marginally more complete and directly usable due to its full XAML context and namespace/accessibility guidance; B is concise and correct but has less complete setup cont...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — maui-data-binding (gpt-5.6-luna)

Why: Net win +100.0% (4W/0T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +55.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 4W/0T/0L; d=4; p=0.063; net +100.0%

Overfit: Moderate (score 0.48)

Repeated-run reliability (not used by the gate): 4 paired runs (4W/0T/0L).

⚠️ Underpowered — maui-safe-area (claude-sonnet-5)

Why: Net win +100.0% (4W/0T/0L over 4 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +85.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5, and this eval won every one of them) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 4W/0T/0L; d=4; p=0.063; net +100.0%

Overfit: Moderate (score 0.47)

Repeated-run reliability (not used by the gate): 4 paired runs (4W/0T/0L).

⚠️ Underpowered — maui-safe-area (gpt-5.6-luna)

Why: Net win +50.0% (3W/0T/1L over 4 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +65.0% across 4 paired run(s) — underpowered (4 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=4; 3W/0T/1L; d=4; p=0.312; net +50.0%

Warnings: Activation-only stop: isolated 1 failed run

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 4 paired runs (3W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Handle notch and status bar safe areas on iOS Eligible +100.0% +100.0% 1/0/0
▼ Keyboard avoidance with safe area for chat UI Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Keyboard avoidance with safe area for chat UI: Response A wins because it provides a more complete and practical answer by explicitly warning about a common mistake—applying SoftInput directly to ScrollView/CollectionView—which is a critical caveat that developers would encounter. While Response B is more concise and uses ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — mtp-hot-reload (gpt-5.6-luna)

Why: Net win +40.0% (7W/0T/3L over 10 preference-eligible stimulus vote(s), sign test p=0.172), mean preference +29.1% across 11 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=10; 7W/0T/3L; d=10; p=0.172; net +40.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 11 paired runs (8W/0T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Decline Test Explorer hot reload integration Excluded (activation contract) +100.0% +40.0% 1/0/0
▼ Enable a configured host that does not react to edits Eligible -100.0% -40.0% 0/0/1
▲ Filter an xUnit v3 hot-reload host to one test Eligible +100.0% +40.0% 1/0/0
▼ Suggest hot reload for failing test in MTP project (SDK 10) Eligible -100.0% -40.0% 0/0/1
▼ Suggest launchSettings.json configuration for hot reload Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Enable a configured host that does not react to edits: While both responses correctly identify the missing TESTINGPLATFORM_HOTRELOAD_ENABLED=1 setting and provide valid relaunch commands, Response A provides a more complete answer by offering both bash and PowerShell examples. This cross-platform guidance is more helpful for users...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.code-testing-generator (claude-sonnet-5)

Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible (sign test p=0.500 > 0.05) — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%

Warnings: Activation: isolated 2/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (3W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Generate a project-wide pytest suite across modules Eligible -100.0% -40.0% 0/0/1
▲ Generate collaborating Go package tests Eligible +100.0% +40.0% 1/0/0
▲ Generate layered Vitest coverage for an async cart Eligible +100.0% +40.0% 1/0/0
▼ Generate project-wide xUnit tests for a .NET library Eligible -100.0% -40.0% 0/0/1
▲ Preserve a classic MSTest project while adding broad coverage Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Generate a project-wide pytest suite across modules: Both are strong, passing, comprehensive test-suite submissions with useful evidence maps. A has a marginally more thorough reported behavioral test set, especially around edge-case boundaries and rollover effects, while B presents the evidence more cleanly.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.code-testing-generator (gpt-5.6-luna)

Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 5 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Generate collaborating Go package tests Eligible -100.0% -40.0% 0/0/1
▼ Preserve a classic MSTest project while adding broad coverage Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Preserve a classic MSTest project while adding broad coverage: While Response B claims more test methods (20 vs 16), Response A delivers significantly higher quality and reliability. Response A provides concrete, verifiable evidence with specific line numbers and actual source code verification shown in the session output. It preserves al...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.test-quality-auditor (claude-sonnet-5)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +10.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 1 dormancy excluded

Warnings: Activation: isolated 1/5; plugin 0/5

Repeated-run reliability (not used by the gate): 6 paired runs (4W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assertion quality analysis Eligible -100.0% -100.0% 0/0/1
▲ Comprehensive test quality audit of weak test suite Eligible +100.0% +40.0% 1/0/0
= Decline request to generate new tests Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose test smells and propose a repair order Eligible +100.0% +40.0% 1/0/0
▲ Identify behavior gaps that existing tests would miss Eligible +100.0% +40.0% 1/0/0
▲ Targeted anti-pattern review Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Assertion quality analysis: A answers the user's question with concrete, test-by-test assessment and actionable assertion-variety guidance. B fails to recover from a file-viewing issue and gives no substantive analysis.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.test-quality-auditor (gpt-5.6-luna)

Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +13.3% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%; 1 dormancy excluded

Warnings: Activation: isolated 1/5; plugin 1/5

Repeated-run reliability (not used by the gate): 6 paired runs (4W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Assertion quality analysis Eligible +100.0% +40.0% 1/0/0
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
▼ Identify behavior gaps that existing tests would miss Eligible -100.0% -40.0% 0/0/1
▲ Targeted anti-pattern review Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Diagnose test smells and propose a repair order: Response A delivers a more comprehensive and practical assessment. It identifies additional coverage gaps beyond the core test smells (RemoveItem, duplicate-ID merging, discount edge cases), explains risks with more behavioral detail, and provides a significantly more granular...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.testability-migration (claude-sonnet-5)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +36.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05) — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%

Warnings: Activation: isolated 0/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Full pipeline: detect statics and recommend migration plan Eligible +100.0% +40.0% 1/0/0
▲ Inventory static dependencies without modifying the project Eligible +100.0% +40.0% 1/0/0
▼ Migrate time dependencies and add deterministic tests Eligible -100.0% -40.0% 0/0/1
▲ Replace filesystem statics without touching unrelated dependencies Eligible +100.0% +40.0% 1/0/0
▲ Targeted request: just migrate DateTime to TimeProvider Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Migrate time dependencies and add deterministic tests: Both implementations satisfy the task well. A is marginally stronger overall because its test suite adds an explicit after-expiry case in addition to the requested exact-boundary coverage. B has a slightly cleaner required-constructor DI shape, but the difference is minor.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.testability-migration (gpt-5.6-luna)

Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +44.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%

Warnings: Activation: isolated 0/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Full pipeline: detect statics and recommend migration plan Eligible +100.0% +40.0% 1/0/0
= Inventory static dependencies without modifying the project Eligible +0.0% +0.0% 0/1/0
▲ Migrate time dependencies and add deterministic tests Eligible +100.0% +40.0% 1/0/0
▲ Replace filesystem statics without touching unrelated dependencies Eligible +100.0% +40.0% 1/0/0
▲ Targeted request: just migrate DateTime to TimeProvider Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Inventory static dependencies without modifying the project: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — assertion-quality (claude-sonnet-5)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +32.5% across 16 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%

Warnings: Activation: isolated 7/8; plugin 8/8; Activation-only stop: isolated 1 failed run

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 16 paired runs (12W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Judge assertion strength in a shallow Jest suite Eligible +0.0% +0.0% 1/0/1
▼ Recognize structural and interaction checks in Jest while flagging vacuous tests Eligible -50.0% -20.0% 0/1/1

Illustrative judge evidence:

  • Judge assertion strength in a shallow Jest suite: A inspected the relevant code and delivered a precise, criterion-complete diagnosis plus actionable replacement assertions. B merely requested file paths despite the files being discoverable and provided no analysis.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — assertion-quality (gpt-5.6-luna)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.3% across 16 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 16 paired runs (9W/6T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Distinguish weak and meaningful assertions in a pytest suite Eligible +0.0% +0.0% 0/2/0
▼ Recognize structural and interaction checks in Jest while flagging vacuous tests Eligible -50.0% -20.0% 0/1/1

Illustrative judge evidence:

  • Distinguish weak and meaningful assertions in a pytest suite: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — code-testing-agent (claude-sonnet-5)

Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.2% across 18 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%

Warnings: Activation: isolated 6/9; plugin 3/9

Overfit: High (score 0.51)

Repeated-run reliability (not used by the gate): 18 paired runs (14W/0T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add focused xUnit tests for one reservation class Eligible +100.0% +40.0% 2/0/0
▼ Expand a healthy existing pytest suite to every ledger boundary Eligible -100.0% -40.0% 0/0/2
▲ Generate a layered Vitest suite for an async shopping cart Eligible +100.0% +40.0% 2/0/0
= Generate a project-wide Go suite across collaborating packages Eligible +0.0% +0.0% 1/0/1
▲ Generate a project-wide pytest suite across multiple modules Eligible +100.0% +40.0% 2/0/0
= Generate project-wide tests for an SDK-style xUnit library Eligible +0.0% +0.0% 1/0/1

Illustrative judge evidence:

  • Expand a healthy existing pytest suite to every ledger boundary: Both are strong, passing suite extensions with Decimal-based boundary assertions. A is marginally better on substantive test completeness, particularly its explicit positive-balance coverage with and without a limit, while B's superior baseline verification does not outweigh t...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — coverage-analysis (claude-sonnet-5)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +20.0% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 4 dormancy excluded

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 24 paired runs (13W/7T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Reconcile a coverage target spread across several members Eligible +0.0% +0.0% 0/2/0
▼ Refactoring safety assessment from coverage data Eligible -100.0% -40.0% 0/0/2
= Stay dormant for behavioral gap analysis Excluded (activation contract) +0.0% +0.0% 1/0/1
▼ Stay dormant for one-member CRAP analysis Excluded (activation contract) -50.0% -20.0% 0/1/1
= Stay dormant for static source-to-test pairing Excluded (activation contract) +0.0% +0.0% 0/2/0

Illustrative judge evidence:

  • Reconcile a coverage target spread across several members: Both responses are correct, complete, self-contained, and directly address the coverage arithmetic, insufficiency of Apply alone, and practical next steps. B adds a correct reconciliation of the one uncovered line outside the listed methods, while A gives a similarly sound ris...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — coverage-analysis (gpt-5.6-luna)

Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.5% across 24 paired run(s), 4 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 4 dormancy excluded

Warnings: Activation: isolated 7/8; plugin 7/8

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 24 paired runs (14W/8T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Coverage plateau diagnosis Eligible -50.0% -20.0% 0/1/1
= Distinguish partially covered branches from covered lines Eligible +0.0% +0.0% 1/0/1
= Project-wide coverage analysis with existing Cobertura data Eligible +0.0% +0.0% 0/2/0

Illustrative judge evidence:

  • Coverage plateau diagnosis: Response A delivers more actionable guidance with specific test recommendations (Criterion 4: much-better), while Response B provides clearer diagnostic structure (Criterion 2: slightly-better). Both fail to calculate the impact of fixing CalculateGpa (Criterion 3: tie) and bo...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — crap-score (claude-sonnet-5)

Why: Net win +44.4% (5W/3T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +17.8% across 9 paired run(s) — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 5W/3T/1L; d=6; p=0.109; net +44.4%

Overfit: Moderate (score 0.27)

Repeated-run reliability (not used by the gate): 9 paired runs (5W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Generate coverage then compute CRAP score Eligible +0.0% +0.0% 0/1/0
▼ Recognize when complexity alone blocks the CRAP threshold Eligible -100.0% -40.0% 0/0/1
= Recompute complexity instead of trusting a stale source comment Eligible +0.0% +0.0% 0/1/0
= Report a fully covered method at its complexity floor Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Generate coverage then compute CRAP score: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — detect-static-dependencies (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Avoid false positives when ambient resources are already abstracted Eligible -100.0% -40.0% 0/0/1
= Detect statics inside lambda expressions and LINQ queries Eligible +0.0% +0.0% 0/1/0
= Detect time-related statics and recommend TimeProvider Eligible +0.0% +0.0% 0/1/0
▼ Stay dormant for Python timezone review Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Avoid false positives when ambient resources are already abstracted: Both responses reached the correct conclusion (0 dependencies requiring new seams) and properly identified all injected dependencies and deterministic code. However, Response A provides accurate line numbers (9, 10, 11) while Response B reports incorrect line numbers (18, 21, ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — generate-testability-wrappers (claude-sonnet-5)

Why: Net win -28.6% (2W/1T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -0.8% across 24 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/1T/4L; d=6; p=0.344; net -28.6%; 1 dormancy excluded

Warnings: Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.46)

Repeated-run reliability (not used by the gate): 24 paired runs (11W/3T/10L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Generate TimeProvider adoption for DateTime.UtcNow Eligible +0.0% +0.0% 1/1/1
▼ Generate a minimal process runner wrapper Eligible -33.3% -53.3% 1/0/2
▼ Generate custom Environment wrapper Eligible -33.3% -13.3% 1/0/2
▼ Generate only the console members a prompt uses Eligible -100.0% -80.0% 0/0/3
▼ Make time controllable in a library that has no DI container Eligible -33.3% -13.3% 1/0/2
▲ Recommend System.IO.Abstractions for file system calls Eligible +100.0% +80.0% 3/0/0

Illustrative judge evidence:

  • Generate TimeProvider adoption for DateTime.UtcNow: Both deliver the right .NET 10 abstraction, compile, and include working deterministic tests. A is stronger because it demonstrates advancing FakeTimeProvider and verifying changed time-dependent behavior, directly satisfying the more demanding test criterion. B's explicit pro...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — maui-collectionview (claude-sonnet-5)

Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%

Warnings: Activation-only stop: isolated 1 failed run

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Avoid ListView and ViewCell mistakes Eligible +0.0% +0.0% 0/1/0
= Basic CollectionView with data binding and DataTemplate Eligible +0.0% +0.0% 0/1/0
▼ Selection and pull-to-refresh with CollectionView Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Avoid ListView and ViewCell mistakes: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — maui-collectionview (gpt-5.6-luna)

Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Avoid ListView and ViewCell mistakes Eligible +0.0% +0.0% 0/1/0
▼ Grid layout with CollectionView Eligible -100.0% -40.0% 0/0/1
= ItemSizingStrategy placement for uniform items Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Avoid ListView and ViewCell mistakes: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — maui-dependency-injection (claude-sonnet-5)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Warnings: Activation-only stop: isolated 2 failed runs

Overfit: High (score 0.53)

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Avoid AddScoped pitfall in MAUI Eligible -100.0% -40.0% 0/0/1
▲ Diagnose a page whose injected dependencies are missing Eligible +100.0% +40.0% 1/0/0
= Register services with correct lifetimes in MauiProgram.cs Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Avoid AddScoped pitfall in MAUI: Both give useful, concise EF Core guidance including a context factory, transient lifetime, and explicit scopes. A is better because its explanation of MAUI scoped lifetime is accurate, whereas B incorrectly asserts automatic per-window IServiceScopes.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — maui-shell-navigation (claude-sonnet-5)

Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +32.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%

Overfit: High (score 0.62)

Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Set up Shell navigation with tabs and flyout Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Set up Shell navigation with tabs and flyout: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — maui-shell-navigation (gpt-5.6-luna)

Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%

Overfit: High (score 0.62)

Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose common Shell navigation mistakes Eligible +0.0% +0.0% 0/1/0
▼ Set up Shell navigation with tabs and flyout Eligible -100.0% -40.0% 0/0/1
= Stable routes for deep linking into tabs Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose common Shell navigation mistakes: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — maui-theming (gpt-5.6-luna)

Why: Net win +80.0% (4W/1T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +44.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/1T/0L; d=4; p=0.063; net +80.0%

Overfit: High (score 0.56)

Repeated-run reliability (not used by the gate): 5 paired runs (4W/1T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect and respond to system theme changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect and respond to system theme changes: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — migrate-static-to-wrapper (gpt-5.6-luna)

Why: Net win +50.0% (6W/3T/1L over 10 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +26.0% across 10 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=10; 6W/3T/1L; d=7; p=0.063; net +50.0%

Repeated-run reliability (not used by the gate): 10 paired runs (6W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Migrate a static helper class without breaking its callers Eligible +0.0% +0.0% 0/1/0
= Migrate only in scoped files, leaving others untouched Eligible +0.0% +0.0% 0/1/0
▼ Preserve DateTimeOffset values during TimeProvider migration Eligible -100.0% -40.0% 0/0/1
= Preserve local calendar semantics when migrating DateTime.Now Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Migrate a static helper class without breaking its callers: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — run-tests (gpt-5.6-luna)

Why: Net win +13.6% (6W/13T/3L over 22 preference-eligible stimulus vote(s), sign test p=0.254), mean preference +8.2% across 22 paired run(s) — not credible (sign test p=0.254 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=22; 6W/13T/3L; d=9; p=0.254; net +13.6%

Warnings: Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 1 failed run

Repeated-run reliability (not used by the gate): 22 paired runs (6W/13T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Collect a crash dump on an MTP project (SDK 9) Eligible -100.0% -100.0% 0/0/1
= Collect coverage on an SDK 9 MTP bridge Eligible +0.0% +0.0% 0/1/0
= Collect coverage with VSTest Eligible +0.0% +0.0% 0/1/0
= Enable a diagnostic log for a VSTest project Eligible +0.0% +0.0% 0/1/0
= Enable diagnostic logs for a native MTP project Eligible +0.0% +0.0% 0/1/0
= Filter NUnit tests on an SDK 9 MTP bridge Eligible +0.0% +0.0% 0/1/0
= Filter TUnit tests by class using treenode-filter Eligible +0.0% +0.0% 0/1/0
= Filter one NUnit class on VSTest Eligible +0.0% +0.0% 0/1/0
▼ Filter xUnit v3 tests by class on MTP Eligible -100.0% -40.0% 0/0/1
▼ Filter xUnit v3 tests by class pattern and trait using query filter language Eligible -100.0% -100.0% 0/0/1
= Generate TRX from a VSTest project Eligible +0.0% +0.0% 0/1/0
= MTP project on SDK 10 passes args directly Eligible +0.0% +0.0% 0/1/0
= Run one VSTest invocation without rebuilding Eligible +0.0% +0.0% 0/1/0
= Run tests in a VSTest MSTest project Eligible +0.0% +0.0% 0/1/0
= Select one target framework in a multi-targeted project Eligible +0.0% +0.0% 0/1/0
= Use the repository unit-test entry point Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Collect a crash dump on an MTP project (SDK 9): Response A provides the correct, documented command using standard dotnet CLI options (--blame-crash --blame-crash-dump-type full) for collecting crash dumps from test runs. Response B provides an incorrect command using a non-existent --crashdump parameter with inaccurate...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 5 results are in Full Results.

Details for 26 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1202 in dotnet/skills, download eval artifacts with gh run download 36708764237 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://github.com/raw/dotnet/skills/b4fba51282c6a0f28029015f1f094256da2e4c31/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file .NET Pull requests that update .NET code ready-to-merge PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants