镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

fix(handlers): SAO-17638 LLM span double-encoding and workflow input - #259

Merged
etserend merged 5 commits into
mainfrom
fix/sao-17638-openai-agents-tracing-processor
Oct 8, 2026
Merged

etserend merged 5 commits into
mainfrom
fix/sao-17638-openai-agents-tracing-processor

Conversation

@etserend

@etserend etserend commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes two issues in SplunkAOTracingProcessor (OpenAI Agents SDK) from SAO-17638:

  • LLM span messages were double-encoded. _extract_llm_data serialized
    GenerationSpanData input and output to a JSON string, but LoggedLlmSpan
    already converts message lists, so the whole conversation collapsed into one
    message. The lists now pass through; for output, the first choice is kept
    because LoggedLlmSpan accepts a single output message.
  • Workflow/agent spans showed "<Type> Step" as input. Agent, turn and
    workflow spans have no input of their own. They now use the latest user message
    in the first LLM call's input, on both the incremental (default) and
    ingestion-hook paths.

Scope

Chat Completions (GenerationSpanData), which the reported setup uses. Responses
API (ResponseSpanData) handling is unchanged.

Not changed (by design)

  • Intermediate Turn spans show the tool result: in that Turn the LLM only calls a tool, and the answer comes in the next Turn.
  • Duplicate "Agent workflow" rows come from the OpenAI Agents SDK's own task span, not from this SDK.

Test plan

  • poetry run pytest tests/test_openai_agents.py tests/test_openai_agents_utils.py -n 0
    (2 network-dependent VCR tests also fail locally on main)
  • ruff check and ruff format --check on changed files, invoke type-check
  • Live repro through LiteLLM and Azure OpenAI: chat spans export
    system/user/assistant/tool messages; all workflow spans show the user prompt

🤖 Generated with Claude Code

@etserend
etserend marked this pull request as ready for review October 7, 2026 02:22

@opikalova opikalova left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 This review was generated by the Astra agent (claude-opus-5). It may contain mistakes.

Verdict: request_changes — The I/O fix skips the default Responses API path, and the new user-input helper corrupts structured message content.

General Comments

  • 🟠 major (testing): The new tests only drive generation_span. The OpenAI Agents SDK selects OpenAIResponsesModel by default, so real traces carry ResponseSpanData. No test drives a response_span with a list[dict] input through the processor and asserts the exported gen_ai.input.messages. Please add that case. Without it, the regression this change targets stays untested on the path most users hit.

  • 🟡 minor (other): ruff format --check fails on all 4 changed files. The repository runs ruff-format as a pre-commit hook, so please run poetry run pre-commit run --files <changed-files>. The reported items are:

  • src/splunk_ao/handlers/openai_agents/handler.py line 36: a stray third blank line after _logger.

  • src/splunk_ao/handlers/openai_agents/handler.py line 359: line over 120 characters.

  • src/splunk_ao/utils/openai_agents.py line 291: line over 120 characters.

  • tests/test_openai_agents.py lines 138, 170, 193, 205.

  • 🟡 minor (documentation): CHANGELOG.md keeps an empty [Unreleased] section. This change alters the exported content of OpenAI Agents LLM and workflow spans, which users observe directly. Please add a Fixed entry under [Unreleased].

Comment thread src/splunk_ao/utils/openai_agents.py
Comment thread src/splunk_ao/utils/openai_agents.py Outdated
Comment thread src/splunk_ao/utils/openai_agents.py Outdated
Comment thread src/splunk_ao/handlers/openai_agents/handler.py
Comment thread src/splunk_ao/handlers/openai_agents/handler.py Outdated
Comment thread tests/test_openai_agents.py Outdated
Comment thread tests/test_openai_agents.py Outdated
@etserend etserend changed the title fix(handlers): fix LLM span I/O and workflow input in SplunkAOTracingProcessor fix(handlers): SAO-17638 fix LLM span I/O and workflow input in SplunkAOTracingProcessor Oct 7, 2026
@etserend etserend changed the title fix(handlers): SAO-17638 fix LLM span I/O and workflow input in SplunkAOTracingProcessor fix(handlers): SAO-17638 fix OpenAI Agents span I/O and workflow rollup Oct 7, 2026
@etserend
etserend force-pushed the fix/sao-17638-openai-agents-tracing-processor branch from 2d26692 to 367269c Compare October 7, 2026 18:49
@etserend etserend changed the title fix(handlers): SAO-17638 fix OpenAI Agents span I/O and workflow rollup fix(handlers): SAO-17638 LLM span double-encoding and workflow input Oct 7, 2026
etserend and others added 3 commits October 7, 2026 16:22
…plunkAOTracingProcessor (SAO-17638)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Claude <noreply@anthropic.com>
…SAO-17638)

Use the same user-message fallback in _update_owned_root, drop drive-by
whitespace and comment edits, apply ruff format, and add a CHANGELOG entry.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@etserend
etserend force-pushed the fix/sao-17638-openai-agents-tracing-processor branch from 367269c to 5c0dcf1 Compare October 7, 2026 21:23
@etserend

etserend commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up: the valid review points outside this PR's scope are tracked in SAO-18193:

  • Responses API path (the SDK default): LLM input and output are still exported as one JSON blob, and workflow spans still show "Workflow Step".
  • Multimodal user input (for example text plus an image), on both the Chat Completions and Responses paths.

@opikalova
opikalova requested a review from pradystar October 8, 2026 10:06

@opikalova opikalova left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 This review was generated by the Astra agent (claude-opus-5-5). It may contain mistakes.

Verdict: approve — The scoped Chat Completions fix is correct and the earlier blocking points are fixed or tracked in SAO-18193. Two minor gaps remain.

Comment thread src/splunk_ao/handlers/openai_agents/handler.py
Comment thread CHANGELOG.md Outdated
opikalova
opikalova previously approved these changes Oct 8, 2026
@pradystar

Copy link
Copy Markdown
Collaborator

Changes look good, but there is one regression introduced. See below for details -

Normalize all supported output sequences - GenerationSpanData.output accepts Sequence[Mapping[str, Any]], but this branch only unwraps lists. A custom model supplying tuple output, such as ({"role": "assistant", "content": "Hello"},), now passes that tuple to LoggedLlmSpan, whose output validator rejects it. The handler discards the completed LLM span, losing its output and token metrics. Previously the tuple was serialized to a valid string. Please normalize supported sequences before selecting the first output message.

This affects custom models using tuple output; OpenAI's built-in Chat Completions model supplies lists. Since this is a regression and affects GenerationSpanData I would fix this as part of current changes instead of follow up.

…7638)

GenerationSpanData.output is typed Sequence[Mapping], so a custom model may
pass a tuple. Only lists were unwrapped, so a tuple reached LoggedLlmSpan,
failed validation, and the LLM span was dropped. Unwrap any sequence and keep
the first choice. Add the reviewer's tests for tuple output and for input
assigned after the span starts, as OpenAIChatCompletionsModel does.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…OG (SAO-17638)

OpenAIChatCompletionsModel assigns the generation input after the span starts,
so the workflow input fallback is only set at span end. Add a test that mirrors
that and checks the workflow span shows the user message; it fails if the
span-end fallback is removed. Scope the CHANGELOG entry to Chat Completions
models, since the Responses API path is unchanged here.
@etserend

etserend commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Fixed: _extract_llm_data unwraps any non-string Sequence (tuples included) before taking the first choice. Tuple output is exported as an assistant message again instead of failing LoggedLlmSpan validation.

@etserend
etserend merged commit e51ae74 into main Oct 8, 2026
13 checks passed
@etserend
etserend deleted the fix/sao-17638-openai-agents-tracing-processor branch October 8, 2026 21:38
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 8, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants