Skip to content

feat(genai-openai): instrument Responses.parse and AsyncResponses.parse - #786

Open
RichardoMrMu wants to merge 18 commits into
open-telemetry:mainfrom
RichardoMrMu:responses-parse-instrumentation
Open

RichardoMrMu wants to merge 18 commits into
open-telemetry:mainfrom
RichardoMrMu:responses-parse-instrumentation

Conversation

@RichardoMrMu

Copy link
Copy Markdown

What changed

Instruments Responses.parse and AsyncResponses.parse on the OpenAI SDK, so
applications using the Responses API structured-output helper emit a GenAI span.

  • __init__.py: wrap Responses.parse / AsyncResponses.parse with the
    existing responses_create / async_responses_create wrappers, guarded by a
    _is_responses_parse_supported() availability check (older SDKs lack it), and
    unwrap them in _uninstrument.
  • Adds sync + async tests.

Fixes #659.

On the operation mapping (answering @eternalcuriouslearner)

Responses.parse maps to the same inference operation as Responses.create —
there is no separate operation. parse() is just the structured-output helper:
it sends the same /v1/responses request (with a text_format) and returns a
ParsedResponse, whose telemetry-relevant fields (model, usage, output,
finish reason) are identical to a Response. The existing responses_create
wrapper already declares ResponseResult = Union[ParsedResponse[Any], Response]
and its extract_params(**kwargs) absorbs the extra text_format kwarg, so
reusing that wrapper produces a correct span with no new logic.

This mirrors exactly how the package already instruments
chat.completions.parse — it reuses the chat_completions_create_v_new
wrapper for the same reason. The one difference the issue calls out is that
Responses.parse does not delegate to the instrumented Responses.create
internally, so instrumenting create alone does not cover it; hence the
separate wrap here.

Scope

Minimal and additive: reuses the existing wrappers and matches the existing
chat.completions.parse precedent — no change to span attributes, operation
naming, or the create path. Streaming parse is out of scope (parse is a
structured-output convenience over a non-streaming request).

Tests

  • test_responses_parse_basic / test_async_responses_parse_basic: call
    parse(..., text_format=...) and assert a single Responses GenAI span with the
    expected attributes. They reuse the existing responses_create_basic
    cassettes (parse issues the same /v1/responses request, matched by VCR),
    following the cassette-reuse pattern already used in this test module.
  • test_responses_parse_wrapping_lifecycle: asserts instrument() wraps both
    parse methods and uninstrument() restores them.
  • All tests are guarded to skip on SDKs without Responses.parse.

Verified locally that a parse() call under instrumentation emits exactly one
GenAI span carrying the response id (via a mocked SDK response, no network), and
that wrap/unwrap round-trips cleanly. ruff format --check / ruff check clean.

@opentelemetry-pr-dashboard

opentelemetry-pr-dashboard Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Pull request dashboard status

Waiting on the author · refreshed 2026-09-27 05:49 UTC

Respond to 2 review items (e.g. link a commit, explain why not, ask a follow-up):

  • Inline threads: 1, 2
Status above doesn't look right?
  • Just replied or pushed? Anything around or after the refresh time above may not be picked up yet — give it a few minutes.
  • Should this be with reviewers? Comment /dashboard route:reviewers to route it to them.
  • Anything wrong — including the routing? Report it with what you expected; it helps us improve the dashboard.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One or more issues must be addressed before approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity · 1 Medium severity

Open (3)
What changed in this PR

Adds OpenAI Responses API structured-output instrumentation for synchronous and asynchronous parse methods.

Changes:

  • Wraps Responses.parse and AsyncResponses.parse.
  • Adds lifecycle and telemetry tests.
  • Reuses Responses create extraction logic.
File Description
instrumentation/​opentelemetry-instrumentation-genai-openai/​tests/​test_responses.py Updated as part of this pull request.
instrumentation/​opentelemetry-instrumentation-genai-openai/​tests/​test_async_responses.py Updated as part of this pull request.
instrumentation/​opentelemetry-instrumentation-genai-openai/​src/​opentelemetry/​instrumentation/​genai/​openai/​__init__.py Updated as part of this pull request.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +1624 to +1626
with vcr.use_cassette(
"test_async_responses_create_basic[content_mode0].yaml"
):

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — this was addressed in the follow-up commit on this PR. Both parse tests now use dedicated cassettes (test_responses_parse_basic[content_mode0].yaml and the async equivalent) whose output_text is valid CalendarEvent JSON ({"name":"science fair",...}) instead of the shared "This is a test." create cassette, so the SDK's Pydantic post-parser succeeds before the span assertions. The two parse tests pass.

"""
_skip_if_not_latest()

with vcr.use_cassette("test_responses_create_basic[content_mode0].yaml"):

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — this was addressed in the follow-up commit on this PR. Both parse tests now use dedicated cassettes (test_responses_parse_basic[content_mode0].yaml and the async equivalent) whose output_text is valid CalendarEvent JSON ({"name":"science fair",...}) instead of the shared "This is a test." create cassette, so the SDK's Pydantic post-parser succeeds before the span assertions. The two parse tests pass.

wrap_function_wrapper(
"openai.resources.responses.responses",
"Responses.parse",
responses_create(handler),

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is covered in the latest version of the PR. extract_params() now accepts text_format and response_extractors._extract_output_type_from_text_format() maps a Pydantic/dataclass type (what Responses.parse(text_format=...) receives) to gen_ai.output.type=json, mirroring the chat.completions.parse path. The mapping is asserted by both parse tests via _assert_request_attrs(span, output_type="json"), and test_responses_create_output_type_unchanged_by_parse keeps a plain create() (no text_format) from spuriously recording an output type.

Responses.parse takes the caller's Pydantic model as text_format and only
converts it to a JSON-schema text.format inside the SDK, after the wrapper has
read the request kwargs, so reuse of the responses_create wrapper dropped the
structured-output format and never emitted gen_ai.output.type (Copilot review
on open-telemetry#786). Map text_format like the chat.completions.parse path: a
structured-output type is reported as json. Only format metadata is recorded,
never the caller's schema.
… keep create unchanged

- assert parse spans report gen_ai.output.type=json for text_format
- add a regression test that Responses.create still omits the attribute
- reuse the existing parse/create cassettes
… write

The previous push replaced test_responses.py with a re-generated variant of
the file. Restore the exact on-disk content, which keeps the only intended
change: the gen_ai.output.type assertion for Responses.parse and the
create-behaviour regression test.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

Instrument Responses.parse and AsyncResponses.parse

2 participants