feat(genai-openai): instrument Responses.parse and AsyncResponses.parse - #786
RichardoMrMu wants to merge 18 commits into
Conversation
Pull request dashboard statusWaiting on the author · refreshed 2026-09-27 05:49 UTC Respond to 2 review items (e.g. link a commit, explain why not, ask a follow-up): Status above doesn't look right?
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
One or more issues must be addressed before approval.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 2
Open (3)
What changed in this PR
Adds OpenAI Responses API structured-output instrumentation for synchronous and asynchronous parse methods.
Changes:
- Wraps
Responses.parseandAsyncResponses.parse. - Adds lifecycle and telemetry tests.
- Reuses Responses create extraction logic.
| File | Description |
|---|---|
| instrumentation/opentelemetry-instrumentation-genai-openai/tests/test_responses.py | Updated as part of this pull request. |
| instrumentation/opentelemetry-instrumentation-genai-openai/tests/test_async_responses.py | Updated as part of this pull request. |
| instrumentation/opentelemetry-instrumentation-genai-openai/src/opentelemetry/instrumentation/genai/openai/__init__.py | Updated as part of this pull request. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| with vcr.use_cassette( | ||
| "test_async_responses_create_basic[content_mode0].yaml" | ||
| ): |
There was a problem hiding this comment.
Good catch — this was addressed in the follow-up commit on this PR. Both parse tests now use dedicated cassettes (test_responses_parse_basic[content_mode0].yaml and the async equivalent) whose output_text is valid CalendarEvent JSON ({"name":"science fair",...}) instead of the shared "This is a test." create cassette, so the SDK's Pydantic post-parser succeeds before the span assertions. The two parse tests pass.
| """ | ||
| _skip_if_not_latest() | ||
|
|
||
| with vcr.use_cassette("test_responses_create_basic[content_mode0].yaml"): |
There was a problem hiding this comment.
Good catch — this was addressed in the follow-up commit on this PR. Both parse tests now use dedicated cassettes (test_responses_parse_basic[content_mode0].yaml and the async equivalent) whose output_text is valid CalendarEvent JSON ({"name":"science fair",...}) instead of the shared "This is a test." create cassette, so the SDK's Pydantic post-parser succeeds before the span assertions. The two parse tests pass.
| wrap_function_wrapper( | ||
| "openai.resources.responses.responses", | ||
| "Responses.parse", | ||
| responses_create(handler), |
There was a problem hiding this comment.
This is covered in the latest version of the PR. extract_params() now accepts text_format and response_extractors._extract_output_type_from_text_format() maps a Pydantic/dataclass type (what Responses.parse(text_format=...) receives) to gen_ai.output.type=json, mirroring the chat.completions.parse path. The mapping is asserted by both parse tests via _assert_request_attrs(span, output_type="json"), and test_responses_create_output_type_unchanged_by_parse keeps a plain create() (no text_format) from spuriously recording an output type.
Responses.parse takes the caller's Pydantic model as text_format and only converts it to a JSON-schema text.format inside the SDK, after the wrapper has read the request kwargs, so reuse of the responses_create wrapper dropped the structured-output format and never emitted gen_ai.output.type (Copilot review on open-telemetry#786). Map text_format like the chat.completions.parse path: a structured-output type is reported as json. Only format metadata is recorded, never the caller's schema.
… keep create unchanged - assert parse spans report gen_ai.output.type=json for text_format - add a regression test that Responses.create still omits the attribute - reuse the existing parse/create cassettes
… write The previous push replaced test_responses.py with a re-generated variant of the file. Restore the exact on-disk content, which keeps the only intended change: the gen_ai.output.type assertion for Responses.parse and the create-behaviour regression test.


What changed
Instruments
Responses.parseandAsyncResponses.parseon the OpenAI SDK, soapplications using the Responses API structured-output helper emit a GenAI span.
__init__.py: wrapResponses.parse/AsyncResponses.parsewith theexisting
responses_create/async_responses_createwrappers, guarded by a_is_responses_parse_supported()availability check (older SDKs lack it), andunwrap them in
_uninstrument.Fixes #659.
On the operation mapping (answering @eternalcuriouslearner)
Responses.parsemaps to the same inference operation asResponses.create—there is no separate operation.
parse()is just the structured-output helper:it sends the same
/v1/responsesrequest (with atext_format) and returns aParsedResponse, whose telemetry-relevant fields (model, usage, output,finish reason) are identical to a
Response. The existingresponses_createwrapper already declares
ResponseResult = Union[ParsedResponse[Any], Response]and its
extract_params(**kwargs)absorbs the extratext_formatkwarg, soreusing that wrapper produces a correct span with no new logic.
This mirrors exactly how the package already instruments
chat.completions.parse— it reuses thechat_completions_create_v_newwrapper for the same reason. The one difference the issue calls out is that
Responses.parsedoes not delegate to the instrumentedResponses.createinternally, so instrumenting
createalone does not cover it; hence theseparate wrap here.
Scope
Minimal and additive: reuses the existing wrappers and matches the existing
chat.completions.parseprecedent — no change to span attributes, operationnaming, or the create path. Streaming
parseis out of scope (parse is astructured-output convenience over a non-streaming request).
Tests
test_responses_parse_basic/test_async_responses_parse_basic: callparse(..., text_format=...)and assert a single Responses GenAI span with theexpected attributes. They reuse the existing
responses_create_basiccassettes (parse issues the same
/v1/responsesrequest, matched by VCR),following the cassette-reuse pattern already used in this test module.
test_responses_parse_wrapping_lifecycle: assertsinstrument()wraps bothparsemethods anduninstrument()restores them.Responses.parse.Verified locally that a
parse()call under instrumentation emits exactly oneGenAI span carrying the response id (via a mocked SDK response, no network), and
that wrap/unwrap round-trips cleanly.
ruff format --check/ruff checkclean.