test(e2e): Add node-anthropic-send-to-sentry test app - #24856
RulaKhaled wants to merge 1 commit into
Conversation
An e2e app that uses the Anthropic integration the way a user does and sends the data to a real Sentry project. A plain Sentry.init on an express app, no tunnel, and the stock @anthropic-ai/sdk client making real requests through OpenRouter's Anthropic-compatible endpoint: a message, a streamed message, one through the messages.stream() helper, and a forced tool use. The tests read the gen_ai spans back through the Sentry API, like the other *-send-to-sentry apps, and check op, name, origin, status, token usage, stop reasons and the recorded prompts and answers. A latest variant runs the same suite against @anthropic-ai/sdk@latest. The stream-helper test asserts a single gen_ai.chat span; that is what caught the duplicate span on current SDK versions fixed in the commit this sits on. The streamed tests do not assert input token usage: OpenRouter reports 0 in message_start and the real count in message_delta, which the streaming instrumentation does not read. Closes #24748 Co-Authored-By: Claude Fable 5.1 <[email protected]>
size-limit report 📦
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 3b49445. Configure here.
|
|
||
| // The request span is the segment; it ends last, so it can land after its children. | ||
| await expect.poll(() => isModelSpanUnderRequestSpan(traceId, span.event_id!), EVENT_POLLING_OPTIONS).toBe(true); | ||
| }); |
There was a problem hiding this comment.
Test timeout too short for polls
Medium Severity
The first test runs two sequential 180s Sentry polls, but the Playwright timeout is only 210s. If the gen_ai.chat span becomes queryable late, the follow-up parent-span poll is cut short and the test fails even when the http.server span would still arrive.
Additional Locations (2)
Triggered by project rule: PR Review Guidelines for Cursor Bot
Reviewed by Cursor Bugbot for commit 3b49445. Configure here.
| return undefined; | ||
| } | ||
|
|
||
| if (response.status === 401 || response.status === 403) { |
There was a problem hiding this comment.
Bug: The waitForSpan poll can succeed prematurely because fetchSpanAttributes may return a truthy empty object {}, causing subsequent assertions to fail.
Severity: MEDIUM
Suggested Fix
Update the condition in waitForSpan to ensure the attributes object is not empty before considering the poll successful. A suggested implementation is: found = span && attributes && Object.keys(attributes).length > 0 ? { span, attributes } : undefined;.
Prompt for AI Agent
Review the code at the location below. A potential bug has been identified by an AI
agent. Verify if this is a real issue. If it is, propose a fix; if not, explain why it's
not valid.
Location:
dev-packages/e2e-tests/test-applications/node-anthropic-send-to-sentry/tests/utils/sentry-api.ts#L113
Potential issue: The `fetchSpanAttributes` function can return an empty object `{}` if
the Sentry API returns an empty `attributes` array, which can happen due to race
conditions during data ingestion. In the `waitForSpan` function, the truthiness check
`span && attributes` evaluates to true because an empty object is truthy in JavaScript.
This causes the polling mechanism to succeed prematurely. Consequently, subsequent
assertions that expect specific properties on the `attributes` object fail with
misleading errors, as they are attempting to access properties on an empty object,
leading to flaky tests.
Also affects:
dev-packages/e2e-tests/test-applications/node-anthropic-send-to-sentry/tests/utils/sentry-api.ts:89~89
Did we get this right? 👍 / 👎 to inform future reviews.


An e2e app that uses the Anthropic integration the way a user does and sends the data to a real Sentry project, per #24748. Same shape as the OpenAI one in #24824.
Stacked on #24855. The
latestvariant needs that fix; this PR targets its branch and retargets todeveloponce it merges.The app. A plain
Sentry.initon an express app, preloaded withnode --import, no tunnel. The stock@anthropic-ai/sdkclient points at OpenRouter's Anthropic-compatible endpoint (bearer token viaauthToken) and four routes make one real request each: a message, a streamed message viacreate({ stream: true }), one via themessages.stream()helper, and a forced tool use.The tests. Each one polls Sentry for the request's
gen_ai.chatspan through the organization trace endpoint, then reads its attributes through the project trace-items endpoint, the same way the other*-send-to-sentryapps look events up.tests/utils/sentry-api.tsis the node-express copy plusfetchSpanAttributes. Asserted per span: op, name, origin, status, provider, request and response model, response id, output and total tokens, and that the prompts and answers arrived. The stream tests also check stop reasons and the streaming flag, and the helper test checks there is exactly onegen_ai.chatspan in the trace. The plain-message test checks the span sits below thehttp.serverspan.Variants.
@anthropic-ai/sdkis pinned to 0.63.0, the version the integration suite uses. Thenode-anthropic-send-to-sentry (latest)variant runs the same suite against@anthropic-ai/sdk@latest. The app is optional, like the other send-to-sentry apps.What the
latestvariant found. On 0.129 the stream helper produced two nestedgen_ai.chatspans. The SDK lowercased the header the integration uses to skip the helper's internalcreate; fixed in #24855.Two things the tests do not assert, on purpose.
input_tokens: 0inmessage_startand the real count inmessage_delta; the streaming instrumentation reads input usage only frommessage_start, so the stored value is 0. The Anthropic API also carries cumulative usage onmessage_delta, so reading it there would be a small improvement. Separate from this PR.gen_ai.response.tool_callsandgen_ai.request.max_tokens, which ingest does not store under those names (same as for OpenAI).Verified. Both variants pass against a real project: 4 of 4 on 0.63.0 and, with #24855 applied, 4 of 4 on 0.129.0.
Closes #24748
🤖 Generated with Claude Code