Skip to content

Sentry not logging traces for FastAPI's StreamingResponse endpoint #5002

Description

@Macok

Environment

SaaS (https://sentry.io/)

Steps to Reproduce

I initialised sentry_sdk with

sentry_sdk.init(
        dsn=os.environ['SENTRY_DSN'],
        environment=os.getenv("SENTRY_ENV", "production"),
        send_default_pii=True,
        traces_sample_rate=1.0,
    )

In the Trace Explorer, I can see traces from all endpoints of my app, but not below one:

@app.post("/generate_summary")
def generate_summary(params: GenerateSummaryParams,
                     user: str = Depends(get_current_user)):
    conversation = fetch_conversation(user, params.conversation_id)
    return StreamingResponse(do_generate_summary(conversation, params.doc_ids, user),
                             media_type="text/event-stream")

I tried to manually add transaction inside do_generate_summary() generator like below:

    with sentry_sdk.start_transaction(name="my_transaction"):
        with sentry_sdk.start_span(name="my_span"):
            ...

With above code, something even more strange happens. Now, I can see transactions created by sentry_sdk for my /generate_summary endpoint, but my_transaction transaction and my_span span are missing!

Expected Result

Traces for FastAPI endpoints returning a StreamingResponse should be sent to Sentry.

Actual Result

Traces for FastAPI endpoints returning a StreamingResponse are not being sent to Sentry.

Version

sentry-sdk==2.42.1
fastapi==0.119.1
Python 3.12.12

Activity

  1. linear commented on Oct 23, 2025

    @linear
  2. getsantry commented on Oct 23, 2025

    @getsantry

    Assigning to @getsentry/support for routing ⏲️

  3. moved this to Waiting for: Support in GitHub Issues with 👀 3on Oct 23, 2025
  4. moved this from Waiting for: Support to Waiting for: Product Owner in GitHub Issues with 👀 3on Oct 23, 2025
  5. linear commented on Oct 23, 2025

    @linear
  6. alexander-alderman-webb commented on Oct 24, 2025

    @alexander-alderman-webb
    Contributor

    Hi @Macok,

    Thank for your report. The last comment you made is really useful. Based on the remark, it sounds like you are able to get transactions for endpoints returning StreamingResponse, but not together with your do_generate_summary() function.

    With above code, something even more strange happens. Now, I can see transactions created by sentry_sdk for my /generate_summary endpoint, but my_transaction transaction and my_span span are missing!

    You could be reaching size limits on the ingestion. To confirm can you add debug=True when initializing Sentry, like so

    sentry_sdk.init(
          dsn=os.environ['SENTRY_DSN'],
          environment=os.getenv("SENTRY_ENV", "production"),
          send_default_pii=True,
          traces_sample_rate=1.0,
    +     debug=True,
      )

    Then you should see a message like the following if the serialized transaction is too large:

    ERROR: Unexpected status code: 400 (body: b'{"detail":"envelope exceeded size limits for type \'event\' (https://develop.sentry.dev/sdk/envelopes/#size-limits)"}')
  7. moved this from Waiting for: Product Owner to No status in GitHub Issues with 👀 3on Oct 24, 2025
  8. Macok commented on Oct 24, 2025

    @Macok
    Author

    That was it! Thanks for quick help!

    Inside do_generate_summary, I'm streaming response from OpenAI chunk by chunk, and streaming the chunks to my frontend:

        response = openai.chat.completions.create(
            model=model,
            messages=messages,
            stream=True
        )
        for chunk in response:
            ...
            yield f'event: CHUNK\ndata: {chunk.choices[0].delta.content}'
    

    Apparently, this creates hundreds of spans within Sentry transaction, leading to envelope exceeded size limits error.

    Can you suggest some solution?
    Maybe there's a way to have my chunk processing logic grouped as a single span which doesn't create any child spans?
    Something like:

    with sentry_sdk.create_span(name='chunk_processing', allow_child_spans=False):
        # chunk processing logic showed above goes here
    
  9. 4 remaining items

  10. Macok commented on Oct 24, 2025

    @Macok
    Author

    Thanks for the answer! But I'm already on 2.42.1.
    I was able to fix envelope exceeded size limits by removing the yield line from my code.

        response = openai.chat.completions.create(
            model=model,
            messages=messages,
            stream=True
        )
        for chunk in response:
            ...
            yield f'event: CHUNK\ndata: {chunk.choices[0].delta.content}'   # <-- WORKS FINE WITHOUT THIS LINE 
    

    It seems like sentry_sdk creates a span for each chunk of data streamed from my backend to the frontend.

    Is this something the Sentry team would consider fixing? Streaming chunks to the frontend is rather common for LLM-powered apps.
    For now, do you see any walkaround for me?

  11. moved this to Waiting for: Product Owner in GitHub Issues with 👀 3on Oct 24, 2025
  12. Macok commented on Oct 29, 2025

    @Macok
    Author

    @alexander-alderman-webb
    Did you manage to take a look at this?

  13. alexander-alderman-webb commented on Oct 29, 2025

    @alexander-alderman-webb
    Contributor

    Hi @Macok,

    We are aware of the problem and there is an initiative on the way that solves the size restriction in the long-term.

    In terms of a work-around, you must cut down the size of the transaction holding the spans generated by our integrations for FastAPI and OpenAI.

    What you choose to omit depends on which telemetry you care about most. Storing the LLM's response text likely accounts for most of the transaction's size, so I would write a before_send_transaction() callback that truncates the response text.

    You can experiment more with the before_send_transaction() callback. Only the return value of the function will be emitted from sentry_sdk.

    import sentry_sdk
    from sentry_sdk.consts import SPANDATA, OP
    
     def before_send_transaction(transaction, hint):
        for span in transaction["spans"]:
            if span["op"] == OP.GEN_AI_CHAT:
                span["data"][SPANDATA.GEN_AI_RESPONSE_TEXT] = "Redacted"            
    
        return transaction
    
    sentry_sdk.init(
          dsn=os.environ['SENTRY_DSN'],
          environment=os.getenv("SENTRY_ENV", "production"),
          send_default_pii=True,
          traces_sample_rate=1.0,
    +     debug=True,
    +     before_send_transaction=before_send_transaction,
      )
  14. moved this from Waiting for: Product Owner to No status in GitHub Issues with 👀 3on Oct 29, 2025
  15. Macok commented on Oct 31, 2025

    @Macok
    Author

    Ok, I'll experiment with before_send_transaction 👍
    Thanks for help!

  16. moved this to Waiting for: Product Owner in GitHub Issues with 👀 3on Oct 31, 2025
  17. sentrivana commented on Nov 11, 2025

    @sentrivana
    Contributor

    Work is underway to solve this out of the box, in the meantime cutting down on the event size in before_send_transaction is the way to go -- I'll close.

  18. zycon commented on Mar 18, 2026

    @zycon

    @alexander-alderman-webb Thank you, do we have a possible fix on sentry.io anytime soon?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions