Skip to main content

Azure foundry models comms are broken

I have to disable streaming responses when using Azure foundry and mistral medium 3.5. Ideally it shouldn’t be required. The sequence that is occurring is that the stream delivers the reply normally across 7 chunks

  1. The endpoint sends a final chunk: {"choices":[], "usage":{...}}

  2. BoltAI reads choices[0].delta without checking the array is non-empty → undefined is not an object (evaluating 'g.delta')

That trailing usage-only chunk is standard OpenAI streaming behaviour, and I confirmed the Foundry Mistral deployment emits it unconditionally on all three endpoint shapes — /openai/v1/, and the deployment path on both 2024-10-21 and 2025-01-01-preview. So there's no URL or api-version that avoids it; disabling streaming is the only client-side workaround.

The means replies appear all at once rather than typing out. Everything else is unaffected.

Status: Completed2 comments

Log in to comment and vote

Comments2

  • Daniel Nguyen changed status to Completed
    Team•

    Sep 15

    Pinned

    I fixed this in v2.16. Can you confirm? Thanks

  • Daniel Nguyen

    Team•

    Sep 11

    Thanks for the detailed report. I tested the trailing usage-only chunk through BoltAI’s Azure provider, but couldn’t reproduce the error.

    Are you using the built-in Azure OpenAI provider or a Custom OpenAI-compatible provider? Could you share your BoltAI version/build and a screenshot of the provider settings, with the API key and any other credentials hidden? I’d like to test the same configuration.