Skip to main content
Streaming is supported by both Chat Completions (completion() / acompletion()) and Responses (responses() / aresponses()).
A ready-to-run example is available here!
Enable real-time display of LLM responses as they’re generated, token by token. This guide demonstrates how to use streaming callbacks to process and display tokens as they arrive from the language model.

How It Works

Streaming allows you to display LLM responses progressively as the model generates them, rather than waiting for the complete response. This creates a more responsive user experience, especially for long-form content generation.
1

Enable Streaming on LLM

Configure the LLM with streaming enabled:
2

Define Token Callback

Create a callback function that processes streaming chunks as they arrive:
The callback receives a ModelResponseStream object containing:
  • choices: List of response choices from the model
  • delta: Incremental content changes for each choice
  • content: The actual text tokens being streamed
3

Register Callback with Conversation

Pass your token callback to the conversation:
The token_callbacks parameter accepts a list of callbacks, allowing you to register multiple handlers if needed (e.g., one for display, another for logging).

Responses Streaming

For direct LLM calls, pass stream=True and an on_token callback. The call returns a complete LLMResponse after the stream finishes; callbacks receive ModelResponseStream chunks while it is being read.
Use await llm.aresponses(...) for the asynchronous equivalent. It accepts synchronous or asynchronous callbacks. Without a callback, the SDK normally falls back to a non-streaming request. Endpoints that require streaming, such as ChatGPT subscription endpoints, still drain the stream and return the complete response without a callback. Instrumentation wrappers do not need to inherit from LiteLLM’s stream classes. The SDK accepts synchronous iterables and asynchronous iterables on the async path, including synchronous wrappers returned to an async caller. A completed Responses event yielded by the stream remains valid if the wrapper’s completed_response attribute is absent or None. A non-null wrapper completion takes precedence after the stream has been read. A stream with no completion raises LLMNoResponseError and follows the configured retry policy.

Ready-to-run Example

This example is available on GitHub: examples/01_standalone_sdk/29_llm_streaming.py
examples/01_standalone_sdk/29_llm_streaming.py
You can run the example code as-is.
The model name should follow the LiteLLM convention: provider/model_name (e.g., anthropic/claude-sonnet-4-5-20250929, openai/gpt-4o). The LLM_API_KEY should be the API key for your chosen provider.
ChatGPT Plus/Pro subscribers: You can use LLM.subscription_login() to authenticate with your ChatGPT account and access Codex models without consuming API credits. See the LLM Subscriptions guide for details.

Next Steps