how-to-stream-llm-responses.md — opusjake_os ARTICLE
// OPUSJAKE BLOG · LLM STREAMING

How to Stream LLM Responses: SSE, Proxies, Cancels, and the Render Loop

2026-10-098 MIN READBY · OPUSJAKE
FIRST TOKEN
llm streamingserver-sent eventsclaude apiopenai apiai ux

To stream LLM responses, set stream: true on the API call, relay the provider's server-sent events through your backend without buffering, and read them in the client with a stream reader that appends each text delta as it lands. The answer takes the same total time, but the first words show up in under a second instead of after the full generation finishes.

That is the whole idea. The work is in the plumbing: proxies that buffer, events that split across network chunks, tool calls that arrive as broken JSON, and users who hit stop while you keep paying for tokens. This guide walks the full path from model to screen.

TL;DR

  • Streaming changes perceived speed, not real speed. A 600-token answer at 60 tokens per second takes 10 seconds either way. Streaming shows the first word at roughly 0.5 seconds instead of 10.
  • It is server-sent events over one open HTTP response. text/event-stream, event: and data: lines, a blank line between events.
  • Your backend must relay, not collect. Forward each delta as it arrives and disable buffering at every hop: framework, proxy, compression, client.
  • Wire cancellation end to end. A stop button that only closes the browser tab keeps the model generating, and you keep paying.
  • Never act on partial tool calls. Accumulate argument fragments and parse when the block closes.

What streaming actually changes

Two numbers describe how fast an LLM call feels. Time to first token (TTFT) is how long until the model emits anything. Throughput is how many tokens per second it produces after that. Without streaming, the user waits TTFT plus the entire generation before seeing a single character. With streaming, they wait TTFT, then read along as the rest arrives.

People read at roughly 4 to 5 words per second. Most models generate faster than that, so a streamed answer stays ahead of the reader the whole time. The wait effectively disappears, which is why every serious chat product streams.

Streaming also matters for long outputs in a way that has nothing to do with UX. A non-streamed request that generates for several minutes risks idle HTTP timeouts at load balancers and SDKs, and Anthropic's SDK will steer you to streaming for very large max_tokens values. Even for a backend job with no user watching, a long generation is safer as a stream you collect.

Buffered versus streamed response timingFor the same 600-token answer generated in about 10 seconds, a buffered response shows nothing until second 10, while a streamed response shows the first words at about half a second and fills in while the user reads, so total time is equal but perceived wait drops by about 95 percent.// FIG · SAME ANSWER, TWO WAITSStreaming moves the first word, not the lastBUFFEREDspinner · nothing on screenALLSTREAMEDtext appears as it is generated0s0.5s10sstreamed: first word at 0.5sbuffered: first word at 10sSame tokens, same bill. The user only feels the first gap.

The anatomy of a stream

Every major provider streams over server-sent events (SSE). The response has Content-Type: text/event-stream, the connection stays open, and the body is a sequence of events separated by a blank line. Each event has an optional event: name and a data: line carrying JSON.

On the Anthropic Messages API, one streamed reply looks like this, in order:

  1. message_start carries the message ID, model, and initial usage.
  2. content_block_start opens a block: text, a tool call, or thinking.
  3. content_block_delta arrives many times. For text the delta is text_delta with a few characters. For a tool call it is input_json_delta with a fragment of the arguments.
  4. content_block_stop closes the block. Only now is a tool call's JSON complete.
  5. message_delta carries the stop_reason and final output token count.
  6. message_stop ends the stream.

You will also see ping events (ignore them) and possibly an error event mid-stream, such as an overloaded error. That last one matters: the HTTP status was already 200 when the stream began, so errors after that point only show up as events. Your client has to handle them.

OpenAI follows the same shape with different names. The Responses API emits typed events like response.output_text.delta and response.completed. Chat Completions sends choices[0].delta.content chunks and ends with data: [DONE], and you need stream_options: {"include_usage": true} to get token counts at the end.

Relay the stream from your server

Never call a model API directly from the browser, because that exposes your key. Your backend makes the call and relays events to the client. The whole trick is forwarding each piece the moment it arrives instead of collecting the full answer first.

Here is a minimal relay in a Node route handler using the Anthropic SDK:

import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();

export async function POST(req: Request) {
  const { messages } = await req.json();
  const stream = client.messages.stream(
    { model: "claude-sonnet-5-5", max_tokens: 2048, messages },
    { signal: req.signal }, // client disconnect aborts the upstream call
  );

  const body = new ReadableStream({
    async start(controller) {
      const enc = new TextEncoder();
      const send = (event: string, data: unknown) =>
        controller.enqueue(enc.encode(`event: ${event}\ndata: ${JSON.stringify(data)}\n\n`));
      try {
        for await (const ev of stream) {
          if (ev.type === "content_block_delta" && ev.delta.type === "text_delta") {
            send("text", { t: ev.delta.text });
          }
        }
        const final = await stream.finalMessage();
        send("done", { stop: final.stop_reason, usage: final.usage });
      } catch (err) {
        send("error", { message: "generation failed" });
      } finally {
        controller.close();
      }
    },
  });

  return new Response(body, {
    headers: {
      "Content-Type": "text/event-stream",
      "Cache-Control": "no-cache, no-transform",
      "X-Accel-Buffering": "no",
    },
  });
}

Three details carry most of the weight. You re-emit your own small event schema (text, done, error) instead of piping raw provider events, which keeps the client simple and lets you swap providers later. The final done event carries usage so you can log cost per request. And the request's abort signal is passed upstream, which is what makes cancellation real.

Kill the buffering at every hop

The most common streaming bug: everything works locally, then in production the answer appears all at once after ten seconds. Something in the path is buffering. Check each hop in order.

  • Framework. Return a streamed Response or use the framework's streaming helper. Some route types and older serverless runtimes buffer the whole body by design. Vercel functions and edge routes support streaming; confirm your route type does.
  • Reverse proxy. nginx buffers upstream responses by default. Send X-Accel-Buffering: no from the app or set proxy_buffering off for that location.
  • Compression. Gzip middleware waits to fill a buffer before flushing. Exclude text/event-stream from compression, which no-transform also signals to well-behaved intermediaries.
  • CDN. Make sure the route is not cached and the CDN passes streamed bodies through.
  • Client. await res.json() or await res.text() waits for the end. You have to read res.body as a stream.

A quick test from your terminal: curl -N against the endpoint. The -N flag disables curl's own buffering, so if events print one by one there, the server side is fine and the problem is the client.

Read it in the client and render without jank

The browser's EventSource only does GET requests with no body, which does not fit an LLM call. Use fetch and read the body stream yourself:

const ctrl = new AbortController();
const res = await fetch("/api/chat", {
  method: "POST",
  body: JSON.stringify({ messages }),
  signal: ctrl.signal,
});
const reader = res.body!.pipeThrough(new TextDecoderStream()).getReader();
let buf = "";
while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  buf += value;
  const events = buf.split("\n\n");
  buf = events.pop()!; // keep the incomplete tail for the next chunk
  for (const raw of events) handleEvent(raw);
}

The buf = events.pop() line is the one people skip. A network chunk can end halfway through an event, or even halfway through a multi-byte character, so you keep the incomplete tail and prepend it to the next chunk. TextDecoderStream handles the character split for you.

For rendering, do not re-render the whole message on every delta. Deltas can arrive 50 or more times per second, and re-parsing markdown on each one makes long answers stutter. Append deltas to a string, then render on requestAnimationFrame at most once per frame. Expect incomplete markdown mid-stream (an unclosed code fence, half a table) and either tolerate it or close open fences before rendering.

Agents that run multi-step loops benefit from the same pattern: stream the model's text for each step and emit a separate event when a tool runs, so the user sees progress instead of silence. The Write Loops, Not Prompts guide covers the loop structure that sits behind that UI.

Stream tool calls without acting on half a call

When the model calls a tool, its arguments arrive as JSON string fragments: {"ci, then ty": "Tor, then onto"}. None of those is valid on its own.

Accumulating streamed tool-call argumentsA streamed tool call arrives as three partial JSON fragments that are each invalid on their own. The client appends them into a buffer per content block and only parses, validates, and executes the tool after the block stop event confirms the arguments are complete.// FIG · PARTIAL JSONBuffer the fragments, run on block stopdelta 1 {"cidelta 2 ty": "Tordelta 3 onto"}each one: invalid JSONBUFFER[block 1]append every input_json_deltaCONTENT_BLOCK_STOPPARSE + VALIDATE + RUN{"city": "Toronto"}safe to executeShow progress from partials if you like. Execute only on stop.

Keep one buffer per content block index, append each argument delta, and parse on the block's stop event. The SDK helpers (stream.finalMessage() on Anthropic, the accumulated response on OpenAI) do this for you if you do not need the raw events. Streaming structured output works the same way: you can show fields filling in with a partial JSON parser, but you validate against your schema and write to anything real only after the stream completes.

Cancellation, errors, and when not to stream

Cancellation. Put a stop button on every streamed answer, wire it to AbortController.abort() on the client, and pass the server request's abort signal to the provider call, as in the relay above. When the upstream request aborts, the model stops generating and output billing stops with it. If you skip the server half, closing the tab leaves your backend happily paying for a 4,000-token answer nobody will read.

Errors. Because the 200 status is already sent, a failure halfway through arrives as an error event or a dropped connection. Show the partial text with a clear "response interrupted" marker and a retry button. Do not silently retry and stream a second, different answer on top of the first.

Metrics. Log TTFT and output tokens per second per request, along with the usage from the final event. TTFT regressions are the first thing users notice, usually from a bloated prompt or a slow upstream tool call before the model even starts.

When not to stream. Skip it for batch jobs and background agents nobody is watching (unless the output is long enough to risk timeouts), for short classification or extraction calls where the whole answer is ten tokens, and for anything that must pass validation or moderation before a human sees it. Streaming unvetted text to a screen is the same as showing it.

The bottom line

To stream LLM responses well, turn on stream, relay small events from your server the moment they arrive, strip buffering from every hop with curl -N as your test, parse SSE with a carry-over buffer in the client, act on tool calls only at block stop, and wire the stop button all the way to the provider. Total generation time does not change. The first word moving from second ten to second one is the whole win.

I write up the build patterns I actually ship, with the code, in the OpusJake newsletter.

// FREQUENTLY ASKED
What does it mean to stream an LLM response?

Streaming means the API sends the model's output in small pieces as it is generated instead of waiting for the full answer. You set a stream flag on the request, and the provider holds the HTTP connection open and pushes server-sent events, each carrying a few tokens of text. Your app appends those pieces to the screen as they arrive. Total generation time is the same, but the user sees the first words in a fraction of a second instead of staring at a spinner for ten or twenty seconds. That first visible token is what makes a chat interface feel fast.

Does streaming make an LLM faster or cheaper?

Neither, in raw terms. The model generates tokens at the same speed and you pay the same per token whether you stream or not. What streaming changes is time to first visible output, which drops from the full generation time to roughly the time to first token, often under a second. It can save money indirectly: if a user cancels a bad answer early and your server aborts the upstream request, generation stops and you stop paying for output tokens past that point. Without streaming and proper cancellation, you pay for the whole answer even if nobody reads it.

Why is my LLM stream arriving all at once?

Something between the model and the browser is buffering. The usual suspects are a reverse proxy like nginx with proxy buffering on, a compression middleware that waits for enough bytes to gzip, a serverless platform or framework route that does not support streamed responses, or client code that calls response.json() or await response.text() instead of reading the body stream. Fix it by sending Content-Type text/event-stream with Cache-Control no-cache, adding X-Accel-Buffering no for nginx, excluding the route from compression, and reading the body with a stream reader on the client.

Can I stream tool calls and JSON output?

Yes, but you cannot act on them mid-stream. Providers stream tool-call arguments and structured output as partial JSON string fragments, so the text you have halfway through is usually not valid JSON. Accumulate the fragments per content block and parse only when the block-stop or done event arrives. If you want to show progress, use a partial JSON parser to render fields as they fill in, but never execute a tool or write to a database until the complete arguments have parsed and validated.

Should I use EventSource or fetch to consume an LLM stream?

Use fetch with a body stream reader. The browser's EventSource API only supports GET requests and cannot send a JSON body or custom auth headers, while an LLM call needs a POST with the conversation and usually a token. With fetch you get response.body as a ReadableStream, decode it with TextDecoder, split events on blank lines, and handle each data line. You also get an AbortController for clean cancellation, which EventSource does not give you in a form that maps to a POST request.

// BUILD WITH OPUSJAKE

OpusJake is Jake Schincariol's operating system for building with AI: agents, workflows, prompts, and the free resources behind them. Get the next move every week.

STATUS · ONLINE · OPUSJAKE © OPUSJAKE // CRT V1