TIL / Token-aware batching beats fixed-size batching for embedding calls

Token-aware batching beats fixed-size batching for embedding calls

Azure OpenAIEmbeddingsReliabilityRAG

The problem

An embedding endpoint enforces a tokens-per-minute limit, not just a requests-per-minute one. Batching by a fixed item count (say, 100 texts per call) works fine until a run of unusually long chunks pushes one batch over the token limit and the call comes back 429. Under peak indexing load, that turned into a steady trickle of failed batches.

The fix

Track a running token estimate per batch and flush it when either a text-count cap or a token cap is hit, whichever comes first. On a 429, honor the Retry-After header instead of a fixed backoff.

def batch_texts(texts, max_items=100, max_tokens=15_000, estimate_tokens=len_tokens):
    batch, batch_tokens = [], 0
    for text in texts:
        t = estimate_tokens(text)
        if batch and (len(batch) >= max_items or batch_tokens + t > max_tokens):
            yield batch
            batch, batch_tokens = [], 0
        batch.append(text)
        batch_tokens += t
    if batch:
        yield batch

Gotcha

A token estimate (even a rough one, like len(text) // 4) is enough to keep batches under budget in practice - don’t reach for the real tokenizer on the hot path just to save a rare retry. The bigger win is honoring Retry-After exactly rather than guessing a sleep duration; a fixed backoff that’s shorter than the server’s own window just produces the same 429 again.