DistilCodeAll techniques

Technique 02 / 07

Prompt and context caching

Mark the stable parts of a request so a supporting model provider can reuse work across turns.

Interactive model

See the context change.

Toggle the same follow-up request between an uncached and cache-aware path.

Cache-aware request1 new block

Reuse the stable prefix

  1. 01

    Cached system instructions

  2. 02

    Cached repository conventions

  3. 03

    Cached prior conversation

  4. 04

    Process new follow-up

Conceptual demonstration. Counts illustrate the flow, not measured production savings.

How it works

Less context. Same thread of work.

Coding sessions repeatedly send the same system instructions and much of the same conversation prefix. When the provider supports prompt caching, keeping those blocks stable lets it reuse a previously processed prefix rather than charge and compute for it as entirely new input.

DistilCode adds provider-specific cache markers to selected system and recent messages. Its gateway integration can also forward cache keys, TTLs, and skip-cache options. The provider ultimately decides whether a cache entry is created or hit, so caching is an optimisation opportunity rather than a guarantee.

In practice

The operating loop

  1. 01

    Find stable boundaries

    System instructions and selected message boundaries are marked without rewriting their content.

  2. 02

    Speak the provider dialect

    Anthropic, Bedrock, OpenAI-compatible, OpenRouter, Copilot, and Alibaba options are represented in the format each integration expects.

  3. 03

    Measure real hits

    Usage records distinguish ordinary input from cache-read and cache-write tokens.