Gisting: Compressing LLM Agent context to ↑ throughput and ↓ cost

Gisting compresses context into a set of learned tokens, preserving its quality while making the model faster and cheaper.

ai

Sources