Gisting: Compressing LLM Agent context to ↑ throughput and ↓ cost
Gisting compresses context into a set of learned tokens, preserving its quality while making the model faster and cheaper.
Sources
- T1Gisting: Compressing LLM Agent context to ↑ throughput and ↓ costStripe / Shopify Eng / Netflix TechBlog / Dropbox Tech / Slack Eng / Spotify Eng