ice_age
pomagrenate/ice_age
Overview
Universal IDE plugin & proxy that cuts LLM token consumption by up to 70% using deterministic AST pruning & context compression. Written in Go.
Technologies & Concepts
Project Analysis
🔴 Problem
AI agent token inefficiency: Modern AI coding agents spend excessive tokens on pleasantries, repeated explanations, filler words, hedging, and redundant context rather than delivering direct technical answers.
Cost impact: For developers using AI agents throughout the day, verbose responses lead to significant API token costs, especially for teams spending thousands on LLM usage.
Communication overhead: AI responses optimized for general conversation rather than high-frequency software engineering create unnecessary cognitive load and slower scanning for developers.
No optimization layer: Existing AI coding tools lack a compression layer to reduce token consumption while maintaining technical accuracy and information density.
⚪ Baseline
Starting point: Normal AI agent responses with full conversational style
Baseline characteristics:
- Pleasantries and greetings ("Sure! I'd be happy to help...")
- Long-form explanations for simple technical answers
- Filler words and hedging ("just", "really", "basically")
- Redundant context and repeated information
- Articles and unnecessary words
Example baseline response: "Sure! I'd be happy to help you understand this issue. The reason this happens is that you're creating a new object reference on every render. React sees that the reference has changed, even if the actual values inside the object remain the same. To fix this issue, you can use the useMemo hook to memoize the object and prevent a new reference from being created on every render."
🔵 Change
Built ice_age: Universal IDE plugin & proxy that makes AI coding agents communicate in compressed, technical, high-signal prose
Key architectural changes:
- Prompt optimization system that removes filler, pleasantries, and unnecessary articles
- Deterministic AST pruning to remove irrelevant code context
- Multiple intensity levels (lite, full, ultra, wenyan)
- Auto-clarity fallback for security warnings and ambiguous situations
- Multi-agent support (Claude Code, Cursor, Windsurf, Cline, Copilot, Gemini CLI)
- Session persistence and mode tracking
Implementation approach: Go-based plugin architecture with hooks, skills, and rule files for different AI coding agents
🟣 Measurement
Test environment: gemini-3.7-flash model with 2 trials per prompt
Benchmark methodology:
- 10 diverse prompts across categories (debugging, bugfix, setup, explanation, refactor, architecture, code-review, devops)
- Comparison: ice_age vs normal (uncompressed) context
- Token count measurement for both input and output
- Category-wise performance analysis
Metrics collected: Average token savings, category-specific performance, perfect compression cases, failure modes
🟢 Result
Overall token savings: 34% average reduction in output tokens (167 tokens → 78 tokens average)
Best performing categories: Setup questions (74% savings), explanation questions (72% savings), bugfix questions (63% savings)
Perfect compression cases: 100% savings for refactor and implementation prompts where ice_age provided answers while normal mode failed
Performance range: 0% to 100% savings depending on prompt category and complexity
Real-world impact: For teams spending thousands on API tokens, 34% average savings represents substantial cost reduction while maintaining code quality and developer productivity
Performance Visualizations
Savings Percentage by Category

Summary Statistics Comparison

Token Comparison

🟡 Lesson
Prompt optimization effectiveness: The 34% average savings demonstrates that removing conversational overhead from AI responses can significantly reduce token consumption without sacrificing technical accuracy or information density.
Category performance variance: I learned that effectiveness varies significantly by category - setup and explanation questions showed the highest savings (60-74%), while some debugging questions showed 0% or negative savings. This suggests the approach works best for well-structured technical content.
Deterministic vs ML-based approaches: The deterministic rule-based approach provides consistent, predictable results unlike ML-based compression. This reliability is crucial for development workflows where consistency matters.
Trade-offs and limitations: Some prompts showed 0% or negative savings, indicating that compression isn't universally beneficial. For debugging questions requiring full context, aggressive pruning might remove relevant information, requiring auto-clarity fallbacks.
Perfect compression insights: The 100% savings cases (async-refactor, error-boundary) were particularly interesting - ice_age provided complete answers while normal mode failed, suggesting that compression can sometimes improve context handling by removing noise.
Communication design matters: The project taught me that AI responses should be optimized for engineering workflows (precise → technical → scannable → minimal) rather than general conversation (greeting → disclaimer → repetition → explanation → conclusion → actual answer).