Estimated monthly savings
$5,765
Approved or applied
$0
Findings
6
Clear AI Spend does not change your production configuration. Approve and Mark applied only record your decision here; apply the change in your own systems.
Switch SupportClassifier to a lower-cost model
Expensive model for simple task · SupportClassifier
Ticket routing is a short classification task running on a premium reasoning model.
Current configuration
claude-opus-class · temperature 0.2 · full ticket body
$0.040 / task · current success rate 98.7%
Recommended configuration
claude-haiku-class · temperature 0 · trimmed ticket body + label schema
$0.0060 / task · estimated success rate 98.4%
Estimated savings: $1,480/month
Enable prompt caching on ReplyDrafter preamble
Missing caching · ReplyDrafter
A 6.1k-token static preamble is re-billed on every one of ~41k monthly calls.
Current configuration
No prompt caching · full preamble per call
$0.052 / task · current success rate 97.2%
Recommended configuration
Provider prompt caching enabled · 5 min TTL
$0.031 / task · estimated success rate 97.2%
Estimated savings: $1,240/month
Add structured output + retry ceiling on LeadEnricher
Excessive retries · LeadEnricher
Schema validation failures trigger up to 4 retries on a premium model.
Current configuration
Free-form JSON parsing · max 4 retries · same model on retry
$0.118 / task · current success rate 91.3%
Recommended configuration
Native structured outputs · max 1 retry · efficient model on retry
$0.061 / task · estimated success rate 94.4%
Estimated savings: $1,105/month
Replace full-log context with top-k retrieval
Oversized context · IncidentSummarizer
Average 11.2k input tokens per summary where 1.6k retrieved chunks match quality.
Current configuration
Entire incident log inlined
$0.074 / task · current success rate 95.8%
Recommended configuration
Vector retrieval, top 8 chunks, 1.6k token budget
$0.027 / task · estimated success rate 95.1%
Estimated savings: $840/month
Add a loop breaker to PipelineTriager
Tool-call loops · PipelineTriager
12% of runs exceed 8 tool calls without converging, then fail.
Current configuration
Unbounded tool loop · premium model each hop
$0.163 / task · current success rate 88.1%
Recommended configuration
Max 4 tool hops · escalate to human after budget breach
$0.098 / task · estimated success rate 89.7%
Estimated savings: $690/month
Fingerprint-dedupe SQLCopilot requests
Duplicate requests · SQLCopilot
Identical query requests within 60s are billed twice.
Current configuration
No deduplication layer
$0.021 / task · current success rate 96.5%
Recommended configuration
Request fingerprint cache · 120s window
$0.017 / task · estimated success rate 96.5%
Estimated savings: $410/month