Recommendations

Configuration changes with projected savings, quality impact and risk

Demo data. No AI provider or gateway is connected yet — every figure on this screen is seeded sample data, not production telemetry.

Estimated monthly savings

$5,765

Estimate, not measured savings

Approved or applied

$0

0 / 6

Findings

6

Clear AI Spend does not change your production configuration. Approve and Mark applied only record your decision here; apply the change in your own systems.

Switch SupportClassifier to a lower-cost model

Expensive model for simple task · SupportClassifier

Confidence 94%
Risk: LowProposed

Ticket routing is a short classification task running on a premium reasoning model.

Current configuration

claude-opus-class · temperature 0.2 · full ticket body

$0.040 / task · current success rate 98.7%

Recommended configuration

claude-haiku-class · temperature 0 · trimmed ticket body + label schema

$0.0060 / task · estimated success rate 98.4%

Estimated savings: $1,480/month

Enable prompt caching on ReplyDrafter preamble

Missing caching · ReplyDrafter

Confidence 97%
Risk: LowProposed

A 6.1k-token static preamble is re-billed on every one of ~41k monthly calls.

Current configuration

No prompt caching · full preamble per call

$0.052 / task · current success rate 97.2%

Recommended configuration

Provider prompt caching enabled · 5 min TTL

$0.031 / task · estimated success rate 97.2%

Estimated savings: $1,240/month

Add structured output + retry ceiling on LeadEnricher

Excessive retries · LeadEnricher

Confidence 89%
Risk: LowProposed

Schema validation failures trigger up to 4 retries on a premium model.

Current configuration

Free-form JSON parsing · max 4 retries · same model on retry

$0.118 / task · current success rate 91.3%

Recommended configuration

Native structured outputs · max 1 retry · efficient model on retry

$0.061 / task · estimated success rate 94.4%

Estimated savings: $1,105/month

Replace full-log context with top-k retrieval

Oversized context · IncidentSummarizer

Confidence 82%
Risk: MediumProposed

Average 11.2k input tokens per summary where 1.6k retrieved chunks match quality.

Current configuration

Entire incident log inlined

$0.074 / task · current success rate 95.8%

Recommended configuration

Vector retrieval, top 8 chunks, 1.6k token budget

$0.027 / task · estimated success rate 95.1%

Estimated savings: $840/month

Add a loop breaker to PipelineTriager

Tool-call loops · PipelineTriager

Confidence 78%
Risk: MediumProposed

12% of runs exceed 8 tool calls without converging, then fail.

Current configuration

Unbounded tool loop · premium model each hop

$0.163 / task · current success rate 88.1%

Recommended configuration

Max 4 tool hops · escalate to human after budget breach

$0.098 / task · estimated success rate 89.7%

Estimated savings: $690/month

Fingerprint-dedupe SQLCopilot requests

Duplicate requests · SQLCopilot

Confidence 93%
Risk: LowProposed

Identical query requests within 60s are billed twice.

Current configuration

No deduplication layer

$0.021 / task · current success rate 96.5%

Recommended configuration

Request fingerprint cache · 120s window

$0.017 / task · estimated success rate 96.5%

Estimated savings: $410/month