Estimated monthly waste
$13,418
Share of tracked spend
31.5%
Findings
7
Premium model used for binary classification
Expensive model for simple task
68% of SupportClassifier tasks are single-label routing decisions under 400 tokens, but run on a premium reasoning model.
Retry storms on malformed JSON output
Excessive retries
LeadEnricher retries up to 4x when schema validation fails. Retries account for 31% of its spend.
Identical system prompts re-sent without prompt caching
Missing caching
A 6.1k-token system preamble is re-billed on every call across three agents. Prompt caching is not enabled.
Full document dumps instead of retrieval
Oversized context
IncidentSummarizer sends whole log files (avg 11.2k tokens) where top-k retrieval of 1.5k tokens matches quality.
Spend on tasks that never produced an outcome
Failed tasks
Failed tasks still consume premium model calls before aborting. No early-exit guard is configured.
Same request fingerprint within 60 seconds
Duplicate requests
Client-side double submits produce identical completions billed twice.
Agent loops between search and reasoning steps
Tool-call loops
PipelineTriager exceeds 8 tool calls in 12% of runs without converging; no loop breaker configured.