Desperdicio mensual estimado
13.418 US$
Proporción del gasto rastreado
31,5 %
Hallazgos
7
Premium model used for binary classification
Modelo caro para una tarea simple
68% of SupportClassifier tasks are single-label routing decisions under 400 tokens, but run on a premium reasoning model.
Retry storms on malformed JSON output
Reintentos excesivos
LeadEnricher retries up to 4x when schema validation fails. Retries account for 31% of its spend.
Identical system prompts re-sent without prompt caching
Falta de caché
A 6.1k-token system preamble is re-billed on every call across three agents. Prompt caching is not enabled.
Full document dumps instead of retrieval
Contexto sobredimensionado
IncidentSummarizer sends whole log files (avg 11.2k tokens) where top-k retrieval of 1.5k tokens matches quality.
Spend on tasks that never produced an outcome
Tareas fallidas
Failed tasks still consume premium model calls before aborting. No early-exit guard is configured.
Same request fingerprint within 60 seconds
Solicitudes duplicadas
Client-side double submits produce identical completions billed twice.
Agent loops between search and reasoning steps
Bucles de llamadas a herramientas
PipelineTriager exceeds 8 tool calls in 12% of runs without converging; no loop breaker configured.