Skip to main content

budget-aware-cost-management

Implements session-level budget quotas, cost monitoring with threshold alerts, and ROI tracking to enforce AI agent spending limits and optimize return on investment.

Zur Installation springen

Quellinformationen

Repository
paulpas/agent-skill-router
Letzte Quellaktivität
9. Juni 2026 um 01:54
Erkannte Sprache von SKILL.md
Englisch
Sterne
6
Forks
0

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
budget-aware-cost-management
description
Implements session-level budget quotas, cost monitoring with threshold alerts, and ROI tracking to enforce AI agent spending limits and optimize return on investment.
license
MIT
compatibility
opencode
metadata
{"version":"1.0.0","domain":"agent","role":"implementation","scope":"infrastructure","output-format":"analysis","triggers":"budget quota, cost monitoring, token budget, spending limits, ROI tracking, AI cost optimization, how do i control agent spending","archetypes":["tactical","orchestration"],"anti_triggers":["model selection","query complexity classification","performance benchmarking"],"response_profile":{"verbosity":"medium","directive_strength":"high","abstraction_level":"operational"},"related-skills":"resource-optimization, goal-setting-monitoring, evaluation-monitoring"}
# Budget-Aware Cost Management for AI Agents Implements budget quota enforcement, real-time cost monitoring, and ROI tracking to ensure AI agent operations stay within financial constraints while maximizing output value. ## TL;DR Checklist - [ ] Define per-session budget with hard/soft limits - [ ] Implement token-level cost tracking per operation - [ ] Set threshold alerts at 50%, 75%, 90% budget utilization - [ ] Configure adaptive retry strategies based on remaining budget - [ ] Track ROI (output quality / cost spent) per agent session - [ ] Enforce priority-based budget allocation in multi-agent deployments - [ ] Generate cost-utilization reports with optimization recommendations --- ## When to Use Use this skill when: - Deploying multi-agent systems where total token spend needs hard caps - Building production AI services with per-user or per-session budget limits - Running experiments where ROI tracking determines continuation vs shutdown - Managing agent teams that share a pooled compute budget - Implementing cost-aware scheduling in resource-constrained environments (edge devices, mobile) ## When NOT to Use Avoid this skill for: - Simple single-agent scripts with negligible cost (overhead outweighs benefit) - Research/exploratory work without budget constraints (use unconstrained reasoning) - Offline/local model deployments with zero marginal cost - Situations where quality must be guaranteed regardless of cost (use `resource-optimization` instead) --- ## Core Workflow 1. **Define Budget Architecture** — Establish per-session budgets with hard limits (never exceed) and soft limits (trigger optimization). Classify all planned operations by cost tier. **Checkpoint:** Verify budget cap is enforced at the infrastructure level, not just tracked. 2. **Implement Cost Tracking Layer** — Deploy token-level monitoring that counts input/output tokens per API call, aggregates session totals, and maintains a real-time utilization counter. **Checkpoint:** Confirm tracking granularity matches billing granularity (per-token, per-model). 3. **Configure Threshold Alerts** — Set up warnings at 50% (informational), 75% (prepare optimization), 90% (activate conservative mode), and 100% (graceful shutdown trigger). **Checkpoint:** All alert handlers must be non-blocking; alerts should not consume budget tokens. 4. **Deploy Adaptive Retry Logic** — Implement budget-aware retry strategies: aggressive retries when >75% remaining, single retry at 50-75%, no retries below 25%. Classify errors as retryable vs terminal before consuming additional budget. **Checkpoint:** Total retries × estimated cost per retry ≤ remaining budget margin. 5. **Track ROI Metrics** — Measure output quality (success rate, user satisfaction, accuracy) against cost spent per operation. Calculate ROI = quality_score / (total_tokens × cost_per_token). Log metrics for trend analysis. **Checkpoint:** ROI calculation must use normalized quality scores, not raw outputs. 6. **Generate Cost Reports** — Produce utilization summaries showing budget consumed vs allocated, top-cost operations, optimization opportunities, and ROI trends. Provide actionable recommendations for quota adjustment. **Checkpoint:** Reports must include both absolute costs and per-operation averages for comparison. --- ## Implementation Patterns ### Pattern 1: Session Budget Quota Manager Enforces hard and soft budget limits per session with graceful degradation. ```python from dataclasses import dataclass, field from enum import Enum from typing import Optional class LimitType(Enum): HARD = "hard" # Never exceed — triggers shutdown SOFT = "soft" # Warn + optimize — does not block @dataclass class BudgetLimit: token_budget: int limit_type: LimitType cost_per_million_tokens: float @dataclass class SessionBudget: """Per-session budget quota with hard/soft limits and real-time tracking.""" budget: BudgetLimit tokens_used: int = 0 cost_incurred: float = 0.0 alerts_triggered: list[str] = field(default_factory=list) is_active: bool = True @property def utilization_pct(self) -> float: return (self.tokens_used / self.budget.token_budget) * 100 if self.budget.token_budget > 0 else 0.0 @property def tokens_remaining(self) -> int: return max(0, self.budget.token_budget - self.tokens_used) def estimate_operation_cost(self, estimated_tokens: int) -> float: """Estimate cost for a planned operation.""" return (estimated_tokens / 1_000_000) * self.budget.cost_per_million_tokens def can_afford(self, estimated_tokens: int) -> bool: """Check if session can afford estimated operation without hitting hard limit.""" return (self.tokens_used + estimated_tokens) <= self.budget.token_budget def record_consumption(self, tokens: int) -> bool: """Record token consumption. Returns False if hard limit exceeded.""" new_total = self.tokens_used + tokens if self.budget.limit_type == LimitType.HARD and new_total > self.budget.token_budget: self.is_active = False self.alerts_triggered.append(f"HARD LIMIT EXCEEDED: {new_total} > {self.budget.token_budget}") return False self.tokens_used = new_total self.cost_incurred += self.estimate_operation_cost(tokens) self._check_threshold_alerts() return True def _check_threshold_alerts(self): """Generate threshold alerts without consuming budget tokens.""" utilization = self.utilization_pct if 75 <= utilization < 90 and "SOFT_LIMIT_75" not in self.alerts_triggered: self.alerts_triggered.append("SOFT_LIMIT_75") elif 90 <= utilization < 100 and "CONSERVATIVE_MODE_90" not in self.alerts_triggered: self.alerts_triggered.append("CONSERVATIVE_MODE_90") elif utilization >= 100 and "HARD_LIMIT_TRIGGERED" not in self.alerts_triggered: self.is_active = False self.alerts_triggered.append("HARD_LIMIT_TRIGGERED") # --- Usage Example --- def run_agent_session(session_id: str) -> SessionBudget: """Create a session with $5 budget at $0.03/M tokens.""" budget = BudgetLimit( token_budget=166_666_667, # ~$5 at $0.03/M tokens limit_type=LimitType.HARD, cost_per_million_tokens=0.03, ) session = SessionBudget(budget=budget) # Before each operation: check affordability estimated_tokens = 5_000 if not session.can_afford(estimated_tokens): raise BudgetExhaustedError(f"Session {session_id} cannot afford estimated operation") # ... execute operation ... actual_tokens = perform_llm_call(session_id, estimated_tokens) # Record consumption — may deactivate session at hard limit session.record_consumption(actual_tokens) return session class BudgetExhaustedError(Exception): """Raised when a session's budget is exhausted and no operations can proceed.""" pass ``` ### Pattern 2: Real-Time Cost Monitor with Threshold Alerts Monitors token consumption in real-time with non-blocking alerts. ```python import threading from dataclasses import dataclass, field from datetime import datetime, timedelta from typing import Optional @dataclass class CostAlert: """Non-blocking cost alert that does not consume budget tokens.""" timestamp: datetime threshold_pct: float message: str recommended_action: str @dataclass class CostMonitor: """Real-time cost monitoring with configurable threshold alerts. Alerts are sent to a callback queue and never block the agent execution path. Monitoring overhead is kept below 0.1% of total compute budget. """ alert_thresholds: list[float] = field(default_factory=lambda: [50.0, 75.0, 90.0, 100.0]) alerts: list[CostAlert] = field(default_factory=list) _last_alerted_pct: float = 0.0 _lock: threading.Lock = field(default_factory=threading.Lock) def check_and_alert(self, current_tokens: int, max_tokens: int, alert_callback=None): """Check utilization against thresholds and trigger alerts if needed.""" if max_tokens == 0: return utilization_pct = (current_tokens / max_tokens) * 100 for threshold in sorted(self.alert_thresholds): if utilization_pct >= threshold and self._last_alerted_pct < threshold: action_map = { 50.0: "Log cost trend — no action required", 75.0: "Switch to lower-cost model or simplify queries", 90.0: "Disable non-critical operations, activate conservative retry mode", 100.0: "Graceful shutdown: complete in-flight request then stop", } alert = CostAlert( timestamp=datetime.utcnow(), threshold_pct=threshold, message=f"Cost utilization reached {utilization_pct:.1f}% (threshold: {threshold}%)", recommended_action=action_map.get(threshold, "Review budget allocation"), ) self.alerts.append(alert) if alert_callback: alert_callback(alert) self._last_alerted_pct = max(self._last_alerted_pct, utilization_pct) # --- Usage Example --- monitor = CostMonitor( alert_thresholds=[50.0, 75.0, 90.0, 100.0] ) def on_cost_alert(alert: CostAlert): """Non-blocking handler — does not consume agent budget tokens.""" print(f"[{alert.timestamp}] {alert.message}") print(f" Recommended: {alert.recommended_action}") # During agent execution: monitor.check_and_alert(tokens_used, total_budget, alert_callback=on_cost_alert) ``` ### Pattern 3: Budget-Aware Retry Logic Adaptive retry strategies based on remaining budget margin. ```python from enum import IntEnum import random class RetryMode(IntEnum): AGGRESSIVE = 3 # >75% budget remaining — retry up to 3 times MODERATE = 2 # 50-75% remaining — retry up to 2 times CONSERVATIVE = 1 # 25-50% remaining — retry only once NONE = 0 # <25% remaining — no retries class ErrorClassification(IntEnum): RETRYABLE = 1 # Transient: rate limit, timeout, network error TERMINAL = 0 # Permanent: invalid input, schema mismatch, auth failure def classify_error(exception: Exception) -> ErrorClassification: """Classify error as retryable or terminal to avoid wasting budget.""" transient_errors = (TimeoutError, ConnectionError, RateLimitError) return ( ErrorClassification.RETRYABLE if isinstance(exception, transient_errors) else ErrorClassification.TERMINAL ) def adaptive_retry( operation: callable, session_budget: SessionBudget, estimated_cost_per_retry: int = 1_000, max_retries: Optional[int] = None, ) -> object: """Execute operation with budget-aware adaptive retry logic. Args: operation: Callable that performs the LLM/API call. session_budget: Active SessionBudget to check against. estimated_cost_per_retry: Token budget reserved for one retry attempt. max_retries: Override default (derived from remaining budget). Returns: Operation result. Raises: BudgetExhaustedError: If no retries are affordable. FinalError: If all retries exhausted or error classified as terminal. """ # Determine retry mode based on remaining budget if max_retries is None: utilization = session_budget.utilization_pct if utilization < 25: max_retries = RetryMode.NONE elif utilization < 50: max_retries = RetryMode.CONSERVATIVE elif utilization < 75: max_retries = RetryMode.MODERATE else: max_retries = RetryMode.AGGRESSIVE if max_retries == 0: raise BudgetExhaustedError( f"Budget below 25% threshold ({session_budget.utilization_pct:.1f}%). " f"No retries allowed. Complete current operation and shut down gracefully." ) last_exception: Optional[Exception] = None for attempt in range(max_retries + 1): try: return operation() except Exception as exc: error_type = classify_error(exc) if error_type == ErrorClassification.TERMINAL: raise FinalError(f"Terminal error on attempt {attempt + 1}: {exc}") from exc # Check if we can afford another retry before attempting it if attempt < max_retries and not session_budget.can_afford(estimated_cost_per_retry): raise BudgetExhaustedError( f"Cannot afford retry #{attempt + 2}. Remaining budget: " f"{session_budget.tokens_remaining:,} tokens." ) # Add jitter to avoid thundering herd on rate limits jitter = random.uniform(0.1, 1.0) * (2 ** attempt) last_exception = exc import time; time.sleep(jitter) raise FinalError(f"All {max_retries} retries exhausted. Last error: {last_exception}") class RateLimitError(Exception): """Transient error raised when API rate limit is exceeded.""" pass class FinalError(Exception): """Non-retryable error after all retry attempts exhausted.""" pass ``` ### Pattern 4: Multi-Agent Budget Orchestrator Priority-based budget allocation across agent teams sharing a pooled budget. ```python from dataclasses import dataclass, field from enum import IntEnum class AgentPriority(IntEnum): CRITICAL = 0 # Always funded first (e.g., safety monitoring) HIGH = 1 # Funded when budget allows (e.g., user-facing operations) LOW = 2 # Deferred when budget constrained (e.g., background analysis) @dataclass class AgentBudgetAllocation: """Budget allocation for a single agent or agent team.""" agent_id: str priority: AgentPriority max_tokens: int tokens_used: int = 0 @property def utilization_pct(self) -> float: return (self.tokens_used / self.max_tokens) * 100 if self.max_tokens > 0 else 0.0 @dataclass class MultiAgentBudgetOrchestrator: """Priority-based budget allocation across multiple agents sharing a pooled budget.""" total_token_budget: int allocations: list[AgentBudgetAllocation] = field(default_factory=list) tokens_remaining: int = 0 def __post_init__(self): self.tokens_remaining = self.total_token_budget def allocate(self, agent_id: str, max_tokens: int, priority: AgentPriority) -> bool: """Allocate budget to an agent. Returns False if insufficient pooled budget.""" # Check pool availability if max_tokens > self.tokens_remaining: return False self.allocations.append(AgentBudgetAllocation( agent_id=agent_id, priority=priority, max_tokens=max_tokens, )) self.tokens_remaining -= max_tokens return True def get_next_agent(self) -> Optional[AgentBudgetAllocation]: """Get the next agent to run, ordered by priority then allocation size.""" if not self.allocations: return None # Sort by priority (CRITICAL first), then by tokens remaining (largest first) eligible = [ a for a in self.allocations if a.tokens_used < a.max_tokens and self.tokens_remaining > 0 ] if not eligible: return None eligible.sort(key=lambda a: (a.priority, -(a.max_tokens - a.tokens_used))) return eligible[0] def record_usage(self, agent_id: str, tokens_consumed: int) -> bool: """Record token consumption for an agent. Returns False if agent not found.""" for alloc in self.allocations: if alloc.agent_id == agent_id: new_total = alloc.tokens_used + tokens_consumed if new_total > alloc.max_tokens: return False alloc.tokens_used = new_total self.tokens_remaining += tokens_consumed # Refund unused allocation return True return False def utilization_summary(self) -> dict: """Get budget utilization summary across all agents.""" total_allocated = sum(a.max_tokens for a in self.allocations) total_used = sum(a.tokens_used for a in self.allocations) return { "total_budget": self.total_token_budget, "total_allocated": total_allocated, "total_consumed": total_used, "pool_remaining": self.tokens_remaining, "allocation_utilization": (total_used / total_allocated * 100) if total_allocated > 0 else 0, "agents": [ {
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen