SovAIHub
ModulesSAI-260
SAI-260 table of contents
Concept1 min readDraft

Capacity, cost, and unit economics

Measure constrained capacity, demand, waste, unit cost, and accepted outcomes without optimizing spend in isolation.

Last content review 2026-08-03Included in SAI-260

Connect resources to accepted outcomes

AI cost can be expressed per request, token, model-hour, accelerator-hour, document indexed, evaluation run, supported answer, or accepted task completion. Choose units that reveal decisions; low cost per token can coexist with expensive unusable outcomes.

Cost model

Include compute and accelerators, storage, network and transfer, platform licenses, artifact and model lifecycle, evaluation, observability, security, facilities, support, operations labor, reserved idle capacity, recovery, and depreciation where relevant.

Capacity model

Measure demand distribution, concurrency, context and output lengths, model loading, batching, queueing, memory, hardware availability, failure headroom, maintenance, and growth. Separate theoretical peak, tested sustained capacity, and approved operational capacity.

Optimization guardrails

Evaluate smaller models, quantization, caching, batching, routing, retrieval, prompt reduction, scheduling, and hardware changes against quality, permission, safety, latency, resilience, and sovereignty requirements. Savings are not accepted if they shift unacceptable risk or hidden labor elsewhere.

Report assumptions, measurement conditions, allocation rules, confidence range, and sensitivity. Recalculate after material workload or architecture change.