SovAIHub
ModulesSAI-260
SAI-260 table of contents
Assessment4 min readDraft

SAI-260 knowledge check

Verify telemetry, quality, service objectives, incidents, recovery, capacity, cost, and change decisions.

Last content review 2026-08-09Included in SAI-260

How to use this assessment

Complete the questions without referring to the chapter text. Then review the guidance and update your operations package where your answer exposes a gap.

This public knowledge check supports learning. It does not supply universal SLO targets or a production monitoring configuration.

Scenario questions

1. Available but unhealthy

An endpoint responds to every request within its latency target, yet a growing share of answers are unsupported or incorrect. Explain why "available" does not mean "healthy" here.

2. Signal without context

A dashboard shows a single number for "retrieval quality" with no further detail. Explain what context is missing before that number can be trusted.

3. Biased feedback signals

A team uses thumbs-up rates and a model-based grader as their primary quality measure. Explain why both can be misleading without further work.

4. Meaningful objectives

A team defines a service objective by excluding every request that previously failed. Explain why this makes the objective less useful.

5. Fallback needs its own approval

During an outage, traffic is automatically redirected to a secondary model. Explain why this fallback needs separate evaluation and authorization from the primary model.

6. When recovery is complete

After an incident, all affected files and services have been restored. Explain why this alone does not mean recovery is complete.

7. Costs beyond accelerators

A cost report includes only accelerator-hours. Name categories of cost it is missing.

8. Sustained versus theoretical capacity

A capacity plan is built entirely from the hardware vendor's peak throughput specification. Explain why this overstates what the system can reliably deliver.

Answer guidance

  1. Availability measures whether the endpoint answers, not whether the answer is correct, permitted, or safe — an endpoint can be reachable and fast while failing on retrieval support, policy outcomes, or output quality.
  2. The number needs its definition, source, dimensions, unit, sampling method, collection delay, retention, access, known quality limits, owner, and the decision it supports — a raw figure without this context cannot be trusted or acted on correctly.
  3. User approval clicks and model-based grader scores both carry their own biases and blind spots. Without calibration and tracked disagreement or sampling bias, they can systematically over- or under-report real quality.
  4. Silently excluding every difficult or previously failed request can make the objective trivially satisfiable — exclusions need to be stated and justified, with an error budget and a defined response for when it is consumed.
  5. A fallback model or endpoint is a different system serving requests; it changes behavior and risk, so it needs its own evaluation and authorization rather than inheriting the primary path's approval.
  6. Recovery is complete only when the system is authorized, healthy, and producing acceptable, evidenced outcomes — restoring files or services does not by itself confirm rollback, restore, and rebuilt behavior have been validated.
  7. Accelerator-only accounting misses storage, network and transfer, platform licenses, artifact and model lifecycle costs, evaluation, observability, security, facilities, support, operations labor, reserved idle capacity, recovery, and depreciation.
  8. A vendor's peak specification assumes ideal conditions; tested sustained capacity accounts for queueing, memory pressure, failure headroom, and maintenance windows — planning to peak numbers overstates what the system can reliably deliver.

Completion rubric

Mark the operations package complete only when:

  • Health, behavior, policy outcomes, and user outcomes are measured as connected signals, not proxies for each other.
  • Measurement conditions and privacy limits are stated for every signal.
  • Service objectives are meaningful, with exclusions justified rather than hidden.
  • Incidents, recovery, capacity, and cost are traceable to evidence.
  • Alerts lead to owned, tested decisions rather than remaining purely informational.

Completion outcome

SAI-260 is complete when the learner can explain why an available endpoint is not the same as a healthy AI service, and can produce the operations package from the AI operations workshop.