| name | model-serving-deployment |
| description | Deploy and operate model inference with explicit batching, concurrency, memory, routing, health, rollout, fallback, and compatibility behavior. |
Model Serving Deployment
Use when this procedure is the primary professional method needed for the assignment.
Procedure
- Confirm the decision or outcome this work must support, its scope, owner, constraints, and definition of success.
- Establish the evidence baseline using model size, tokenizer/config, traffic shape, GPU/CPU capacity, latency SLO, and serving runtime. Do not fill material gaps with assumptions when they can change the result.
- Benchmark representative load, size replicas/accelerators, configure timeouts/batching, expose health/metrics, canary rollout, and validate failure/fallback.
- Exercise realistic edge, failure, transition, or exception cases that could invalidate the result; record unresolved uncertainty explicitly.
- Validate the output against the original outcome and any neighboring professional contracts so this skill does not silently absorb another specialist's authority.
- Record the resulting artifact, measurements, decisions, provenance, and handoff information needed for another owner to reproduce or continue the work.
Quality gate
Serving meets representative quality/latency/throughput goals and has a proven rollback/fallback path.