The problem. Getting an AI model to reply is easy. Running it as a real service — without it getting abused, breaking, or burning a huge bill — is the hard part, and almost no one shows you how.
You could ask an AI model how. You'll get a plausible answer — but is it correct? current? actually safe? It won't tell you what it skipped, and it never ran anything.
This cookbook is the opposite. Each recipe is a real AI deployment you run yourself in minutes — already set up the careful way (a token to get in, limits against abuse, timeouts, a locked-down container), with plain-English notes on why each piece is there and honest about what it does not cover. Every recipe is proven to run (verified in CI) — the way people who actually operate these systems harden them, written down so you don't learn it the painful way. No API key or cost to start.
In one line: a vendor-neutral, runnable cookbook for deploying AI workloads — apps, agents, and model serving — with the safeguards built in and the trade-offs made explicit.
Maintained by Nimesha Jinarajadasa. Contributions — including vendor-authored recipes — are welcome; see Contribute a recipe.
Whether you're about to put an AI service online or just want to understand what that safely takes, here's the path — about 10 minutes:
- Run a safe AI service. Four commands (below) start recipe 01 on your machine. Free, no API key.
- See it work. One
curlreturns a reply. - Prove it's safe. The readiness check reports which safeguards are present (
8/8) — the auth, limits, and timeouts that stop abuse and runaway cost. - Understand the choices. Deployment decisions explains each safeguard in plain words, with its limits and what you'd add before going public.
- Make it yours. Start from recipe 01's patterns for your own service, and re-run the readiness check as you adapt it, to catch anything you dropped.
You walk away with a hardened, copyable starting point and a tool to verify a deployment has the safeguards most tutorials skip — the difference between "a model replied" and "a service I can safely put online."
| Workload | Recipe | Status |
|---|---|---|
| AI applications | 01 · Containerized text assistant | Available. App tests, container build/smoke, and non-root/read-only checks pass in CI. Live-provider run still pending. |
| Agent workloads | Durable tasks, retries, and recovery | Planned |
| Model serving | GPU inference, capacity, and overload handling | Planned |
See the roadmap for what's next and how it's prioritized.
Prerequisites: Python 3.12+ (setup/smoke scripts) and Docker with Compose v2.
cd recipes/01-ai-app
python3 scripts/init_local.py
cp .env.example .env # PowerShell: Copy-Item .env.example .env
docker compose up --build -d --wait
python3 scripts/smoke.pyCall it over HTTP — in mock mode (the default) it returns a fixed reply and calls no model, so no API key and no cost:
curl -s http://localhost:8000/api/chat \
-H "Authorization: Bearer $(cat .secrets/app_token)" -H 'Content-Type: application/json' \
-d '{"message": "Hello"}'Then prove its safeguards with the shared readiness tool:
python3 ../../tools/readiness-check/readiness_check.py http://localhost:8000 --token "$(cat .secrets/app_token)"Stop with docker compose down. Live mode, Python-only development, and troubleshooting are in the recipe README.
Recipe 01 is a stateless, single-turn text assistant (an HTTP API) that demonstrates the deployment safeguards every recipe here is held to:
- shared-token authentication; request-body, message, rate, and concurrency limits
- upstream timeouts plus an overall deadline; no hidden retries; bounded provider responses
- separate liveness/readiness; request IDs; logs that omit prompts, answers, and secrets
- non-root, read-only container with dropped capabilities and resource limits
- hash-pinned dependencies, automated tests, and CI with a container smoke test
You can prove them yourself — run the readiness check against the endpoint, or trigger each one by hand in the see-the-safeguards walkthrough — and see how they map to the OWASP LLM Top 10 & API Security.
The scope is stated plainly: this targets a single-instance learning deployment or internal prototype, not a public multi-user service. Deployment decisions pairs each choice with its boundary and next step; the verification record states exactly what was and wasn't tested.
Architecture
flowchart TD
B[API client] --> A[Authentication and input limits]
A --> G[Rate and concurrency admission]
G --> M[Mock response]
G --> H[Bounded HTTP model adapter]
H --> P[Configured model endpoint]
L[Health probes] --> S[Local process and readiness]
The model endpoint belongs to the operator; users cannot choose a URL or model through a request.
The cookbook grows by workload. Anyone can add a recipe — including vendors publishing one for their own tool — held to the same standard: it must actually run, document its failure behavior and limitations, include verification evidence, and carry no marketing. Start from the recipe template; governance explains how recipes are reviewed and why neutrality is protected.
A recipe is a runnable deployment (an API or a CLI) + a README + verification. It ships its own tests and smoke check — that's the primary validation; for HTTP request/response recipes, the shared readiness check also reports its safeguards.
- Docs index — what each document is, and a reading order
- Readiness check — a tool that probes any recipe's endpoint for safeguards
- Recipe 01 — run, deploy, verify, troubleshoot
- See the safeguards yourself · OWASP mapping
- Deployment decisions · Verification record
- Roadmap · Contributing · Governance · Code of Conduct · Security · License: MIT