- Open-weight models
- Families such as gpt-oss, Qwen, Mistral, Gemma, DeepSeek, Phi and Llama, each checked against its licence. Open-weight does not mean unrestricted: some licences depend on company size or on reselling access.
- Serving
- vLLM, llama.cpp or Ollama on your hardware: NVIDIA, AMD or Apple GPUs, or a CPU for small workloads.
- Private cloud
- Amazon Bedrock with private VPC endpoints, Microsoft Azure AI Foundry with private endpoints, or Google Cloud with VPC Service Controls.
- Cloud models
- OpenAI, Anthropic, Google Gemini or xAI Grok models on paid tiers, with each provider’s training, retention and processing-location terms confirmed in writing.
- Fast inference
- Groq, the inference provider that runs open-weight models on its own chips. Not to be confused with Grok, the model family from xAI.
- Routing and caching
- Rules handle what can be written down, caches handle repeats, and larger models handle only what needs them.
Model names and versions change quickly, so proposals name the exact model, version and licence.