Groq delivers fast, low-cost AI inference with custom LPU chips and GroqCloud for real-time, scalable, enterprise-grade deployments.

Self-serve pay-as-you-go tier for developers and startups. Requires a credit card. Provides higher rate limits, access to the Batch API (50% discount), Prompt Caching (50% discount on cached input tokens), Flex Service Tier, Spend Limits, and Chat Support. Pricing is per-token for LLMs and per-hour or per-character for audio models.

Enterprise tier of GroqCloud API for businesses requiring custom solutions at large scale. Includes everything in the Developer plan plus Custom Models, Regional Endpoint Selection, Performance Tier, Scalable Capacity, Dedicated Support, and LoRA Fine-Tunes. Pricing is custom/negotiated.