Skip to content
FA

Fireworks AI

Inference

Official site consultedAI and speech (global tools)

Last verified: 1 October 2026

Fireworks AI offers serverless inference billed per token, dedicated GPU deployments and model training. Over 100 models cover text, vision, audio, image and embeddings, with function calling, structured JSON outputs and batch inference. The API is OpenAI-compatible.

Services
  • Serverless per-token inference
  • Dedicated GPU deployments with autoscaling
  • Training and fine-tuning
  • Function calling and structured JSON outputs
  • Embeddings with reranking, batch inference
Access requirements

The page gives no sign-up link, pricing or free credits.

Sources