FriendliAI
GPU inference cloud for deploying open-weight and custom AI models with high throughput and pay-per-token pricing.
What it does
FriendliAI is an AI inference cloud built for fast, cost-efficient serving of generative models. It offers pay-per-token model APIs, dedicated GPU endpoints, and self-hosted containers, using kernel-level optimizations like continuous batching and speculative decoding to boost throughput, and provides one-click deployment of hundreds of thousands of Hugging Face models.
Core features
Pay-per-token frontier model APIs
Dedicated GPU endpoints billed per second
One-click deployment of 580,000+ Hugging Face models
Custom and fine-tuned model hosting
Performance optimizations (custom kernels, caching, speculative decoding)
99.99% uptime SLA with geo-distributed infrastructure
Best for
→Serving LLMs for coding agents and chatbots
→Running semantic search and visual/audio understanding
→Deploying custom fine-tuned models at scale
→Reducing GPU cost while scaling inference
Pricing
A100 80GB
$2.9/hour
H100 80GB
$3.9/hour
H200 141GB
$4.5/hour
B200 180GB
$8.9/hour