toolspool
FriendliAI logo

FriendliAI

verifiedPaidAPIfriendli.ai

GPU inference cloud for deploying open-weight and custom AI models with high throughput and pay-per-token pricing.

What it does

FriendliAI is an AI inference cloud built for fast, cost-efficient serving of generative models. It offers pay-per-token model APIs, dedicated GPU endpoints, and self-hosted containers, using kernel-level optimizations like continuous batching and speculative decoding to boost throughput, and provides one-click deployment of hundreds of thousands of Hugging Face models.

Core features

Pay-per-token frontier model APIs
Dedicated GPU endpoints billed per second
One-click deployment of 580,000+ Hugging Face models
Custom and fine-tuned model hosting
Performance optimizations (custom kernels, caching, speculative decoding)
99.99% uptime SLA with geo-distributed infrastructure

Best for

Serving LLMs for coding agents and chatbots
Running semantic search and visual/audio understanding
Deploying custom fine-tuned models at scale
Reducing GPU cost while scaling inference

Pricing

A100 80GB
$2.9/hour
H100 80GB
$3.9/hour
H200 141GB
$4.5/hour
B200 180GB
$8.9/hour