nfer

Learn

AI deployment guides

Practical, numbers-first explanations of provider pricing, deployment shapes, and the trade-offs behind the monthly bill.

LLM provider comparison 2026

Twenty providers across API, dedicated, and GPU rental. Cheapest provider per popular model, EU-sovereign options, per-provider profiles, and a buyer's decision tree.

Read guide →

How to deploy open-source LLMs cheaply

API, dedicated, GPU rental, or self-host — pick the cheapest shape with a decision tree, break-even math, and three worked examples.

Read guide →

Self-hosted Llama 3 vs Claude API

Real cost breakdown for self-host vs frontier API: compute, ops, sovereignty, SLA. Three worked scenarios from 5M to 500M tokens/day and a hybrid routing pattern that beats either pure strategy.

Read guide →

How we modelled the inference market

Why 'what does it cost to run model X?' has no single answer — and how nfer makes per-token APIs, GPU rentals, and provisioned throughput honestly comparable.

Read guide →

Popular models

  • DeepSeek-V3.2
  • Llama-3.1-8B-Instruct
  • DeepSeek-R1
  • Mistral-7B-Instruct-v0.3
  • Qwen3-Coder-30B-A3B-Instruct
  • gemma-3-4b-it
  • See all models →

Providers

  • Nebius
  • AWS
  • Azure
  • Verda
  • CoreWeave
  • OVHcloud

Learn

  • All guides →
  • LLM provider comparison 2026
  • How to deploy open-source LLMs cheaply
  • Self-hosted Llama 3 vs Claude API
  • How we modelled the inference market
  • Methodology

About

  • About
  • FAQ
  • Book a consultation
nfer© 2026 nfer