跳到正文
原文
Google AI:DEV 作者专属(RSS)· Fejun·· 3 小时前AI 评分53

用 LiteLLM 在 10 分钟内自建私有 LLM 网关教程

Stop Bleeding AI Costs: Self-Host Your Own Private LLM Gateway (LiteLLM) in Under 10 Minutes

AI 导读

作者发布一篇在 Vultr 上 10 分钟内部署 LiteLLM 私有 LLM 网关的教程。网关用统一 OpenAI SDK 结构调用 100+ LLM,支持虚拟密钥与每日花费限额、负载均衡与故障转移,以及基于 Redis 的语义缓存(宣称可省最多 40% API 成本)。

正文

As developers, we’ve all been there: you build an AI-powered feature, deploy it to production, and suddenly get slapped with a massive OpenAI or Anthropic API bill because a user ran a recursive loop, or because your team shared raw API keys across staging environments. 💸

Sharing master API keys is a major security risk, and tracking individual user costs across multiple models (GPT-4o, Claude 3.5 Sonnet, Llama-3) is a management nightmare.

There is a better way. By self-hosting LiteLLM, you can spin up a unified, private LLM Gateway that acts as a proxy between your applications and your AI providers.

With LiteLLM, you get:

  • One Unified API: Call 100+ LLMs using the exact same OpenAI SDK structure.
  • Granular Cost Tracking: Create virtual keys for teams, developers, or users with strict spending limits (e.g., max $5/day).
  • Load Balancing & Failover: Automatically route requests to backup API keys or models if rate limits are hit.
  • Semantic Caching: Cache identical queries in Redis to slash API costs by up to 40%.

In this guide, we'll deploy a production-ready LiteLLM Instance with a beautiful UI Admin Dashboard on a high-speed Vultr High-Performance Cloud Instance in under 10 minutes.


Why Host on Vultr?

When proxying API requests, latency is everything. Adding a proxy layer shouldn't introduce lag. Vultr's High-Performance Cloud offers high-frequency CPU cores and NVMe storage in over 32 global locations. This means your proxy can be co-located right next to your target audience (or your main application servers) to ensure sub-millisecond routing overhead.


Step 1: Provision Your Vultr Instance

  1. Sign up or log into your account via Vultr.
  2. Click Deploy Server and select Cloud Compute.
  3. Choose High Performance (AMD or Intel) to ensure ultra-fast request routing.
  4. Select a server location closest to your main applications.
  5. Choose Ubuntu 24.04 LTS as your Operating System.
  6. Select the 2 GB RAM / 1 vCPU plan (which is more than enough to handle thousands of concurrent proxy requests).
  7. Add your SSH key and click Deploy Now.

Once your instance is ready, copy the IP address and connect to it via your terminal:

ssh root@YOUR_VULTR_IP

Step 2: Install Docker and Docker Compose

Update your system packages and install Docker to manage our containers easily:

sudo apt update && sudo apt upgrade -y
sudo apt install -y docker.io docker-compose

Verify the installation:

docker --version && docker-compose --version

Step 3: Create the LiteLLM Configuration

Create a dedicated directory for LiteLLM:

mkdir litellm-gateway && cd litellm-gateway

Now, create a litellm_config.yaml file. This is where you configure the models you want to proxy. For this demo, we'll configure OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet:

model_list:
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: "os.environ/OPENAI_API_KEY"
  - model_name: claude-3-5-sonnet
    litellm_params:
      model: anthropic/claude-3-5-sonnet-20240620
      api_key: "os.environ/ANTHROPIC_API_KEY"

# Configure global database for key management and logs
router_settings:
  routing_strategy: latency-based-routing

general_settings:
  master_key: "os.environ/LITELLM_MASTER_KEY"

Step 4: Define the Docker Compose Stack

We will set up three components:

  1. LiteLLM Core: The proxy service.
  2. PostgreSQL: Database to store your custom API keys, user budgets, and audit logs.
  3. LiteLLM Admin UI: A clean dashboard to create keys, set limits, and view live metrics.

Create a docker-compose.yml file in the same directory:

version: '3.8'

services:
  db:
    image: postgres:16-alpine
    container_name: litellm-db
    environment:
      POSTGRES_DB: litellm
      POSTGRES_USER: litellm_user
      POSTGRES_PASSWORD: super_secret_db_password
    volumes:
      - pgdata:/var/lib/postgresql/data
    ports:
      - "5432:5432"
    restart: always

  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    container_name: litellm-proxy
    ports:
      - "4000:4000"
    volumes:
      - ./litellm_config.yaml:/app/config.yaml
    environment:
      - DATABASE_URL=postgresql://litellm_user:super_secret_db_password@db:5432/litellm
      - LITELLM_MASTER_KEY=sk-your-super-secure-master-admin-key-12345
      - OPENAI_API_KEY=your_actual_openai_api_key_here
      - ANTHROPIC_API_KEY=your_actual_anthropic_api_key_here
    depends_on:
      - db
    command: ["--config", "/app/config.yaml", "--detailed_debug"]
    restart: always

volumes:
  pgdata:

💡 Note: Replace your_actual_openai_api_key_here and your_actual_anthropic_api_key_here with your real API keys, and change the LITELLM_MASTER_KEY to a secure, random string.


Step 5: Start Your Gateway

Run the Docker stack in detached mode:

docker-compose up -d

Check if everything is running correctly:

docker-compose ps

Step 6: Accessing the Admin Dashboard

Your gateway is now active!

  1. Open your browser and go to http://YOUR_VULTR_IP:4000/ui.
  2. Log in using the LITELLM_MASTER_KEY you configured in your docker-compose.yml (e.g., sk-your-super-secure-master-admin-key-12345).

From this dashboard, you can:

  • Create Virtual Keys for external developers or internal microservices.
  • Set Budgets: Limit a key to a specific dollar amount (e.g., $10 total limit, or resetting daily).
  • View Real-time Analytics: Track exactly which developer, key, or model is consuming the most tokens.

How to Use Your Private Gateway in Code

To consume your self-hosted gateway, you simply point your existing OpenAI SDK to your Vultr server. No code changes required!

Python Example:

from openai import OpenAI

client = OpenAI(
    api_key="sk-the-virtual-key-you-created-in-dashboard",
    base_url="http://YOUR_VULTR_IP:4000"
)

# Call any model you defined in your LiteLLM config!
response = client.chat.completions.create(
    model="claude-3-5-sonnet",
    messages=[{"role": "user", "content": "Explain quantum computing in 2 sentences."}]
)

print(response.choices[0].message.content)

Secure Your Setup for Production

Before pointing production applications to your gateway, ensure you secure it with SSL. You can easily install Nginx and Certbot on your Vultr instance to get a free Let's Encrypt SSL certificate:

sudo apt install -y nginx certbot python3-certbot-nginx

Configure Nginx to reverse proxy traffic from port 80/443 to http://localhost:4000 and run sudo certbot --nginx to enable HTTPS.

Ready to get your team's AI costs under lock and key? Build your high-performance, private LLM gateway today on Vultr High-Performance Cloud!


Liked this resource? Join our daily Telegram channel for more developer tools and cloud insights: @Libretech2026

来源:Google AI:DEV 作者专属(RSS) · dev.to