3AM Marketer
AI

Self-Hosted AI Agents: GPU Sizing & Runtime Guide

A technical deployment blueprint for self-hosted AI agents covering precise GPU VRAM sizing thresholds and NVIDIA Container Toolkit runtime configurations.

Self-Hosted AI Agents: GPU Sizing & Runtime Guide

Self-hosted AI agents allow organizations to run autonomous workflows where the source code, model selection, and data remain entirely on their own infrastructure. Moving agents to private infrastructure requires sizing GPU VRAM for local models and configuring container runtimes to maintain hardware access.4

Self-Hosted AI Agent Architecture and Tool Orchestration

Self-hosted AI agents commonly utilize vector databases like Weaviate or Qdrant for semantic retrieval, or S3-compatible object storage to prevent data loss from ephemeral container restarts. LangChain and LangGraph are open source under the MIT license, with LangGraph functioning as an orchestration layer for stateful multi-step agents and memory. Flowise Community Edition and Dify self-hosted platforms are both open source under the Apache 2.0 license. Observer AI runs autonomous local AI agents powered by local LLMs via Ollama and executes Python code via a Jupyter server.35

VRAM Requirements: Sizing Model Weights for Local Agent Inference

A consumer RTX 4090 GPU provides 24GB of VRAM. 4-bit quantization using GPTQ or AWQ reduces VRAM footprint to roughly 25%–30% of FP16 requirements, delivering a 65%–70% reduction.1

ModelFP16 VRAM4-bit VRAMHardware Target
Llama 3.1 8B~16GB~5GBRTX 4090 or L4
DeepSeek-R1 8B~16GB~5GBRTX 4090
Mistral 7B~14GB~4.5GBRTX 4090 or RTX 3090
Llama 3.1 70B~140GB~40GB2× A100 80GB (FP16) / A100 80GB or H100 SXM (4-bit)
Qwen 2.5 72B~144GB~42GB2× A100 80GB or H100 SXM

GPU Container Runtime Configuration with NVIDIA Container Toolkit

The nvidia-ctk command reads /etc/nvidia-container-runtime/config.toml by default and requires --in-place or --output to modify the file directly. Setting container runtime candidates via nvidia-ctk config can configure crun as the primary low-level runtime with runc retained as a fallback.2

  • Docker container runtime configuration via nvidia-ctk runtime configure --runtime=docker updates the /etc/docker/daemon.json file on the host.
  • Configuring containerd for Kubernetes using nvidia-ctk creates a drop-in configuration file at /etc/containerd/conf.d/99-nvidia.toml.
  • Configuring CRI-O runtime using nvidia-ctk generates a drop-in file at /etc/crio/conf.d/99-nvidia.toml.
  • For Podman container engines, NVIDIA recommends using CDI for accessing NVIDIA devices inside containers.2

On Linux systems where systemd cgroup drivers are used, a known issue causes containers to lose GPU access when systemctl daemon reload is executed.2

For fully air-gapped environments or local development, you can run the full stack on local hardware. Ollama provides the simplest path for local development and testing, while production environments requiring concurrent request handling and resource management operate better with vLLM. If your deployment requires 70B parameter models, budget for A100 or H100 GPUs rather than consumer hardware. Sizing hardware against model weight requirements and choosing robust container runtimes ensures reliable local agent performance.1

Questions

Can I host my own AI agent?

Yes, you can host your own AI agent using open-source tools such as LangGraph, Flowise, or Observer AI. These platforms execute on local machines or private cloud infrastructure, ensuring your data never leaves your environment.

Can I have my own personal AI agent?

Yes, projects like Observer AI allow individuals to run autonomous agents locally using small LLMs via Ollama. It can execute local code and automate tasks directly on your machine without requiring paid cloud subscriptions.

Is there a self-hosted AI?

Yes, open-source language models such as Llama 3.1, Mistral 7B, and DeepSeek-R1 can be fully self-hosted. Using engines like vLLM or Ollama, these models run entirely on your private GPU hardware.