3AM Marketer
AI

Ollama Local LLM Architecture and REST API Guide

A technical reference on running open models locally with Ollama, covering memory hardware dynamics, daemon configuration, REST API endpoints, and nanosecond timing metrics.

Ollama Local LLM Architecture and REST API Guide

Ollama executes open models locally using Georgi Gerganov's llama.cpp backend, processing inference without sending prompts off the host machine. Running local models through Ollama is completely free, relying on automatic hardware detection to orchestrate weights across GPU VRAM or system RAM.123

Hardware Memory Management and 5GB to 10GB Model Storage

Local execution relies on Georgi Gerganov's llama.cpp project as the underlying backend engine. When Ollama initializes, it automatically detects available GPU hardware and adjusts execution settings to match your host resources. A standard 7B parameter model stored in .gguf or .safetensors formats typically occupies between 5GB and 10GB of drive space.13

Memory placement dictates token throughput. If model weights exceed dedicated GPU VRAM, Ollama spills execution over into system RAM, resulting in an immediate decrease in generation speed. In contrast, Apple Silicon M1, M2, and M3 architectures utilize unified memory shared directly between the CPU and GPU, avoiding the latency penalty of system RAM bus spillover.3

Installation Workflows and Port 11434 Daemon Configuration

Installation binaries differ across host platforms. Linux and macOS environments install Ollama using a curl command targeting an install script hosted on ollama.com. Windows systems run an installation command using a PowerShell script, while containerized environments pull the official image distributed under ollama/ollama on Docker Hub.1

Once installed, the CLI command 'ollama serve' starts the local server process. By default, the Ollama API daemon listens locally on port 11434. To expose the HTTP service across a local network or container bridge, setting the environment variable OLLAMA_HOST=0.0.0.0:11434 configures the server to bind to all network interfaces.4

CLI Workflows: Gemma 4 and Model Tag Resolution

Acquiring model weights begins with the command 'ollama pull <model-name>'. Model identifiers follow a model:tag naming structure, such as downloading gemma3:4b. If you omit the tag during acquisition or invocation, Ollama automatically resolves the tag to latest. For interactive command-line sessions, running 'ollama run gemma4' downloads missing layers and immediately starts an interactive chat interface.145

The command line interface also provides preconfigured integration targets for external tools. Running 'ollama launch claude' initializes local integration pathways for Claude Code, while 'ollama launch openclaw' configures an assistant workflow to connect local models across external messaging channels.1

REST Integration Across /api/generate and /api/chat

External applications communicate with the local server over HTTP via port 11434. Ollama provides first-party client libraries for Python and JavaScript under ollama-python and ollama-js. Direct HTTP requests target two primary endpoints: POST /api/generate for discrete completions and POST /api/chat for multi-turn conversational payloads containing a messages array.145

Endpoints stream JSON tokens by default. Supplying stream set to false stops streaming and returns a single JSON object once processing finishes. For raw completion calls, passing raw set to true disables prompt formatting templates, sending input tokens directly to the backend.5

Advanced Parameters: 5-Minute keep_alive and JSON Schemas

Advanced parameters grant granular control over resource management and output formats. The keep_alive parameter controls how long model weights remain loaded in system or GPU memory after an API request, defaulting to 5 minutes. Passing a value to keep_alive prevents continuous model reloading latency between frequent automated tasks.5

Output formatting can be constrained at runtime. Setting format to 'json' activates JSON mode, forcing output to form a valid JSON object. Alternatively, developers can supply an explicit JSON schema inside the format parameter to guarantee that generated outputs conform strictly to predefined field types. For multimodal models, the images parameter accepts a list of base64-encoded images. Thinking models also accept a think parameter, configured as a boolean or set to 'low', 'medium', 'high', or 'max'.5

Performance Telemetry and Nanosecond Duration Metrics

The final JSON response object returned by POST /api/generate includes telemetry measuring compute stages. All duration metrics are reported in nanoseconds. The payload provides four discrete timing attributes: total_duration for the complete request lifecycle, load_duration for loading model weights into memory, prompt_eval_duration for processing context, and eval_duration for token generation.5

Generation speed is calculated from telemetry fields using exact token counts. Dividing eval_count by eval_duration and multiplying the result by 10^9 produces generation throughput in tokens per second. Monitoring this value exposes when resource constraints cause GPU execution to degrade into slower system RAM.5

Questions

What is the best LLM to use locally?

Selecting a local LLM depends entirely on available memory capacity. A standard 7B parameter model requires 5GB to 10GB of storage and dedicated VRAM to run at full speed without spilling into system RAM.

What are the disadvantages of Ollama?

Execution speed drops sharply whenever weights exceed dedicated GPU VRAM and spill into system RAM. Performance is constrained by host hardware bandwidth rather than external cloud compute.

Is Ollama like chatgpt?

Ollama is a local execution engine rather than a centralized cloud service. It runs open-weight models directly on your hardware without sending prompts off the host machine.

Is there a way to run LLM locally?

Yes, Ollama installs via terminal commands across macOS, Linux, and Windows to run quantized models locally. It exposes a local HTTP daemon listening on port 11434 for interactive CLI and programmatic API calls.