Skip to content

Ollama Agent — Your Autonomous Local AI Assistant

License Python 3.11+ DeepAgents LangChain Ollama

Ollama Agent is an autonomous, local-first AI assistant running entirely on your machine via Ollama. It provides an interactive terminal user interface (REPL) and a one-shot command-line interface (CLI) to tackle daily tasks—from deep research, document analysis, and system automation to software development. With zero cloud dependency, no subscriptions, and total data privacy, you stay in complete control of your tools, your prompts, and your data.


Quick Installation

Install Ollama Agent in an isolated Python environment using pipx (recommended) or standard pip:

=== "pipx (Recommended)"

```bash
# Install globally in an isolated environment
pipx install git+https://github.com/arrase/ollama-agent.git

# Upgrade to the latest release
pipx upgrade ollama-agent
```

=== "pip (Virtual Environment)"

```bash
# Install into your active Python 3.11+ virtual environment
pip install git+https://github.com/arrase/ollama-agent.git

# Upgrade existing installation
pip install --upgrade git+https://github.com/arrase/ollama-agent.git
```

Prerequisites

Before launching Ollama Agent, ensure the following are available on your system:

  1. Python 3.11+: Verify with python3 --version.
  2. Ollama: Downloaded, installed, and running (ollama serve).
  3. Tool-Calling Model: A model with tool/function-calling capabilities:
    ollama pull qwen2.5-coder:14b
    # or: ollama pull llama3.1:8b
    
    (If unconfigured, Ollama Agent automatically scans your installed Ollama models and presents an interactive selector).
  4. Embeddings Model (Optional, for Local RAG):
    ollama pull nomic-embed-text
    

60-Second Quick Start

1. Launch the Interactive REPL

Start the full-featured terminal workspace with live markdown streaming, context tracking, and command autocompletion:

ollama-agent

2. Run a One-Off CLI Prompt

Execute a task directly from your shell and stream the answer to standard output:

ollama-agent -p "Summarize git commits made in the last 7 days."

3. Combine Advanced Flags

Specify a dedicated model, reasoning effort, and autonomous execution (YOLO mode):

ollama-agent -m "qwen2.5-coder:14b" -e high -y -p "Audit src/auth.py for security vulnerabilities."

Common CLI Flags at a Glance

  • -p, --prompt: Run in non-interactive single-shot mode and exit when done.
  • -m, --model: Select any installed Ollama model for the session.
  • -e, --effort: Set reasoning effort (dynamically matched to model-supported values, e.g. low, high, max, default).
  • -y, --yolo: Enable YOLO mode (runs tools autonomously without approval prompts).
  • -s, --stealth: Run in-memory without saving conversation history to disk.

Why Ollama Agent? (The Local Advantage)

Most agentic frameworks treat Ollama as a generic OpenAI-compatible proxy. This often causes truncated outputs, missed tool calls, and lost context. Ollama Agent is built natively for Ollama, harnessing the full potential of local LLMs:

Capability Generic OpenAI Proxy Agents Ollama Agent
Data Privacy Frequently routes data to third-party endpoints or telemetry servers. 100% Local & Private. Prompts, code, and documents never leave your machine.
Context Window (num_ctx) Defaults to Ollama's 2,048-token limit, causing premature amnesia. Auto-detects maximum context directly from model metadata (or up to 128k+).
Model Hyperparameters Forces hardcoded defaults (temp=0.7), ignoring creator recommendations. Auto-discovers creator settings from Modelfiles (top_k, min_p, repeat_penalty).
Token Accuracy Relies on inaccurate tiktoken approximations designed for GPT models. Server-native token counting reads exact evaluation metrics from Ollama.
Reasoning Traces Leaks raw <think> tokens into conversation text or fails to parse them. API-driven thinking controls: queries /api/show for supported effort levels and defaults, rendering traces into clean collapsible UI blocks.
Safety Controls All-or-nothing execution without fine-grained user confirmation. Human-in-the-Loop approval before file edits or shell runs, with one-flag YOLO toggle.

Interactive Terminal REPL Non-Interactive Single-Shot CLI
Interactive REPL UI Non-Interactive CLI
Full-featured TUI with live streaming markdown, status header, and tool approvals Streamlined one-shot execution directly in your shell for scripts and CI

How It Works

Ollama Agent operates with a transparent, user-centered execution loop designed to keep you informed and in control:

flowchart LR
    User([User Prompt]) --> Agent{Ollama Agent}
    Agent --> Model[Local Model via Ollama]
    Model --> Plan[Analyze & Select Tools]
    Plan --> Safety{Safety Gate}
    Safety -- "HITL Approval" --> Confirm[User Confirms Action]
    Safety -- "YOLO Flag (-y)" --> RunTool[Run Tool Directly]
    Confirm --> RunTool
    RunTool --> Execute[Files / Shell / RAG / MCP]
    Execute --> Result[Streamed Markdown Response]
    Result --> Memory[(SQLite History & Memory)]
  1. Prompt Ingestion: Type in the interactive REPL, pass a single query with -p, or run a saved task.
  2. Context Resolution: The agent loads project standards (AGENTS.md), user preferences (MEMORY.md), and relevant history.
  3. Local Reasoning: Your local model evaluates instructions and determines necessary tool actions.
  4. Safety Verification: Interactive confirmation prompts appear before running shell commands or modifying files, unless you explicitly enable YOLO mode (-y).
  5. Real-Time Streaming: Tool outputs, reasoning traces, and markdown responses stream live to your screen.

Key Features

Interactive REPL

Rich terminal UI with live syntax-highlighted streaming, multiline editing (\ + Enter), 3-level tab autocompletion, and real-time status updates.

Non-Interactive CLI

Execute single prompts directly from your shell (-p) for automation, scripting, CI/CD pipelines, and quick command-line queries.

Saved Tasks

Store, parameterize, and run repeatable prompt workflows with dynamic variables, dedicated model choices, and custom reasoning levels.

Agent Skills

Equip your assistant with specialized procedural capabilities following the open Agent Skills standard with progressive disclosure.

MCP Tool Support

Connect seamlessly to any external Model Context Protocol server over stdio, http, or sse with live discovery and management.

Local RAG Engine

Index local documentation and codebases into an embedded vector store powered by Ollama embeddings for instant semantic search.

Memory & AGENTS.md

Persistent user preferences across sessions (MEMORY.md) and hierarchical auto-discovery of project guidelines (AGENTS.md).

Custom Subagents

Delegate specialized workflows to subagents with isolated context windows, dedicated system prompts, and exclusive MCP toolsets.

Clipboard Integration

Cross-platform clipboard support across Linux (Wayland/X11), macOS, and Windows to quickly pull snippets into prompts or copy responses.

Multi-Language UI

Fully localized terminal experience across 16 languages with automatic system locale detection and explicit runtime overrides (-l).


Practical Everyday Recipes

1. Codebase Audit & Refactoring

Inspect project files and refactor legacy patterns while keeping changes local:

ollama-agent -m "qwen2.5-coder:14b" -e high -p "Audit src/auth.py for insecure token storage and propose fixes."

2. Interactive @-Mention Attachments

In the interactive REPL, attach local files, directories, or media directly using @:

>>> Explain how authentication works in @src/auth/jwt.py and compare it to @docs/spec.md

3. Local Knowledge Base (RAG) Querying

Create a local semantic search database for your documentation and query it on demand:

# Index local docs into a collection
ollama-agent rag create project-docs ./docs --model nomic-embed-text

# Query the collection via CLI
ollama-agent --rag project-docs -p "How do I configure subagents in settings.yaml?"

4. Running Reusable Tasks with Variables

Run structured, parameterized tasks with dynamic CLI overrides:

ollama-agent task run code-review target_file=src/api.py strict=true -y

5. Instant Clipboard Transformation

Grab clipboard contents, transform them with a local LLM, and output ready-to-use code:

ollama-agent -p "Read the JSON on my clipboard and convert it into a typed Pydantic model."


CLI Reference & Options

Option Short Default Description
--prompt <text> -p None Runs in non-interactive mode with the given prompt and exits.
--model <name> -m settings.yaml Name of the local Ollama model to use for the session.
--effort <level> -e default Reasoning effort (dynamically matched to model-supported values, e.g. low, high, max, default).
--num-ctx <int\|max> -c 10000 Context window size in tokens, or max for model capacity.
--yolo -y False Autonomous mode: bypasses all tool approval prompts.
--stealth -s False Ephemeral mode: does not persist chat history to SQLite.
--rag <collection> — None Preloads a RAG database collection into the agent's context.
--language <code> -l Auto UI language code (e.g. en, es, fr, de, zh, ja).
--builtin-tool-timeout -t 30 Timeout in seconds for individual tool and command executions.
--allow-traversal — False Permits filesystem tools to access files outside current directory.
--config-reset <type> — None Resets configuration files: all, system-prompt, or config-file.

Pro Tips & Best Practices

1. Match Model Size to Your Task

For programming, code review, and script automation, models like qwen2.5-coder:14b or qwen2.5-coder:32b deliver top-tier tool calling and syntax generation. For general writing, research, and analysis, llama3.1:8b or llama3.1:70b provide broad conceptual reasoning.

2. Drop an AGENTS.md into Your Projects

Create an AGENTS.md file in your project root. Ollama Agent automatically discovers it and adopts your project's coding standards, build commands, and architectural preferences without cluttering every prompt.

3. Use Tab Completion Everywhere

In REPL mode, press Tab to autocomplete slash commands (e.g. /model, /session, /context, /rag), entity arguments, and local files prefixed with @.

4. Automatic Context Compaction

During long troubleshooting or exploratory sessions, Ollama Agent tracks your token usage against your model's context window. When usage exceeds 85%, background compaction summarizes past turns while preserving essential working memory. You can also trigger this manually in the REPL with /context compact.


Documentation Roadmap

Dive deeper into Ollama Agent with our dedicated guides:

User Guides

  • CLI & REPL Interface Guide: Master the interactive terminal, slash commands, multiline editor, @-file attachments, and CLI automation.

Extensibility & Tools

Knowledge & Memory

Reference & Internals

  • Configuration Reference: Complete guide to settings.yaml, environment variables, and parameter tuning.
  • System Architecture: In-depth breakdown of state graphs, persistence, tool middleware, and streaming pipelines.
  • Developer Guide: Development environment setup, testing procedures, and contribution guidelines.