Ollama Agent — Your Autonomous Local AI Assistant
Ollama Agent is an autonomous, local-first AI assistant running entirely on your machine via Ollama. It provides an interactive terminal user interface (REPL) and a one-shot command-line interface (CLI) to tackle daily tasks—from deep research, document analysis, and system automation to software development. With zero cloud dependency, no subscriptions, and total data privacy, you stay in complete control of your tools, your prompts, and your data.
Quick Installation
Install Ollama Agent in an isolated Python environment using pipx (recommended) or standard pip:
=== "pipx (Recommended)"
```bash
# Install globally in an isolated environment
pipx install git+https://github.com/arrase/ollama-agent.git
# Upgrade to the latest release
pipx upgrade ollama-agent
```
=== "pip (Virtual Environment)"
```bash
# Install into your active Python 3.11+ virtual environment
pip install git+https://github.com/arrase/ollama-agent.git
# Upgrade existing installation
pip install --upgrade git+https://github.com/arrase/ollama-agent.git
```
Prerequisites
Before launching Ollama Agent, ensure the following are available on your system:
- Python 3.11+: Verify with
python3 --version. - Ollama: Downloaded, installed, and running (
ollama serve). - Tool-Calling Model: A model with tool/function-calling capabilities: (If unconfigured, Ollama Agent automatically scans your installed Ollama models and presents an interactive selector).
- Embeddings Model (Optional, for Local RAG):
60-Second Quick Start
1. Launch the Interactive REPL
Start the full-featured terminal workspace with live markdown streaming, context tracking, and command autocompletion:
2. Run a One-Off CLI Prompt
Execute a task directly from your shell and stream the answer to standard output:
3. Combine Advanced Flags
Specify a dedicated model, reasoning effort, and autonomous execution (YOLO mode):
Common CLI Flags at a Glance
-p, --prompt: Run in non-interactive single-shot mode and exit when done.-m, --model: Select any installed Ollama model for the session.-e, --effort: Set reasoning effort (dynamically matched to model-supported values, e.g.low,high,max,default).-y, --yolo: Enable YOLO mode (runs tools autonomously without approval prompts).-s, --stealth: Run in-memory without saving conversation history to disk.
Why Ollama Agent? (The Local Advantage)
Most agentic frameworks treat Ollama as a generic OpenAI-compatible proxy. This often causes truncated outputs, missed tool calls, and lost context. Ollama Agent is built natively for Ollama, harnessing the full potential of local LLMs:
| Capability | Generic OpenAI Proxy Agents | Ollama Agent |
|---|---|---|
| Data Privacy | Frequently routes data to third-party endpoints or telemetry servers. | 100% Local & Private. Prompts, code, and documents never leave your machine. |
Context Window (num_ctx) |
Defaults to Ollama's 2,048-token limit, causing premature amnesia. | Auto-detects maximum context directly from model metadata (or up to 128k+). |
| Model Hyperparameters | Forces hardcoded defaults (temp=0.7), ignoring creator recommendations. |
Auto-discovers creator settings from Modelfiles (top_k, min_p, repeat_penalty). |
| Token Accuracy | Relies on inaccurate tiktoken approximations designed for GPT models. |
Server-native token counting reads exact evaluation metrics from Ollama. |
| Reasoning Traces | Leaks raw <think> tokens into conversation text or fails to parse them. |
API-driven thinking controls: queries /api/show for supported effort levels and defaults, rendering traces into clean collapsible UI blocks. |
| Safety Controls | All-or-nothing execution without fine-grained user confirmation. | Human-in-the-Loop approval before file edits or shell runs, with one-flag YOLO toggle. |
Screenshot Gallery
| Interactive Terminal REPL | Non-Interactive Single-Shot CLI |
|---|---|
![]() |
![]() |
| Full-featured TUI with live streaming markdown, status header, and tool approvals | Streamlined one-shot execution directly in your shell for scripts and CI |
How It Works
Ollama Agent operates with a transparent, user-centered execution loop designed to keep you informed and in control:
flowchart LR
User([User Prompt]) --> Agent{Ollama Agent}
Agent --> Model[Local Model via Ollama]
Model --> Plan[Analyze & Select Tools]
Plan --> Safety{Safety Gate}
Safety -- "HITL Approval" --> Confirm[User Confirms Action]
Safety -- "YOLO Flag (-y)" --> RunTool[Run Tool Directly]
Confirm --> RunTool
RunTool --> Execute[Files / Shell / RAG / MCP]
Execute --> Result[Streamed Markdown Response]
Result --> Memory[(SQLite History & Memory)]
- Prompt Ingestion: Type in the interactive REPL, pass a single query with
-p, or run a saved task. - Context Resolution: The agent loads project standards (
AGENTS.md), user preferences (MEMORY.md), and relevant history. - Local Reasoning: Your local model evaluates instructions and determines necessary tool actions.
- Safety Verification: Interactive confirmation prompts appear before running shell commands or modifying files, unless you explicitly enable YOLO mode (
-y). - Real-Time Streaming: Tool outputs, reasoning traces, and markdown responses stream live to your screen.
Key Features
Interactive REPL
Rich terminal UI with live syntax-highlighted streaming, multiline editing (\ + Enter), 3-level tab autocompletion, and real-time status updates.
Non-Interactive CLI
Execute single prompts directly from your shell (-p) for automation, scripting, CI/CD pipelines, and quick command-line queries.
Saved Tasks
Store, parameterize, and run repeatable prompt workflows with dynamic variables, dedicated model choices, and custom reasoning levels.
Agent Skills
Equip your assistant with specialized procedural capabilities following the open Agent Skills standard with progressive disclosure.
MCP Tool Support
Connect seamlessly to any external Model Context Protocol server over stdio, http, or sse with live discovery and management.
Local RAG Engine
Index local documentation and codebases into an embedded vector store powered by Ollama embeddings for instant semantic search.
Memory & AGENTS.md
Persistent user preferences across sessions (MEMORY.md) and hierarchical auto-discovery of project guidelines (AGENTS.md).
Custom Subagents
Delegate specialized workflows to subagents with isolated context windows, dedicated system prompts, and exclusive MCP toolsets.
Clipboard Integration
Cross-platform clipboard support across Linux (Wayland/X11), macOS, and Windows to quickly pull snippets into prompts or copy responses.
Multi-Language UI
Fully localized terminal experience across 16 languages with automatic system locale detection and explicit runtime overrides (-l).
Practical Everyday Recipes
1. Codebase Audit & Refactoring
Inspect project files and refactor legacy patterns while keeping changes local:
ollama-agent -m "qwen2.5-coder:14b" -e high -p "Audit src/auth.py for insecure token storage and propose fixes."
2. Interactive @-Mention Attachments
In the interactive REPL, attach local files, directories, or media directly using @:
3. Local Knowledge Base (RAG) Querying
Create a local semantic search database for your documentation and query it on demand:
# Index local docs into a collection
ollama-agent rag create project-docs ./docs --model nomic-embed-text
# Query the collection via CLI
ollama-agent --rag project-docs -p "How do I configure subagents in settings.yaml?"
4. Running Reusable Tasks with Variables
Run structured, parameterized tasks with dynamic CLI overrides:
5. Instant Clipboard Transformation
Grab clipboard contents, transform them with a local LLM, and output ready-to-use code:
CLI Reference & Options
| Option | Short | Default | Description |
|---|---|---|---|
--prompt <text> |
-p |
None | Runs in non-interactive mode with the given prompt and exits. |
--model <name> |
-m |
settings.yaml |
Name of the local Ollama model to use for the session. |
--effort <level> |
-e |
default |
Reasoning effort (dynamically matched to model-supported values, e.g. low, high, max, default). |
--num-ctx <int\|max> |
-c |
10000 |
Context window size in tokens, or max for model capacity. |
--yolo |
-y |
False |
Autonomous mode: bypasses all tool approval prompts. |
--stealth |
-s |
False |
Ephemeral mode: does not persist chat history to SQLite. |
--rag <collection> |
— | None | Preloads a RAG database collection into the agent's context. |
--language <code> |
-l |
Auto | UI language code (e.g. en, es, fr, de, zh, ja). |
--builtin-tool-timeout |
-t |
30 |
Timeout in seconds for individual tool and command executions. |
--allow-traversal |
— | False |
Permits filesystem tools to access files outside current directory. |
--config-reset <type> |
— | None | Resets configuration files: all, system-prompt, or config-file. |
Pro Tips & Best Practices
1. Match Model Size to Your Task
For programming, code review, and script automation, models like qwen2.5-coder:14b or qwen2.5-coder:32b deliver top-tier tool calling and syntax generation. For general writing, research, and analysis, llama3.1:8b or llama3.1:70b provide broad conceptual reasoning.
2. Drop an AGENTS.md into Your Projects
Create an AGENTS.md file in your project root. Ollama Agent automatically discovers it and adopts your project's coding standards, build commands, and architectural preferences without cluttering every prompt.
3. Use Tab Completion Everywhere
In REPL mode, press Tab to autocomplete slash commands (e.g. /model, /session, /context, /rag), entity arguments, and local files prefixed with @.
4. Automatic Context Compaction
During long troubleshooting or exploratory sessions, Ollama Agent tracks your token usage against your model's context window. When usage exceeds 85%, background compaction summarizes past turns while preserving essential working memory. You can also trigger this manually in the REPL with /context compact.
Documentation Roadmap
Dive deeper into Ollama Agent with our dedicated guides:
User Guides
- CLI & REPL Interface Guide: Master the interactive terminal, slash commands, multiline editor,
@-file attachments, and CLI automation.
Extensibility & Tools
- Saved Tasks & Automation: Build and automate repeatable prompt templates with dynamic Jinja2 variables.
- Agent Skills Standard: Extend capabilities with custom procedural workflows and progressive skill discovery.
- Model Context Protocol (MCP): Connect external tool servers over
stdio,http, orssewith hot-reloading. - Specialized Custom Subagents: Configure dedicated subagents with tailored prompts, models, and isolated contexts.
Knowledge & Memory
- Memory, Sessions & Guidelines: Leverage
AGENTS.mdproject guidelines, persistent user memory (MEMORY.md), and session history. - Local RAG Engine: Build, embed, and query private vector databases from your local files.
Reference & Internals
- Configuration Reference: Complete guide to
settings.yaml, environment variables, and parameter tuning. - System Architecture: In-depth breakdown of state graphs, persistence, tool middleware, and streaming pipelines.
- Developer Guide: Development environment setup, testing procedures, and contribution guidelines.

