Skip to content
robotadocs

Local LLM Setup

Robota works with any local inference server that speaks the OpenAI-compatible API. This guide covers Ollama, LM Studio and the llama.cpp server. In Robota these all use the same provider type, gemma, shown in setup as Ollama / LM Studio / llama.cpp.

No API key required. Local models run entirely on your machine. Your code, prompts and conversation history stay on your device.


Quick Start

  1. Start your local model server (see the options below).

  2. Configure the robota CLI to use it:

    robota --configure

    Choose Ollama / LM Studio / llama.cpp (gemma), then answer the three prompts. Press Enter to accept a default:

    PromptDefaultWhat to enter
    Base URLhttp://localhost:1234/v1Your server's URL, including /v1
    Modelsupergemma4-26b-uncensored-v2The model name exactly as your server lists it
    API keylm-studioAny value; local servers do not check it

On a first run with no provider configured, robota offers the same setup: answer "No — use a local model", and after a short LM Studio guide it asks these three prompts.

To configure without prompts (for example in a setup script):

robota --configure-provider local --type gemma \
  --base-url http://localhost:11434/v1 --model llama3.2 --api-key ollama --set-current

The profile is saved in ~/.robota/settings.json, so you configure it once. Run robota --configure again, or use /provider inside a session, to change the provider, URL or model.


Option 1: Ollama

Ollama downloads and serves models and exposes an OpenAI-compatible API.

Install and start

# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh
 
# Pull a model (examples)
ollama pull llama3.2          # small and fast
ollama pull qwen2.5-coder     # strong at TypeScript/Python
 
# Verify the server is running
curl http://localhost:11434/v1/models

Ollama listens on http://localhost:11434.

Configure Robota

Run robota --configure, choose Ollama / LM Studio / llama.cpp (gemma), and enter:

  • Base URL: http://localhost:11434/v1
  • Model: the model name exactly as ollama list shows it (e.g. llama3.2)
  • API key: any value (e.g. ollama)

Models for coding

ModelSizeBest for
qwen2.5-coder:7b7BTypeScript, Python, code review
codellama:13b13BGeneral code generation
llama3.2:3b3BFast responses, simple tasks
deepseek-coder-v2:16b16BComplex reasoning, refactoring

Larger models produce better results but need more memory and run slower.


Option 2: LM Studio

LM Studio is a desktop app for downloading and running models, with a built-in local API server.

Install and start

  1. Download LM Studio from lmstudio.ai.
  2. Search for and download a model (for example a Gemma, Llama or Qwen Coder model).
  3. Start the local server from the Developer tab.

The server runs on http://localhost:1234 by default, which is also Robota's default base URL.

Configure Robota

Run robota --configure, choose Ollama / LM Studio / llama.cpp (gemma), and enter:

  • Base URL: http://localhost:1234/v1 (the default)
  • Model: the model name exactly as LM Studio shows it for the loaded model
  • API key: lm-studio (the default)

Option 3: llama.cpp server

If you build and run llama.cpp yourself:

./llama-server -m models/your-model.gguf --port 8080

Then run robota --configure, choose Ollama / LM Studio / llama.cpp (gemma), and enter http://localhost:8080/v1 as the base URL and any value as the API key.


Using a local model from code

The same provider is available to your own code as GemmaProvider:

import { GemmaProvider } from '@robota-sdk/agent-provider-openai-compatible';
 
const provider = new GemmaProvider({
  apiKey: 'ollama', // not checked by local servers
  baseURL: 'http://localhost:11434/v1',
  defaultModel: 'llama3.2',
});

See Providers for its options.


Troubleshooting

"Connection refused" or "Network error"

  • Check the server is running: curl http://localhost:11434/v1/models (Ollama) or curl http://localhost:1234/v1/models (LM Studio).
  • Check the port in your base URL matches the server.
  • Some servers bind to 127.0.0.1 only; try http://127.0.0.1:<port>/v1.

Model not responding

  • The model name in your profile must match the model loaded in your server exactly.
  • For Ollama, ollama list shows the names; LM Studio shows the name of the loaded model.

Slow responses

  • Local models are slower than hosted APIs, especially on CPU.
  • Smaller quantized models (e.g. q4_k_m variants) run faster.
  • Enable GPU acceleration in Ollama or LM Studio if available.

Tool calling not working

Robota's agent works through tool calls. Some local models do not support the OpenAI function-calling format, and tool calls then fail or never happen. Use a model whose documentation says it supports tool or function calling (for example qwen2.5-coder or llama3.2).


Tips

  • Context windows are smaller. Many local models have far smaller context windows than hosted models. Use /compact when the context fills up.
  • Quality varies with size. Small models handle straightforward tasks; use larger ones for complex refactoring.
  • Offline. Once the model is downloaded, no internet connection is needed.