Local LLM Setup
Robota works with any local inference server that speaks the OpenAI-compatible API. This guide
covers Ollama, LM Studio and the llama.cpp server. In Robota these all use the same
provider type, gemma, shown in setup as Ollama / LM Studio / llama.cpp.
No API key required. Local models run entirely on your machine. Your code, prompts and conversation history stay on your device.
Quick Start
-
Start your local model server (see the options below).
-
Configure the
robotaCLI to use it:robota --configureChoose Ollama / LM Studio / llama.cpp (gemma), then answer the three prompts. Press Enter to accept a default:
Prompt Default What to enter Base URL http://localhost:1234/v1Your server's URL, including /v1Model supergemma4-26b-uncensored-v2The model name exactly as your server lists it API key lm-studioAny value; local servers do not check it
On a first run with no provider configured, robota offers the same setup: answer "No — use a local
model", and after a short LM Studio guide it asks these three prompts.
To configure without prompts (for example in a setup script):
robota --configure-provider local --type gemma \
--base-url http://localhost:11434/v1 --model llama3.2 --api-key ollama --set-currentThe profile is saved in ~/.robota/settings.json, so you configure it once. Run robota --configure
again, or use /provider inside a session, to change the provider, URL or model.
Option 1: Ollama
Ollama downloads and serves models and exposes an OpenAI-compatible API.
Install and start
# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh
# Pull a model (examples)
ollama pull llama3.2 # small and fast
ollama pull qwen2.5-coder # strong at TypeScript/Python
# Verify the server is running
curl http://localhost:11434/v1/modelsOllama listens on http://localhost:11434.
Configure Robota
Run robota --configure, choose Ollama / LM Studio / llama.cpp (gemma), and enter:
- Base URL:
http://localhost:11434/v1 - Model: the model name exactly as
ollama listshows it (e.g.llama3.2) - API key: any value (e.g.
ollama)
Models for coding
| Model | Size | Best for |
|---|---|---|
qwen2.5-coder:7b | 7B | TypeScript, Python, code review |
codellama:13b | 13B | General code generation |
llama3.2:3b | 3B | Fast responses, simple tasks |
deepseek-coder-v2:16b | 16B | Complex reasoning, refactoring |
Larger models produce better results but need more memory and run slower.
Option 2: LM Studio
LM Studio is a desktop app for downloading and running models, with a built-in local API server.
Install and start
- Download LM Studio from lmstudio.ai.
- Search for and download a model (for example a Gemma, Llama or Qwen Coder model).
- Start the local server from the Developer tab.
The server runs on http://localhost:1234 by default, which is also Robota's default base URL.
Configure Robota
Run robota --configure, choose Ollama / LM Studio / llama.cpp (gemma), and enter:
- Base URL:
http://localhost:1234/v1(the default) - Model: the model name exactly as LM Studio shows it for the loaded model
- API key:
lm-studio(the default)
Option 3: llama.cpp server
If you build and run llama.cpp yourself:
./llama-server -m models/your-model.gguf --port 8080Then run robota --configure, choose Ollama / LM Studio / llama.cpp (gemma), and enter
http://localhost:8080/v1 as the base URL and any value as the API key.
Using a local model from code
The same provider is available to your own code as GemmaProvider:
import { GemmaProvider } from '@robota-sdk/agent-provider-openai-compatible';
const provider = new GemmaProvider({
apiKey: 'ollama', // not checked by local servers
baseURL: 'http://localhost:11434/v1',
defaultModel: 'llama3.2',
});See Providers for its options.
Troubleshooting
"Connection refused" or "Network error"
- Check the server is running:
curl http://localhost:11434/v1/models(Ollama) orcurl http://localhost:1234/v1/models(LM Studio). - Check the port in your base URL matches the server.
- Some servers bind to
127.0.0.1only; tryhttp://127.0.0.1:<port>/v1.
Model not responding
- The model name in your profile must match the model loaded in your server exactly.
- For Ollama,
ollama listshows the names; LM Studio shows the name of the loaded model.
Slow responses
- Local models are slower than hosted APIs, especially on CPU.
- Smaller quantized models (e.g.
q4_k_mvariants) run faster. - Enable GPU acceleration in Ollama or LM Studio if available.
Tool calling not working
Robota's agent works through tool calls. Some local models do not support the OpenAI
function-calling format, and tool calls then fail or never happen. Use a model whose documentation
says it supports tool or function calling (for example qwen2.5-coder or llama3.2).
Tips
- Context windows are smaller. Many local models have far smaller context windows than hosted
models. Use
/compactwhen the context fills up. - Quality varies with size. Small models handle straightforward tasks; use larger ones for complex refactoring.
- Offline. Once the model is downloaded, no internet connection is needed.