Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

This feature allows GitHub Copilot Chat and Agent modes to route queries to locally running LLMs (such as LM Studio, Ollama and Docker Model Runner) instead of GitHub’s cloud models. We will use Docker Model Runner with llama.cpp inference engine for our setup.

Docker Model Runner

An integrated feature within Docker Desktop that manages and runs AI models directly via container workflows.

Key features

Local LLM at your service

Pros & Cons

Pros

✅ Free
✅ Data privacy
✅ No network dependency

Cons

❌ Slow especially on CPU
❌ Limited context size
❌ Mostly for small problems