This feature allows GitHub Copilot Chat and Agent modes to route queries to locally running LLMs (such as LM Studio, Ollama and Docker Model Runner) instead of GitHub’s cloud models. We will use Docker Model Runner with llama.cpp inference engine for our setup.
Docker Model Runner¶
An integrated feature within Docker Desktop that manages and runs AI models directly via container workflows.
Key features¶
OpenAI-compliant local HTTP API.
Handles models as OCI artifacts on standard registries
Automatic hardware/GPU handling
Automatic on demand loading / unloading of models
🏃 Inference engines (AKA Runners)
🐳 Docker Model Runner in a nutshell
➡️ Docker Model Runner on Linux
Local LLM at your service¶
1️⃣ Choosing the right model...
If like me, you want to run the local LLM on your laptop with limited GPU VRAM (mine is an NVIDIA GeForce RTX 3050 4GB), llama.cpp is your only choice. It uses the GGUF format, which bakes quantization directly into the weights so a model can shrink from 16-bit floats down to roughly 4 bits per weight with minimal quality loss. The goal is to pick a quantized model whose weights and KV cache fit entirely (or mostly) inside the GPU VRAM to maintain fast tokens-per-second generation. The top picks for coding would be ai/qwen2.5 3B or ai/qwen3.5 4B. And as you can see in the following table, Q4_K_M is the sweet spot:
| Quantization | Bits/weight | Memory | Quality |
|---|---|---|---|
Q4_K_M | ~4.5 | Low | Good |
Q5_K_M | ~5.5 | Moderate | Excellent |
Q8_0 | 8 | High | Near-original |
F16 | 16 | Highest | Original |
2️⃣ Bringing up the endpoint using Docker Model Runner...
Verify NVIDIA GPU & CUDA (if applicable)
nvidia-smiCheck the status of the model runner:
docker model statusIf not enabled, enable it via the GUI Settings AI Enable Docker Model Runner, or run the following command (on Windows, the command line approach only works in PowerShell, not WSL):
docker desktop enable model-runnerRunners are installed automatically, but if you want to have more control you can install them manually. For example:
docker model install-runner --backend llama.cpp --gpu cudaQuery the local OpenAI-compatible endpoint:
curl http://localhost:12434/v1/modelsSearch for your model name in the Docker AI registry. For example:
docker model search qwenIf necessary visit Docker Hub to find the right tagged version.
Pull the weights and test your model. For example
# docker model pull ai/qwen2.5:3B-Q4_K_M
# docker model pull ai/llama3.2:3B-Q4_K_M
docker model pull ai/qwen3.5:4b-q4_K_M
docker model ls
docker model run qwen3.5:4b-q4_K_M "Are you suitable for coding?"
docker model ps9. Change the default context size
docker model configure --context-size 32768 qwen3.5:4b-q4_K_MIn current builds of Docker Desktop, the above command only modifies the runtime context size in active daemon memory; it does not write the change back into the model’s immutable OCI artifact manifest. Consequently, whenever Docker Desktop restarts, the backend resets to the default (typically 4,096 tokens).
So, you have to run the above command after every reboot. You have been warned!
3️⃣ Connecting the endpoint to GitHub Copilot in VS Code...
Once the runner is serving the API on localhost:
In VS Code, open the Command Palette (Ctrl+Shift+P / Cmd+Shift+P or simply F1).
Search for
Chat: Manage Language Models.Select Add Models Custom Endpoint.
Enter
Local LLMfor Group Name.Enter
dummyfor API.Select
Chat Completionsfor API Type.In the opened JSON file, update
"id"and"name"with the values returned for each model bydocker model lscommand.Update
"url"in the same JSON file withhttp://localhost:12434/v1for each model you want.Close the JSON file to save it.
You can now select the models provided by the Local LLM directly inside the Copilot Chat / Agent dropdown.
✨ Instruction for the AI agent
Add support for writing images in TGA file format to stb.hpp using stbi_write_tga() function that's already included in stb_image_write.hPros & Cons¶
Pros
✅ Free
✅ Data privacy
✅ No network dependency
Cons
❌ Slow especially on CPU
❌ Limited context size
❌ Mostly for small problems